Which parts of modern generative modelling actually matter? Starting from IMLE and adding back only the essentials: we achieve FID 2.56 on ImageNet 256 in a single step.
Which parts of modern generative modelling actually matter? Starting from IMLE and adding back only the essentials: we achieve FID 2.56 on ImageNet 256 in a single step.
Generative models have undergone many generations of evolution, from VAEs/GANs to diffusion/flow matching. Along the way, the underlying techniques have become more complicated and various beliefs about what drives strong empirical performance have taken hold. Due to the success of diffusion models and flow matching, one of the more common beliefs is the importance of transforming the noise distribution to the data distribution gradually through many small transformations.
We ask whether this is truly necessary, and take a minimalist approach to designing a competitive generative model. We start with the bare-bones essentials, namely just a training objective and a model. We purposefully make both simple. For the training objective, we choose Implicit Maximum Likelihood Estimation (IMLE), and eschew more complicated alternatives such as variational inference, adversarial training and numerical integration. For the model, we eschew transformers and instead choose a moderately sized convolutional network. Then we judiciously added elements that are truly essential, which surprisingly do not include iterative denoising. The result is a single-step parameter-efficient generative model that produces high quality samples at fast speed: it achieves an FID of 2.56 on ImageNet 256 and simultaneously attains good precision and recall.
Stochastic interpolant models
Their advantage is commonly attributed to how they decompose generation
Implicit Maximum Likelihood Estimation
Training such a generator directly is generally difficult because the likelihood of the data under it is intractable. IMLE provides a workaround, implicitly maximizing that likelihood without requiring it to be known explicitly. IMLE optimizes a likelihood-based objective akin to regression, without the variational bounds and encoders of VAEs, the adversarial training of GANs, or the iterative numerical integration used by stochastic interpolant models.
Formally, IMLE learns a direct mapping \(f_\theta\) from a prior \(\pi(\mathbf{z})\) to the data distribution \(p_\mathrm{data}(\mathbf{x})\) by minimizing
\[\theta_{\text{IMLE}} = \arg\min_\theta \; \mathbb{E}_{\mathbf{z}_1, \ldots, \mathbf{z}_m \sim \pi(\mathbf{z})} \left[ \sum_{i=1}^{n} \min_{j \in [m]} \; d\left( \mathbf{x}_i, f_\theta(\mathbf{z}_j) \right) \right],\]where \(d(\cdot, \cdot)\) is a distance.
To optimize the objective, the IMLE algorithm alternates between two phases. In the matching phase, $m$ latent codes \(\mathbf{z}_1, \ldots, \mathbf{z}_m\) are sampled from the prior and each real image \(\mathbf{x}_i\) is matched to its nearest generated sample in the set \(\{f_\theta(\mathbf{z}_1), \ldots, f_\theta(\mathbf{z}_m)\}\). This establishes a coupling between \(\pi(\mathbf{z})\) and \(p_\mathrm{data}(\mathbf{x})\). Next, in the optimization phase, the parameters \(\theta\) are updated to bring each generated sample closer to its matched real image.
This paper asks a simple question: is all of the machinery that modern generative models have accumulated, from elaborate architectures
Stochastic interpolant methods are commonly understood to perform iterative denoising. At test time, that network is applied repeatedly to convert pure noise into a generated sample. Consider an alternative view of these models under a deterministic sampler (like DDIM or Flow Matching models). Rather than viewing the model as one instance of the denoising network, we unroll the iterations of the sampler and view the model as their composition, \(h_\theta = g^1_\theta \circ \cdots \circ g^T_\theta\), where \(g^t_\theta\) is the update applied at iteration \(t\). This allows us to view the reverse process as a single compositional model \(h_{\theta}\), that maps a sample from the prior \(\mathbf{x}_T\) to a data sample \(\mathbf{x}_0\). Under this view, each \(\mathbf{x}_{t-1} = g^t_{\theta}(\mathbf{x}_t)\) is an intermediate layer of \(h_{\theta}\).
Read this way, the training objective reveals something specific. SI methods supervise the denoising network’s prediction at every time step, with equal weight applied to all steps, so every intermediate layer of \(h_\theta\) receives a direct training signal. We call this per-stage supervision, and we hypothesize that the sample quality of many-step models stems from it rather than from iterative sampling itself.
The IMLE generator can also be seen as a prior-to-data network composed of stages, only executed in a single forward pass rather than through iterative sampling. We build it as a stack of upsampling blocks, \(f_\theta = f_{\theta^{L}} \circ \cdots \circ f_{\theta^{1}}\), in which each stage operates at a progressively higher resolution. This recasts IMLE in a form similar to the diffusion construction, with each stage \(f_{\theta}^{l}\) now analogous to \(g^t_\theta\), only indexed by resolution rather than noise level.
In existing IMLE methods
For the IMLE algorithm, the model distribution changes during training, so the matching between model samples and ground truth data points needs to be recomputed. Due to randomness in the sampling, there may be no samples drawn from a mode of the model distribution, and as a result, a ground truth data point that can otherwise be explained by that mode would not be well matched. Note that this is different from mode collapse: the poor matching results from randomness in the sampling, not a deficiency in the model distribution.
The standard squared $l_2$ loss lets these mismatches dominate the mean loss. In essence this is a problem of outliers: a small number of large-residual mismatched pairs contribute disproportionately to the mean loss compared to a much larger pool of well-matched pairs.
The classical remedy in robust statistics
We call the resulting method ROMS-IMLE, short for Robust Multi-Stage Implicit Maximum Likelihood Estimation. We ablate our contributions on Oxford Flowers
| Configuration | FID ↓ | Precision ↑ | Recall ↑ |
|---|---|---|---|
| Vanilla IMLE | 62.61 ± 2.28 | 0.45 ± 0.03 | 0.46 ± 0.01 |
| IMLE + Multi-Stage | 31.24 ± 1.08 | 0.79 ± 0.01 | 0.72 ± 0.02 |
| IMLE + Multi-Stage + Geman-McClure (ours) | 13.90 ± 0.08 | 0.93 ± 0.00 | 0.73 ± 0.02 |
Our approach works well in both pixel and latent space.
@article{vashist2026romsimle,
title={ROMS-IMLE: A Minimalist Approach to Competitive Single-Step Generative Modelling},
author={Vashist, Chirag and Li, Ke},
journal={arXiv preprint arXiv:2607.19332},
year={2026}
}