One-step visual generation requires efficient supervision that can shape a generator in a single forward pass. Prior distributional training methods match real and generated features in frozen representation spaces, but lack a unifying theory that connects global objectives to pointwise updates. This gap limits understanding of error propagation and hinders scaling to high-fidelity models.

Framework that separates distribution modeling from matching discrepancy

The paper introduces a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. This separation clarifies how matching objectives shape the generator and provides a foundation for new distillation methods. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered as special cases through Gaussian optimal transport and kernel-density-based KL matching, respectively.

MGFlow: Gaussian mixtures at adjustable granularity

The framework motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. This design allows interpolation between coarse global statistics and fine-grained sample-level detail. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates. This coupling addresses mode collapse that mixture expressivity alone does not resolve, because assignment constraints ensure that distinct real features map to distinct generated features.

FD-Loss and Drifting recovered through optimal transport and KL matching

By operating within the Wasserstein gradient flow formalism, the framework naturally produces FD-Loss when distributions are matched via Gaussian optimal transport. When matching is performed with kernel-density estimation, the same formalism yields Gaussian-kernel Drifting. This unifying perspective explains why both methods work and suggests intermediate designs that combine their strengths.

Mass-constrained assignment and paired updates address mode collapse

Mode collapse remains a persistent challenge in one-step generation, even when feature distributions have high expressivity. MGFlow resolves this by coupling mass-constrained sample assignment with paired component updates. The assignment step maps each real feature to a generated feature while preserving total mass, and the paired update step adjusts both components simultaneously. This two-step coupling ensures that the generator learns to produce diverse outputs that cover the real data manifold.

ImageNet 256x256 results: FDr^6 on pMF-H and JiT-H

On ImageNet 256×256, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with FDr^6 of 1.45 on pMF-H and 1.64 on JiT-H. These numbers improve upon the previous baseline and demonstrate that the framework yields generators with both fidelity and diversity. The results hold across two standard evaluation suites, indicating consistent benefit.

Post-training FLUX.2 4B into a one-step generator

For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. This result is notable because it takes a strong existing model and accelerates it to a single generation step without compromising quality. The post-training approach means no retraining from scratch, only a distillation-inspired fine-tuning phase.

Practical use for developer workflows

A working developer can apply the framework to replace multi-step distillation with a single-generation pipeline. The MGFlow post-training recipe offers a path to accelerate existing text-to-image models such as FLUX.2 without collecting new data or redesigning the architecture. The theoretical insights also suggest directions for improving other distributional training methods by examining how distribution modeling and matching discrepancy interact.