Modern video diffusion models require many denoising evaluations over long spatiotemporal token sequences, making inference computationally expensive. Distribution Matching Distillation (DMD) was introduced to reduce the number of function evaluations to just a few, but during training DMD samples can degrade, exhibiting progressive oversaturation and artifacts. This paper identifies the root cause: critic errors that enter successive student updates and accumulate over time. The authors introduce Projected Distribution Matching Distillation (PDMD) to filter these critic errors, requiring only a one-line code change to existing DMD.

Prior distillation methods and the critic error problem

Distribution Matching Distillation (DMD) reduces inference cost by matching the distribution of student outputs to those of a critic network. However, the training process is unstable: critic errors propagate through successive student updates and build up over denoising steps. This accumulation leads to degraded sample quality, including unnatural textures and oversaturation. Prior work had not addressed this error accumulation, limiting the practical reliability of distilled video models.

How PDMD projects out critic error

PDMD projects out the component of the DMD update that is parallel to the student-critic endpoint residual. At a fixed noisy query, this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, the projection removes a constant fraction of critic error while discarding only a vanishing fraction of the ideal DMD signal. The mathematical insight is that by filtering the error direction, the student update stays closer to the true distillation target, stabilizing training and improving sample fidelity.

The implementation is a single-line change to the DMD update rule. No additional loss terms, new network components, data reconfiguration, extra model passes, or multi-stage training are required. This minimal change makes PDMD immediately applicable to any DMD-configured video diffusion model.

VBench and VideoGen-Eval results

With the Wan2.1 video diffusion model, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On the MiniMax-H3 joint video-audio benchmark, PDMD attains a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. These results demonstrate that the error projection consistently improves both visual and audio fidelity across two major video generation systems.

Limitations of the projection approach

The projection relies on high-dimensional assumptions that may not hold for all model architectures or token dimensions. The evaluation covers Wan2.1 and MiniMax-H3; other video diffusion systems have not been tested. While audio metrics improve across the board, the extent of improvement may vary with content type and training configuration.

Using PDMD in practice

A working developer can apply PDMD to any DMD-configured video diffusion model with a single code change. No retraining from scratch, no additional data collection, and no new network architecture are needed. The method stabilizes the distillation training trajectory and produces cleaner samples with fewer artifacts. This makes high-quality video generation accessible at lower computational cost, enabling faster iteration on video content pipelines.

Read the paper on arXiv