Submitted:
03 October 2026
Posted:
08 October 2026
You are already at the latest version
Abstract
Diffusion distillation reduces the number of sampling steps, and its generative images can preserve much of the perceptual quality, but it often degrades Fréchet Inception Distance (FID). Diversity- preserved distribution matching distillation (DP-DMD) attributes the diversity loss of distribution matching distillation (DMD) to reverse KL, and adds a regression loss so that the early inference stage of the student model keeps the low-frequency components of the teacher model. However, recent fast solvers already achieve 5-step inference without modifying the original diffusion model’s parameters, and the original diffusion model’s first-step output is the optimum of that regression loss, so it is unclear whether the regression term should remain only a supplement to reverse KL in DP-DMD. Meanwhile, diversity loss has also been reported for distillation models that do not use reverse KL. We assume and test that using noisy latents from the original diffusion model’s early inference, which is the optimum of that regression loss, can reduce FID whether or not the distillation model uses a reverse-KL objective. EMPURPLE does this with no extra training: the original model generates noisy latents in the early stage, we cache them, and we reuse one of them randomly to start the distilled inference. That algorithm is based on the experiments; a 4-step DDIM inversion of a cached latent, guided by a random caption at low classifier-free guidance (CFG) scale, recovers the noise that can start an inference. 99.92% of the recovered noise lies in the Gaussian typical set, against 99.95% for real Gaussian samples. At low guidance scale, the DDIM forward and inverse maps are nearly identical, so starting the inference of the student from that cached noisy latent is equal to letting the teacher model do one-step early inference, since the experiments’ results show that most of the noisy latent can be generated by an arbitrary prompt if we choose a new noise in the typical set. On COCO 2014 and COCO 2017, EMPURPLE lowers FID for LCM, Flash, SDXL-Turbo, SDXL-Lightning, Hyper-SDXL, and DMD2, including models that do not use reverse KL. On the COCO dataset, the reduction against the official sampler is about 7–19%, and the CLIP score falls in every paired comparison. A tentative PAC argument for this early-stage mismatch is included with the empirical measurements.
Keywords:
diffusion models
; diffusion distillation
; Fréchet inception distance
; mode coverage
; image generation
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.