Diffusion Models
Diffusion models generate data by learning to reverse a gradual noising process: training adds noise to real samples and a network learns to denoise, then generation starts from pure noise and denoises step by step, optionally guided by a text prompt. Denoising Diffusion Probabilistic Models (2020) made the approach practical; latent diffusion (2021) runs the process in a compressed latent space, which made high-resolution image synthesis affordable and underlies Stable Diffusion.
Use
Text-to-image (flux-dot-1, comfyui as a workflow tool), video (google-veo, openai-sora, kling, wan-2-dot-5) and audio are generally described as using diffusion or related methods (flow matching, diffusion transformers), but this was not checked against each model’s own documentation (see Open items); check the model note rather than assuming. Some image models may use autoregressive generation instead.
Related: multimodal-models, world-models, ai-video-generation-tools, ai-slop.
Sources
- https://arxiv.org/abs/2006.11239 (DDPM, 2020)
- https://arxiv.org/abs/2112.10752 (Latent Diffusion Models, 2021)
- Titles and submission dates confirmed via arXiv API 2026-10-07 (DDPM 2020-06-19; Latent Diffusion 2021-12-20).
Open items
- Which specific video/image models (Veo, Sora, Kling, Wan, FLUX) use diffusion or diffusion transformers was not checked against their own papers/docs; the Use paragraph is general.
- Claim that diffusion language models are research/early products is unsourced.
- Share of 2026 image/video models using diffusion vs autoregressive not verified.