Diffusion Models

Diffusion models generate data by learning to reverse a gradual noising process: training adds noise to real samples and a network learns to denoise, then generation starts from pure noise and denoises step by step, optionally guided by a text prompt. Denoising Diffusion Probabilistic Models (2020) made the approach practical; latent diffusion (2021) runs the process in a compressed latent space, which made high-resolution image synthesis affordable and underlies Stable Diffusion.

Use

Text-to-image (flux-dot-1, comfyui as a workflow tool), video (google-veo, openai-sora, kling, wan-2-dot-5) and audio are generally described as using diffusion or related methods (flow matching, diffusion transformers), but this was not checked against each model’s own documentation (see Open items); check the model note rather than assuming. Some image models may use autoregressive generation instead.

Related: multimodal-models, world-models, ai-video-generation-tools, ai-slop.

Sources

Open items

  • Which specific video/image models (Veo, Sora, Kling, Wan, FLUX) use diffusion or diffusion transformers was not checked against their own papers/docs; the Use paragraph is general.
  • Claim that diffusion language models are research/early products is unsourced.
  • Share of 2026 image/video models using diffusion vs autoregressive not verified.