Speculative Decoding

Speculative decoding speeds up LLM generation without changing the output distribution. A small, fast draft model proposes several next tokens; the large target model verifies them in a single parallel forward pass and keeps the longest accepted prefix. Because verification is cheaper than generating the same tokens one at a time, latency drops (Leviathan et al., 2022). Variants add extra prediction heads to the target model itself instead of a separate draft; see multi-token-prediction.

Where it shows up

Serving engines such as vllm implement it (vLLM docs list draft-model, n-gram, suffix, EAGLE and MTP methods as of 2026-10-07), and fast-inference providers (groq, cerebras-inference) compete on latency. It pairs with KV caching and quantization as the standard toolbox of inference-and-serving.

Limits: gain depends on how often the draft agrees with the target; high-temperature or highly creative output accepts fewer drafts.

Related: large-language-model, reasoning-models (long thinking outputs make decoding speed matter more).

Sources

Open items

  • Support in engines other than vLLM, and the Groq/Cerebras latency positioning, were not checked.