Speculative Decoding
Speculative decoding speeds up LLM generation without changing the output distribution. A small, fast draft model proposes several next tokens; the large target model verifies them in a single parallel forward pass and keeps the longest accepted prefix. Because verification is cheaper than generating the same tokens one at a time, latency drops (Leviathan et al., 2022). Variants add extra prediction heads to the target model itself instead of a separate draft; see multi-token-prediction.
Where it shows up
Serving engines such as vllm implement it (vLLM docs list draft-model, n-gram, suffix, EAGLE and MTP methods as of 2026-10-07), and fast-inference providers (groq, cerebras-inference) compete on latency. It pairs with KV caching and quantization as the standard toolbox of inference-and-serving.
Limits: gain depends on how often the draft agrees with the target; high-temperature or highly creative output accepts fewer drafts.
Related: large-language-model, reasoning-models (long thinking outputs make decoding speed matter more).
Sources
- https://arxiv.org/abs/2211.17192 (Fast Inference from Transformers via Speculative Decoding, 2022)
- https://arxiv.org/abs/2302.01318 (Chen et al., Accelerating LLM Decoding with Speculative Sampling, 2023: the independent DeepMind formulation behind the ‘speculative sampling’ alias)
- https://github.com/vllm-project/vllm/blob/main/docs/features/speculative_decoding/README.md (opened 2026-10-07)
- arXiv pages for 2211.17192 and 2302.01318 opened 2026-10-07: Leviathan, Kalman, Matias, submitted 2022-11-30, ICML 2023 oral; reported 2-3x speedup on T5-XXL with identical outputs.
Open items
- Support in engines other than vLLM, and the Groq/Cerebras latency positioning, were not checked.