Model Distillation

Distillation trains a smaller student model to imitate a larger teacher, typically by matching the teacher’s output distribution (soft labels) instead of only hard labels (Hinton et al., 2015). For LLMs it also means fine-tuning a small model on text generated by a stronger one, including reasoning traces: the DeepSeek-R1 paper (2025) says its reasoning patterns can be used to guide smaller models.

Why it matters

  • Aim: smaller, faster models that keep part of a large model’s behaviour (on-device and low-latency use; see gemma-4); how much is kept varies and is not quantified here.
  • A possible route to open-weight reasoners from a strong teacher (reasoning-models, deepseek-family).
  • Terms-of-service and legal questions arise when the teacher is a proprietary API; treat the policy of each provider as the authority.

Distillation differs from quantization (lower numeric precision of the same model) and from low-rank-adaptation (parameter-efficient fine-tuning). Outputs from a teacher are a form of synthetic-data.

Related: large-language-model, open-source-ai, inference-and-serving.

Sources

  • https://arxiv.org/abs/1503.02531 (Distilling the Knowledge in a Neural Network, 2015)
  • https://arxiv.org/abs/2501.12948 (DeepSeek-R1, includes distilled models, 2025)
  • Opened 2026-10-07: arxiv.org/abs/1503.02531 (Hinton, Vinyals, Dean; submitted 2015-03-09; NIPS 2014 workshop) and arxiv.org/abs/2501.12948 (v1 2025-01-22; abstract says the reasoning patterns can be used to guide smaller models; distilled-model release itself not re-checked on the page).

Open items

  • Provider-specific distillation policies not reviewed.
  • The list of released R1-distilled model sizes was not re-verified (abstract page only).