Scaling Laws

Scaling laws are empirical power-law relationships between a model’s loss and the compute, parameters and training data used. Kaplan et al. (2020) found loss falls smoothly as each grows. Hoffmann et al. (2022, “Chinchilla”) argued that earlier models were under-trained and that parameters and training tokens should scale roughly together for a compute-optimal result.

What changed since

  • Beyond compute-optimal: small models are widely reported to be trained on far more tokens than Chinchilla-optimal, with inference cost cited as the reason (not tied to a primary source; see Open items).
  • Post-training and test-time scaling: gains are also reported to come from reinforcement learning and extra inference-time compute (reasoning-models, rlhf-and-rlvr).
  • Data limits: concern about running out of high-quality text is one motivation for synthetic-data (general view, not sourced here).
  • Sparse models: mixture-of-experts decouples total from active parameters (mixture-of-experts-in-agent-swarms covers the agent angle).

Scaling laws are often cited as a rationale for the compute build-out (author’s view): see neocloud and sovereign-ai. Debates over whether scaling leads to artificial-general-intelligence remain unsettled.

Related: large-language-model, generative-pretrained-transformer.

Sources

  • https://arxiv.org/abs/2001.08361 (Scaling Laws for Neural Language Models, 2020)
  • https://arxiv.org/abs/2203.15556 (Training Compute-Optimal LLMs, 2022)
  • https://arxiv.org/abs/2408.03314 (test-time compute scaling, 2024)
  • Opened 2026-10-07: arxiv.org/abs/2001.08361 (Kaplan et al., submitted 2020-01-23; loss follows power laws in model size, data and compute). arxiv.org/abs/2203.15556 and 2408.03314 could not be fetched on 2026-10-07 (fetcher blocked arxiv for those two); they remain as listed from the earlier 2026-10-05 API check.

Open items

  • Statements on current industry training practice (over-training small models) are widely reported but not tied to a primary source here.