Reasoning Models and Test-Time Compute
Reasoning models are language models trained to produce a long intermediate chain of thought before the final answer, spending more computation at inference time on harder problems. The idea builds on chain-of-thought prompting (2022) and self-consistency (sampling several chains and voting). Test-time (inference-time) compute is the umbrella term: more thinking tokens, more samples or search at answer time can improve results, sometimes more cheaply than a bigger model (Snell et al., 2024).
How they are trained
Post-training with reinforcement learning on tasks with checkable answers (maths, code) is the approach the DeepSeek-R1 paper describes for incentivising reasoning. DeepSeek-R1 (2025) showed this in the open; see rlhf-and-rlvr. The s1 paper (2025) showed that even a small fine-tune plus a “budget forcing” control of thinking length yields test-time scaling.
In practice
Some model families expose a reasoning/thinking setting or an adaptive mode (for example, see the vault note claude-opus-4-dot-6-adaptive-reasoning; the wider 2026 picture was not verified); see deepseek-family for open-weight reasoners. Commonly cited trade-offs (not sourced here): higher latency and token cost, and different prompting practice: prompting-reasoning-models, chain-of-thought-prompting.
Related: large-language-model, scaling-laws, distillation, multi-token-prediction, ai-benchmarks-and-evals.
Sources
- https://arxiv.org/abs/2201.11903 (Chain-of-Thought Prompting, 2022)
- https://arxiv.org/abs/2203.11171 (Self-Consistency, 2022)
- https://arxiv.org/abs/2408.03314 (Scaling LLM Test-Time Compute Optimally, 2024)
- https://arxiv.org/abs/2501.12948 (DeepSeek-R1, 2025)
- https://arxiv.org/abs/2501.19393 (s1: Simple test-time scaling, 2025)
- Titles, dates and abstracts of all five arXiv papers re-read on arxiv.org/abs pages on 2026-10-07 (s1 abstract: budget forcing by ending thinking early or appending “Wait”; s1K dataset of 1,000 questions).
Open items
- Which 2026 models expose reasoning modes was not verified here (claim “most frontier families” softened); see the model notes. The statement that RL on checkable tasks teaches verifying/backtracking is a general summary of the DeepSeek-R1 abstract and not separately checked.