Reasoning Models and Test-Time Compute

Reasoning models are language models trained to produce a long intermediate chain of thought before the final answer, spending more computation at inference time on harder problems. The idea builds on chain-of-thought prompting (2022) and self-consistency (sampling several chains and voting). Test-time (inference-time) compute is the umbrella term: more thinking tokens, more samples or search at answer time can improve results, sometimes more cheaply than a bigger model (Snell et al., 2024).

How they are trained

Post-training with reinforcement learning on tasks with checkable answers (maths, code) is the approach the DeepSeek-R1 paper describes for incentivising reasoning. DeepSeek-R1 (2025) showed this in the open; see rlhf-and-rlvr. The s1 paper (2025) showed that even a small fine-tune plus a “budget forcing” control of thinking length yields test-time scaling.

In practice

Some model families expose a reasoning/thinking setting or an adaptive mode (for example, see the vault note claude-opus-4-dot-6-adaptive-reasoning; the wider 2026 picture was not verified); see deepseek-family for open-weight reasoners. Commonly cited trade-offs (not sourced here): higher latency and token cost, and different prompting practice: prompting-reasoning-models, chain-of-thought-prompting.

Related: large-language-model, scaling-laws, distillation, multi-token-prediction, ai-benchmarks-and-evals.

Sources

Open items

  • Which 2026 models expose reasoning modes was not verified here (claim “most frontier families” softened); see the model notes. The statement that RL on checkable tasks teaches verifying/backtracking is a general summary of the DeepSeek-R1 abstract and not separately checked.