Update 2026-10-02

Explicit “think step by step” prompting has diminishing value. Meincke, Mollick et al. (arXiv 2506.07142, 2025-06-08) found CoT gives small average gains on non-reasoning models, adds variability and cost, and only marginal gains on reasoning-capable models. OpenAI’s reasoning best-practices guide says that, since reasoning models reason internally, prompting them to “think step by step” or “explain your reasoning” is unnecessary and may sometimes hinder. Claude’s thinking is adaptive: the model decides per request whether to think, steered mainly by the effort parameter (levels low, medium, high, xhigh, max; medium is the default on Claude Opus 5.5). Gemini 3.x uses thinking_level (low, medium, high) with dynamic thinking by default. See prompting-reasoning-models. The technique still matters for non-reasoning models and small open models.

What it is

Chain-of-thought (CoT) prompting makes a language model write intermediate reasoning steps before the final answer. It is useful for multi-step arithmetic, commonsense and symbolic reasoning, and it makes the answer easier to inspect.

Key papers

  • Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” (arXiv 2201.11903, submitted 2022-01-28): few-shot exemplars containing reasoning steps; a 540B-parameter model with eight exemplars reached state-of-the-art accuracy on GSM8K.
  • Kojima et al., “Large Language Models are Zero-Shot Reasoners” (arXiv 2205.11916, 2022-05-24, NeurIPS 2022): zero-shot CoT by adding “Let’s think step by step”; e.g. MultiArith 17.7% to 78.7% and GSM8K 10.4% to 40.7% with InstructGPT.
  • Wang et al., “Self-Consistency Improves Chain of Thought Reasoning” (arXiv 2203.11171, 2022-03-21, ICLR 2023): sample several reasoning paths and take the majority answer.

Variants and practice

  1. Few-shot CoT (worked examples) versus zero-shot CoT (trigger phrase).
  2. Self-consistency: sample multiple chains, vote.
  3. Iterative refinement: feed the chain back for critique and correction (a general technique; see llm-as-judge-and-evals).
  4. Combining with structure: ask for reasoning in a separate field before the answer when using structured-outputs.
  5. Reasoning models internalise CoT through training and thinking tokens, so the prompt-level technique is largely replaced by effort/thinking-level settings (see the update above).

Caveats

  • Visible reasoning is not guaranteed to be a faithful account of how the model reached its answer (general caveat; no source fetched this session).
  • Extra reasoning tokens cost latency and money; on reasoning models they are billed as output tokens.

Sources