Prompting reasoning models

By 2026 the major vendors’ frontier models reason internally (hidden or summarised thinking). Older advice such as “think step by step” was written for models without this. Below is only what the vendor docs say; the technique notes chain-of-thought-prompting and tree-of-thought-prompting describe the pre-reasoning-model methods.

OpenAI

Reasoning-model guide: keep prompts simple and direct; use delimiters (markdown, XML, section titles); start zero-shot and add examples only if needed; state constraints and success criteria; avoid “think step by step” or “explain your reasoning” because reasoning happens internally; over-engineered prompts can hurt. Use developer messages rather than system messages on current versions. The guide suggests pairing a reasoning model (planning) with a faster GPT model (execution). (The fetched page is written around the o-series; check the current GPT-6 docs for model-specific changes.)

Anthropic (Claude)

  • Current models use adaptive thinking (thinking: {type: "adaptive"}): the model decides per request whether and how much to think. On Fable 5.1, Mythos 5.1, Fable 5, Mythos 5 and Opus 5.5, thinking is always on and adaptive is the only mode. On Claude 4.7 and later, budget_tokens returns a 400 error; older manual budgets remain on Opus/Sonnet 4.6 but are deprecated.
  • Main control is output_config.effort (low, medium, high, xhigh, max; default high on most models, medium on Opus 5.5). Lower effort before trying prompt wording to think less. max_tokens is the hard cap and includes thinking.
  • Steering by prompt works but is wording-sensitive: add “respond directly when in doubt” to reduce thinking, or “this task involves multistep reasoning; think carefully” to increase it. Per-message phrases keep the prompt cache intact, while changing effort invalidates it (prompt-caching).
  • Thinking is billed as output tokens even when display is summarised or omitted. Prefilled assistant turns are not supported from Claude 4.6 on.

Google Gemini

thinking_level (low, medium, high); Gemini 3.8 Flash defaults to medium, models think dynamically by default; thought summaries are optional; thought signatures carry reasoning state across turns. Docs advise low thinking for simple retrieval and maximum for hard coding or maths; thinking tokens are billed.

Practical takeaways

  • Describe the goal, constraints and success criteria; do not script the reasoning steps.
  • Tune effort or thinking level as a cost/latency dial and measure (llm-as-judge-and-evals).
  • Put long inputs first and the question last (context-rot-and-long-context).

Sources