Context rot and long context

Context rot: accuracy and recall degrade as the number of tokens in the window grows. Anthropic’s docs state that “more context isn’t automatically better” and use the term; its engineering post explains it by the model’s finite attention budget.

What research shows

  • Lost in the Middle (Liu et al., arXiv 2307.03172): performance is often highest when relevant information is at the start or end of the input and degrades when it sits in the middle, even for long-context models.
  • Chroma’s Context Rot report (2025-07-14) tested 18 models (Claude, GPT, Gemini, Qwen): performance dropped with longer inputs even on simple tasks, lower question-answer similarity and distractors made it worse, and behaviour differed by family (Claude models more often abstained; GPT models more often hallucinated). Chroma is a vector-database vendor and its own caveat is that the mechanism is unexplained.
  • Google’s Gemini docs: single-fact retrieval is about 99% accurate, but with multiple “needles” the model “does not perform with the same accuracy”.

What long-context models do and do not solve

  • They let you fit whole codebases or document sets (Claude current models: up to 1M tokens, billed at standard pricing; Gemini: “1 million or more”).
  • They do not guarantee the model uses everything equally; cost and latency still grow with input.

Vendor guidance that follows

  • Anthropic: put long documents at the top, query at the end (up to 30 percent better in their tests); wrap multiple documents in XML tags with source metadata.
  • Google: put the question after the context; use context caching to cut cost (prompt-caching).
  • Reduce what enters the window: retrieval, compaction, sub-agents (context-engineering, context-compaction-and-memory).

Sources