AI Context Engineering
Context engineering is deciding what goes into the model’s context window at each step of a task. Anthropic defines it as “the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference” and treats it as the successor to writing a single good prompt, because agents run many turns and accumulate tokens (system prompt, tool definitions, tool results, history, retrieved documents).
Where the term comes from
- Tobi Lutke (Shopify) proposed “context engineering” over “prompt engineering” in June 2025 (“the art of providing all the context for the task to be plausibly solvable by the LLM”); Andrej Karpathy endorsed it days later as “the delicate art and science of filling the context window with just the right information for the next step” (X posts, June 2025; secondary summaries quote both).
- Anthropic’s engineering post “Effective context engineering for AI agents” (2025-09-29) made it the standard reference; Anthropic’s docs now say “more context isn’t automatically better” and cite context rot.
Why it matters: context is finite and degrades
Models have a limited attention budget; accuracy and recall fall as token count grows (context-rot-and-long-context). Even 1M-token windows (Claude and Gemini) do not remove the need to curate.
Main techniques
| Technique | Idea | Note |
|---|---|---|
| Selection / just-in-time retrieval | Keep lightweight references (paths, queries) and load data with tools at runtime instead of pre-loading | model-context-protocol, rag-evaluation |
| Compression / compaction | Summarise history near the limit; clear stale tool results | context-compaction-and-memory |
| Isolation / sub-agents | Give each sub-agent a clean window and return condensed results; costs more tokens overall | Anthropic multi-agent research system (see Sources) |
| Caching | Reuse an identical prompt prefix at a discount; order content static-first | prompt-caching |
| Memory | Persist notes outside the window (files, memory tool, knowledge graphs) | context-compaction-and-memory, ai-memory |
| Tool design | Few, clear, token-efficient tools with good descriptions | tool-description-design |
| Standing instructions | CLAUDE.md, AGENTS.md, skills, rules files | claude-md-and-agent-instructions, agents-md |
Related notes in KB-AI
- Claude Code (harness that applies these techniques), claude-agent-sdk
- agents-md, model-context-protocol (under ai-agentic-systems/protocols-and-standards)
- context-graphs-overview (graph-shaped context), how-to-give-your-app-better-context (2024 note on describing code for LLMs; older, pre-dates the agent framing)
- prompting-reasoning-models, structured-outputs, prompt-injection-and-agent-security (untrusted content in context is an attack surface)
- Vendors: anthropic, openai
Sources
- Anthropic, Effective context engineering for AI agents (2025-09-29): https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents (accessed 2026-09-30)
- Anthropic docs, Context windows: https://platform.claude.com/docs/en/build-with-claude/context-windows (2026-09-30)
- Karpathy, X post on context engineering: https://x.com/karpathy/status/1937902205765607626 (only search-result snippet seen, 2026-09-30)
- Anthropic, How we built our multi-agent research system (2025-06-13): https://www.anthropic.com/engineering/multi-agent-research-system (2026-09-30)
- Manus, Context engineering for AI agents: lessons from building Manus (2025-07-18): https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus (2026-09-30)