Context compaction and memory
Long-running agents outlive one context window. Three complementary mechanisms keep them going; see context-engineering for the wider picture.
Compaction
Summarise the conversation so far and continue from the summary. Anthropic’s docs call server-side compaction (compact_20260112) the primary strategy for long conversations; the context-windows page lists it as beta for Claude 4.6 and later models. Client-side SDK compaction is deprecated in the TypeScript/Ruby SDKs and removed from the Python SDK. Anthropic’s Claude Code docs note project-root CLAUDE.md is re-read from disk after /compact, while instructions given only in chat can be lost.
Context editing
Selective clearing instead of summarising. Beta (header context-management-2025-06-27): clear_tool_uses_20250919 drops old tool results (configurable trigger, how many to keep, excluded tools), clear_thinking_20251015 manages thinking blocks. Clearing happens server-side; the client keeps the full history.
Memory outside the window
- Memory tool (
memory_20250818): Claude writes files to persistent storage before results are cleared and looks them up later. - Structured note-taking (Anthropic’s term): agents keep NOTES.md/to-do files.
- Progress files across windows: Anthropic’s long-running-agent harness uses an init script, a progress file, a JSON feature list with pass/fail status and git commits so a fresh window can recover state.
- Claude Code auto memory: Claude writes its own notes to
~/.claude/projects/<project>/memory/;MEMORY.md(first 200 lines or 25KB) loads every session, topic files load on demand. - Filesystem as context (Manus): treat files as unlimited memory, compress reversibly, keep failed attempts visible, and “recite” a todo list to keep goals in recent attention.
Practical notes
- Anthropic’s prompting guide says that when a window is cleared it can be better to start fresh and let the model rediscover state from files and git than to compact.
- Tell the model compaction exists so it does not stop early near the limit (docs provide sample wording).
- Clearing tool results or changing thinking settings invalidates the prompt cache for later content (prompt-caching).
- Vendor-neutral memory stores and RAG: see ai-memory.
Sources
- https://platform.claude.com/docs/en/build-with-claude/context-editing (2026-09-30)
- https://platform.claude.com/docs/en/build-with-claude/context-windows (2026-09-30)
- https://code.claude.com/docs/en/memory (2026-09-30)
- https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents (2025-11-26; accessed 2026-09-30)
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices (2026-09-30)
- https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus (2026-09-30)