Context compaction and memory

Long-running agents outlive one context window. Three complementary mechanisms keep them going; see context-engineering for the wider picture.

Compaction

Summarise the conversation so far and continue from the summary. Anthropic’s docs call server-side compaction (compact_20260112) the primary strategy for long conversations; the context-windows page lists it as beta for Claude 4.6 and later models. Client-side SDK compaction is deprecated in the TypeScript/Ruby SDKs and removed from the Python SDK. Anthropic’s Claude Code docs note project-root CLAUDE.md is re-read from disk after /compact, while instructions given only in chat can be lost.

Context editing

Selective clearing instead of summarising. Beta (header context-management-2025-06-27): clear_tool_uses_20250919 drops old tool results (configurable trigger, how many to keep, excluded tools), clear_thinking_20251015 manages thinking blocks. Clearing happens server-side; the client keeps the full history.

Memory outside the window

  • Memory tool (memory_20250818): Claude writes files to persistent storage before results are cleared and looks them up later.
  • Structured note-taking (Anthropic’s term): agents keep NOTES.md/to-do files.
  • Progress files across windows: Anthropic’s long-running-agent harness uses an init script, a progress file, a JSON feature list with pass/fail status and git commits so a fresh window can recover state.
  • Claude Code auto memory: Claude writes its own notes to ~/.claude/projects/<project>/memory/; MEMORY.md (first 200 lines or 25KB) loads every session, topic files load on demand.
  • Filesystem as context (Manus): treat files as unlimited memory, compress reversibly, keep failed attempts visible, and “recite” a todo list to keep goals in recent attention.

Practical notes

  • Anthropic’s prompting guide says that when a window is cleared it can be better to start fresh and let the model rediscover state from files and git than to compact.
  • Tell the model compaction exists so it does not stop early near the limit (docs provide sample wording).
  • Clearing tool results or changing thinking settings invalidates the prompt cache for later content (prompt-caching).
  • Vendor-neutral memory stores and RAG: see ai-memory.

Sources