Agent Harness
An agent harness is everything around the model that turns a bare LLM into a working agent: the loop that calls the model, the tool definitions and their execution, permission and approval rules, sandboxing, context management (compaction, memory, caching), instruction files, hooks and logging. The model supplies judgement; the harness supplies hands, guardrails and memory. Comparisons of coding agents often treat “model plus harness” as the unit; in the author’s view the same model can behave differently in different harnesses (not benchmarked here).
Typical components
- Loop and tool layer - function calling, shell/file/web tools, MCP servers.
- Context management - system prompt, instruction files such as AGENTS.md, compaction (context-compaction-and-memory), prompt-caching.
- Permissions and safety - approval prompts, allow-lists, agent-sandboxes, guardrails.
- Extension points - skills, hooks, sub-agents (claude-code-extensibility-guide).
- Observability - traces and evals (agent-evals-and-observability).
In the vault
Harness examples: claude-code, copilot-cli, open-interpreter, hermes-agent, openclaw, deepseek-harness; SDKs for building your own: claude-agent-sdk, openai-agents-sdk, google-adk. Comparison: claude-code-codex-goose-hermes-pi-comparison. Concept context: ai-agents, context-engineering.
Sources
- https://www.anthropic.com/engineering/building-effective-agents (opened 2026-10-07; published 2024-12-19; covers the augmented LLM, tools and the agent-computer interface, but does not use the word ‘harness’)
- Vault notes linked above (this note is a synthesis of them).
Open items
- The cited Anthropic article supports the components (tools, loop, interface design), not the term ‘harness’ itself; no primary source for the term was found.
- No single standard definition of ‘harness’ exists; the component list is a synthesis, not a cited spec.