Summary

Angie Jones argues that agent failures in production are rarely a model-capability problem — they are a system design problem. Her framing example: an agent updates the wrong customer account not because the model “hallucinated” but because “the model did exactly what the surrounding system allowed.” The fix isn’t a smarter model, it’s tighter surrounding infrastructure: policy checks, provenance tracking, scoped permissions, stopping mechanisms, recoverable state, audit trails, and portable controls.

Key ideas

1. Policy enforcement as foundation

Separate the model’s proposal of an action from the system’s authority to execute it. A policy service validates every proposed action against user-approved constraints before it runs — models suggest, policy layers decide.

2. Treat context as untrusted input

Agents pull instructions from prompts, stored memory, and downloaded skills — any of which can be contaminated (directly analogous to prompt injection). Stored context needs provenance: who authored it, when, and what risk tier it carries. Trust warnings alone only partially mitigate contamination — provenance needs to be technically verifiable, not just disclosed.

3. Narrow capability boundaries

Replace broad OAuth-style scopes with narrow, single-purpose capabilities — e.g. “update one shipping address on one specific account, expires after use.” This bounds the blast radius of prompt injection and makes audit trails legible.

4. Stopping mechanisms belong in the workflow, not the model

Restraint has to be engineered into normal operation — routing to a human when confidence drops below a threshold, and kill switches that can interrupt an agent mid-run when a human takes over. Don’t rely on the model to decide when to stop itself.

5. State preservation for recovery

Long-running agents need idempotency keys and confirmed checkpoints. Restarting a form-filling task from scratch risks duplicate submissions; blind resumption risks skipped fields. Sandbox hibernation must properly restore state on wake, not just resume blindly.

6. Audit trails and context graphs

Logs should connect received context → requested action → permitting policy, not just record raw conversation transcripts. A context graph (the article cites a Neo4j-based example) preserves these relationships over time so a later agent or human can query “why did this happen” without replaying the whole conversation. Exclude secrets from logs; apply retention policies.

7. Portable controls across models

Permissions baked into prompt text drift when you swap models or deployments. MCP gives a model-agnostic, shared way to describe and call tools, so the control layer survives a model swap.

Practical recommendations (as a checklist)

  • Policy validation layer separating proposal from execution
  • Context provenance: author, timestamp, risk classification on all stored/retrieved context
  • Scoped, time-limited, action-specific capabilities instead of broad grants
  • Confidence-threshold routing to humans + working kill switches
  • Idempotency keys on all long-running/resumable operations
  • Execution logs linking context → policy → action → outcome
  • MCP (or equivalent) for tool interfaces, not model-specific prompt hacks
  • Secrets excluded from audit trails; retention policy defined

Context

Published 2026-08-25 on the Agentic AI Foundation (AAIF, Linux Foundation) blog by Angie Jones (AAIF VP; creator of Test Automation University). AAIF governs MCP, Goose, AGENTS.md and A2A (MCP, agents-md, a2a-protocol). The post closes by promoting AGNTCon + MCPCon North America, 22-23 October 2026, San Jose McEnery Convention Center (workshops 21 October; confirmed by Linux Foundation event page and AAIF press release). The post itself frames five principles (policy enforcement, context provenance, narrow permissions, safe stopping, state preservation and recovery); the audit-trail/context-graph and portable-controls points above are developed as additions to those, with the context-graph example attributed to Stephen Chin of Neo4j. Related: agent-sandboxes, prompt injection.

Sources