Guardrails, Jailbreaks and Red Teaming
Guardrails are checks placed around a model or agent: input and output classifiers, policy filters, allow-lists, schema validation, human approval for risky actions. Llama Guard (2023) is an example of an LLM-based input/output safeguard model.
Jailbreaks are prompts that make an aligned model ignore its safety training. “Jailbroken” (2023) analysed why safety training fails (competing objectives, mismatched generalisation) and automated adversarial suffixes (Zou et al., 2023) transferred between models, so no prompt-level defence is absolute. Prompt injection is the related attack where untrusted content (web page, email, tool output) carries instructions an agent then follows: see prompt-injection-and-agent-security.
Red teaming is deliberately attacking your own system, manually or with automated adversaries, before release and continuously after.
Practice for agents
Layer defences: least-privilege tools, agent-sandboxes, approval gates in the agent-harness, gateways (agentgateway, arcade-ai-mcp-gateway), monitoring (agent-evals-and-observability). Treat model-based filters as one layer, not the boundary.
Related: ai-safety-and-governance, alignment-and-sycophancy, data-privacy-in-ai, use-of-ai-in-hacking, ai-agents.
Sources
- https://arxiv.org/abs/2312.06674 (Llama Guard, 2023)
- https://arxiv.org/abs/2307.02483 (Jailbroken, 2023)
- https://arxiv.org/abs/2307.15043 (Universal and Transferable Adversarial Attacks, 2023)
- arXiv pages re-opened 2026-10-07: Llama Guard (Inan et al., 2023-12-07, built on Llama2-7b), Jailbroken (Wei, Haghtalab, Steinhardt, 2023-07-05; abstract names competing objectives and mismatched generalization), Zou et al. (2023-07-27; suffixes transfer to black-box models incl. ChatGPT, Bard, Claude).