AI safety and governance frameworks

Two layers: voluntary or certifiable frameworks that organisations adopt (NIST, ISO), and frontier labs’ own safety policies. Binding law is covered in ai-regulation-and-policy. Attack-side security is in prompt-injection-and-agent-security.

Organisational frameworks

  • NIST AI Risk Management Framework 1.0, released 2023-01-26: four functions Govern, Map, Measure, Manage. Generative AI Profile NIST-AI-600-1 released 2024-07-26. NIST says AI RMF 1.0 “is being revised as part of the White House AI Action Plan”, and on 2026-04-07 published a concept note for a critical-infrastructure profile.
  • ISO/IEC 42001: “Information technology - Artificial intelligence - Management system”, developed by ISO/IEC JTC 1/SC 42 and published in December 2023 (secondary sources; the ISO catalogue page returned HTTP 403). Per secondary sources it is an AI management-system standard that organisations can be audited and certified against; the clause structure, combination with ISO/IEC 27001 and certification timing were not verified against ISO. Any 2026 amendment is unverified.

Frontier lab frameworks

  • Anthropic Responsible Scaling Policy: current version 3.4, effective 2026-07-08 (per Anthropic’s page, last updated 2026-08-14, published together with an August 2026 Risk Report): revised automated-R&D threshold, risk reports shared with at least 200 employees, coverage-date flexibility, redaction indications, external review of every part of the unredacted report. Company note: anthropic.
  • Google DeepMind Frontier Safety Framework: version 3.1 updated 2026-04-17 (the page says originally published 2025-09-22): adds a harmful-manipulation Critical Capability Level, updated misalignment protocols, and “Tracked Capability Levels” for earlier warning; safety-case reviews before external launches and coverage of large internal deployments. Company note: deepmind-research-lab.
  • OpenAI Preparedness Framework: version 2 published 2025-04-15 (secondary sources; openai.com returned HTTP 403): tracks Biological and Chemical, Cybersecurity and AI Self-Improvement capabilities at High and Critical thresholds, with research categories such as long-range autonomy and sandbagging (all per secondary sources). On 2026-05-28 OpenAI published a Frontier Governance Framework that maps its practices to California’s SB 53 and the EU GPAI Code of Practice, with Tier 2 and Tier 3 aligned to the Preparedness Framework’s High and Critical thresholds (press and tracker sources, not read on openai.com; unverified). Company note: openai.
  • Others: a July 2026 tracker (Vorp Labs, secondary; figures unverified) counts twelve developers with published frameworks and lists Seoul-summit signatories (including Mistral AI, Zhipu AI and MiniMax) that had not published one by 2026-07-04; it also reports (unverified) xAI dropping its SB 53 compliance statement in a 2026-06-30 rewrite.

Law now pulls these frameworks in

  • California SB 53 (Transparency in Frontier AI Act), effective 2026-01-01: large frontier developers (thresholds of USD 500M revenue and 10^26 FLOPs per law-firm alerts and Wikipedia; statute text not re-read here) must publish, follow and update annually a written frontier AI framework; civil penalties up to USD 1M per violation.
  • New York RAISE Act, signed 2025-12-19, effective 2027-01-01.
  • EU AI Act GPAI obligations since 2025-08-02 (Code of Practice route); see ai-regulation-and-policy.
  • Evidence base: the International AI Safety Report 2026 (2026-02-03, chair Yoshua Bengio) and Stanford’s 2026 AI Index (documented AI incidents 362 in 2025, up from 233) describe the risk picture these frameworks respond to; see ethics-in-ai.

How teams use them

In the author’s suggestion: map product risks to NIST functions, run evals and red teaming (llm-as-judge-and-evals, agent-evals-and-observability), log agent actions, and track lab policy changes when choosing models. General background: ai-research-labs.

Open items

  • OpenAI Preparedness Framework v2 and the 2026-05-28 Frontier Governance Framework: dates and content from secondary sources; openai.com blocked (403).
  • ISO/IEC 42001 edition status and amendments: ISO page blocked (403).
  • NIST AI RMF 1.0 revision: NIST says it is being revised under the White House AI Action Plan; no revised text found as of 2026-10-05.
  • Current Google DeepMind FSF and Anthropic RSP versions were fetched on 2026-10-02 and a 2026-07 tracker agrees on RSP v3.4 (2026-07-08); not re-fetched on 2026-10-05.

Sources