AI safety and governance frameworks
Two layers: voluntary or certifiable frameworks that organisations adopt (NIST, ISO), and frontier labs’ own safety policies. Binding law is covered in ai-regulation-and-policy. Attack-side security is in prompt-injection-and-agent-security.
Organisational frameworks
- NIST AI Risk Management Framework 1.0, released 2023-01-26: four functions Govern, Map, Measure, Manage. Generative AI Profile NIST-AI-600-1 released 2024-07-26. NIST says AI RMF 1.0 “is being revised as part of the White House AI Action Plan”, and on 2026-04-07 published a concept note for a critical-infrastructure profile.
- ISO/IEC 42001: “Information technology - Artificial intelligence - Management system”, developed by ISO/IEC JTC 1/SC 42 and published in December 2023 (secondary sources; the ISO catalogue page returned HTTP 403). Per secondary sources it is an AI management-system standard that organisations can be audited and certified against; the clause structure, combination with ISO/IEC 27001 and certification timing were not verified against ISO. Any 2026 amendment is unverified.
Frontier lab frameworks
- Anthropic Responsible Scaling Policy: current version 3.4, effective 2026-07-08 (per Anthropic’s page, last updated 2026-08-14, published together with an August 2026 Risk Report): revised automated-R&D threshold, risk reports shared with at least 200 employees, coverage-date flexibility, redaction indications, external review of every part of the unredacted report. Company note: anthropic.
- Google DeepMind Frontier Safety Framework: version 3.1 updated 2026-04-17 (the page says originally published 2025-09-22): adds a harmful-manipulation Critical Capability Level, updated misalignment protocols, and “Tracked Capability Levels” for earlier warning; safety-case reviews before external launches and coverage of large internal deployments. Company note: deepmind-research-lab.
- OpenAI Preparedness Framework: version 2 published 2025-04-15 (secondary sources; openai.com returned HTTP 403): tracks Biological and Chemical, Cybersecurity and AI Self-Improvement capabilities at High and Critical thresholds, with research categories such as long-range autonomy and sandbagging (all per secondary sources). On 2026-05-28 OpenAI published a Frontier Governance Framework that maps its practices to California’s SB 53 and the EU GPAI Code of Practice, with Tier 2 and Tier 3 aligned to the Preparedness Framework’s High and Critical thresholds (press and tracker sources, not read on openai.com; unverified). Company note: openai.
- Others: a July 2026 tracker (Vorp Labs, secondary; figures unverified) counts twelve developers with published frameworks and lists Seoul-summit signatories (including Mistral AI, Zhipu AI and MiniMax) that had not published one by 2026-07-04; it also reports (unverified) xAI dropping its SB 53 compliance statement in a 2026-06-30 rewrite.
Law now pulls these frameworks in
- California SB 53 (Transparency in Frontier AI Act), effective 2026-01-01: large frontier developers (thresholds of USD 500M revenue and 10^26 FLOPs per law-firm alerts and Wikipedia; statute text not re-read here) must publish, follow and update annually a written frontier AI framework; civil penalties up to USD 1M per violation.
- New York RAISE Act, signed 2025-12-19, effective 2027-01-01.
- EU AI Act GPAI obligations since 2025-08-02 (Code of Practice route); see ai-regulation-and-policy.
- Evidence base: the International AI Safety Report 2026 (2026-02-03, chair Yoshua Bengio) and Stanford’s 2026 AI Index (documented AI incidents 362 in 2025, up from 233) describe the risk picture these frameworks respond to; see ethics-in-ai.
How teams use them
In the author’s suggestion: map product risks to NIST functions, run evals and red teaming (llm-as-judge-and-evals, agent-evals-and-observability), log agent actions, and track lab policy changes when choosing models. General background: ai-research-labs.
Open items
- OpenAI Preparedness Framework v2 and the 2026-05-28 Frontier Governance Framework: dates and content from secondary sources; openai.com blocked (403).
- ISO/IEC 42001 edition status and amendments: ISO page blocked (403).
- NIST AI RMF 1.0 revision: NIST says it is being revised under the White House AI Action Plan; no revised text found as of 2026-10-05.
- Current Google DeepMind FSF and Anthropic RSP versions were fetched on 2026-10-02 and a 2026-07 tracker agrees on RSP v3.4 (2026-07-08); not re-fetched on 2026-10-05.
Sources
- https://www.nist.gov/itl/ai-risk-management-framework (2026-10-05)
- https://www.anthropic.com/responsible-scaling-policy (2026-10-02)
- https://deepmind.google/discover/blog/strengthening-our-frontier-safety-framework/ (2026-10-02)
- https://vorplabs.com/ai-regulatory-updates/frontier-ai-frameworks (secondary, 2026-10-05)
- Web search 2026-10-05: OpenAI Frontier Governance Framework coverage (startuphub.ai, enterprisedna.co), OpenAI Preparedness Framework v2 coverage, ISO/IEC 42001 coverage (KPMG, ANAB), SB 53 and RAISE Act law-firm alerts (MoFo, Wiley) and Wikipedia’s SB 53 page
- https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026 ; https://hai.stanford.edu/ai-index/2026-ai-index-report (2026-10-05)