Event Summary
In May 2025, Anthropic activated ASL-3 (AI Safety Level 3) safeguards for its most capable models, marking the first time an AI lab implemented a tiered safety framework directly tied to capability thresholds. The Responsible Scaling Policy (RSP) v3.0 defined capability thresholds that trigger increasingly stringent safety measures — from monitoring to deployment restrictions. ASL-3's activation set a precedent for how AI labs could operationalize safety research into enforceable internal governance, rather than relying on voluntary commitments.
Context & Narrative
Anthropic had pioneered the concept of Responsible Scaling Policy in 2023, proposing that as AI models became more capable, labs should progressively tighten safety measures. ASL-2 covered current models with standard monitoring. ASL-3 was designed for models with capabilities that could pose significant risks — including autonomous AI research, bioweapon development assistance, or large-scale persuasion. The May 2025 activation meant that Anthropic assessed its latest models as reaching the ASL-3 capability threshold, triggering enhanced safeguards: asynchronous monitoring classifiers for analyzing model outputs for threats, post-hoc jailbreak detection with rapid response procedures, and deployment restrictions for the most sensitive use cases. The framework was notable for its specificity: it defined concrete capability thresholds (e.g., performance on specific biology benchmarks, autonomous coding capabilities) rather than vague safety commitments. It also included a commitment to not deploy models beyond ASL-3 in uncontrolled settings without further safety evidence. While critics argued that the thresholds were self-defined and self-enforced, ASL-3's activation was a milestone in moving AI safety from abstract principles to operational engineering. It influenced discussions at California's SB 53 hearings and was cited as a model for how AI companies could implement the EU AI Act's requirements for 'systemic risk' models.
Key Findings
-
Fact Grade A
Anthropic activated ASL-3 safeguards in May 2025, implementing the first operational AI safety framework tied to concrete capability thresholds.
Sources [1]
Impact Assessment
-
Paradigm Shift +2 · Long-term
First operational AI safety framework tied to concrete capability thresholds. Created a precedent that safety measures should scale with model capabilities. Influenced regulatory frameworks including SB 53 and EU AI Act implementation.
Affected Groups: AI safety researchers, AI labs, policymakers, regulators
-
Risk Creation -1 · Medium-term
Critics argued thresholds were self-defined and self-enforced, lacking independent oversight. Raised questions about whether AI labs could credibly self-regulate capability thresholds.
Affected Groups: public, ethicists, regulators
Consensus & Sources
-
1
We activated ASL-3 safeguards for relevant models in May 2025 and have been working to improve them ever since.Reference Evidence Citation logged Live source