2 Sources
[1]
Anthropic just made it harder for AI to go rogue with its updated safety policy
The policy, originally introduced in 2023, has evolved with new protocols to ensure that AI models, as they grow more powerful, are developed and deployed safely. This revised policy sets out specific Capability Thresholds -- benchmarks that indicate when an AI model's abilities have reached a
[2]
Anthropic updates policy to address AI risks
Anthropic has developed a framework for assessing different AI capabilities to be better able to respond to emerging risks. Anthropic, the AI safety and research start-up behind the chatbot Claude, has updated its scaling policy, which develops a more flexible approach to assessing and managing AI
Share
Copy Link
Anthropic has updated its Responsible Scaling Policy, introducing new protocols and governance measures to ensure the safe development and deployment of increasingly powerful AI models.

Anthropic, the AI safety and research company behind the chatbot Claude, has announced significant updates to its Responsible Scaling Policy. This policy, initially introduced in 2023, aims to address the growing risks associated with increasingly powerful AI systems
1
2
.The revised policy introduces several new elements:
Capability Thresholds: These are specific benchmarks that indicate when an AI model's abilities have reached a point where additional safeguards are necessary. For example, if a model can assist in creating chemical, biological, or nuclear weapons, it would trigger higher safety standards
1
2
.AI Safety Levels (ASLs): Inspired by U.S. government biosafety standards, these levels range from ASL-2 (current safety standards) to ASL-3 and above (stricter protections for riskier models)
1
.Required Safeguards: These are specific measures implemented when a capability threshold is reached, ensuring appropriate risk mitigation
2
.A key addition to Anthropic's safety framework is the creation of a Responsible Scaling Officer (RSO) role. Jared Kaplan, Anthropic's co-founder and chief science officer, will assume this position, overseeing compliance with the policy and having the authority to pause AI training or deployment if necessary
1
2
.The policy pays particular attention to areas with potential for significant harm:
These areas are subject to stringent monitoring and safeguards to prevent misuse or unintended consequences
1
.Anthropic's updated policy is designed to be "exportable," potentially serving as a blueprint for the broader AI industry. By introducing a structured approach to scaling AI development, Anthropic aims to create a "race to the top" for AI safety
1
.Related Stories
The policy update comes at a time of increasing regulatory scrutiny in the AI industry. Anthropic's framework could serve as a prototype for future government regulations, offering a clear structure for when AI models should be subject to stricter controls
1
.Anthropic states that all its current models meet the ASL-2 standard. The company commits to conducting routine evaluations of its AI models to ensure appropriate safeguards are in place
2
.Anthropic's updated Responsible Scaling Policy represents a significant step in AI governance and risk management. By proactively addressing potential risks and setting industry standards, Anthropic is positioning itself as a leader in responsible AI development, potentially influencing the future direction of AI safety practices across the industry
1
2
.Summarized by
Navi
[2]
26 Feb 2026•Policy and Regulation

15 Aug 2026•Technology

05 Jun 2026•Policy and Regulation

1
Technology

2
Science and Research

3
Technology
