8 Sources
[1]
Anthropic offers $20,000 to whoever can jailbreak its new AI safety system
The company has upped its reward for red-teaming Constitutional Classifiers. Here's how to try. Can you jailbreak Anthropic's latest AI safety measure? Researchers want you to try -- and are offering up to $20,000 if you succeed. On Monday, the company released a new paper outlining an AI safety
[2]
Jailbreak Anthropic's new AI safety system for a $15,000 reward
In testing, the technique helped Claude block 95% of jailbreak attempts. But the process still needs more 'real-world' red-teaming. Can you jailbreak Anthropic's latest AI safety measure? Researchers want you to try -- and are offering up to $15,000 if you succeed. On Monday, the company released
[3]
Anthropic claims new AI security method blocks 95% of jailbreaks, invites red teamers to try
Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Two years after ChatGPT hit the scene, there are numerous large language models (LLMs), and nearly all remain ripe for jailbreaks -- specific prompts and other workarounds
[4]
Anthropic has a new security system it says can stop almost all AI jailbreaks
Tests resulted in more than an 80% reduction in successful jailbreaks In a bid to tackle abusive natural language prompts in AI tools, OpenAI rival Anthropic has unveiled a new concept it calls "constitutional classifiers"; a means of instilling a set of human-like values (literally, a
[5]
Anthropic's New Technique Can Protect AI From Jailbreak Attempts
Constitutional Classifiers were tested on the Claude 3.5 Sonnet Anthropic announced the development of a new system on Monday that can protect artificial intelligence (AI) models from jailbreaking attempts. Dubbed Constitutional Classifiers, it is a safeguarding technique that can detect when a
[6]
Anthropic dares you to jailbreak its new AI model
Even the most permissive corporate AI models have sensitive topics that their creators would prefer they not discuss (e.g., weapons of mass destruction, illegal activities, or, uh, Chinese political history). Over the years, enterprising AI users have resorted to everything from weird text strings
[7]
Constitutional classifiers: New security system drastically reduces chatbot jailbreaks
A large team of computer engineers and security specialists at AI app maker Anthropic has developed a new security system aimed at preventing chatbot jailbreaks. Their paper is published on the arXiv preprint server. Ever since chatbots became available for public use, users have been finding ways
[8]
Anthropic makes 'jailbreak' advance to stop AI models producing harmful results
Artificial intelligence start-up Anthropic has demonstrated a new technique to prevent users from eliciting harmful content from its models, as leading tech groups including Microsoft and Meta race to find ways that protect against dangers posed by the cutting-edge technology. In a paper released
Share
Copy Link
Anthropic introduces a new AI safety system called Constitutional Classifiers, designed to prevent jailbreaking attempts. The company is offering up to $20,000 to anyone who can successfully bypass this security measure.

Anthropic, a leading AI company, has unveiled a novel approach to AI safety called Constitutional Classifiers. This system is designed to prevent "jailbreaking" attempts on large language models (LLMs) like their Claude AI
1
.The Constitutional Classifiers system is based on Anthropic's Constitutional AI approach, which aims to make AI models "harmless" by adhering to a set of principles or "constitution"
2
. Key features include:In initial testing, Anthropic reported significant success:
3
Anthropic is now inviting the public to test their system:
1
While the results are promising, Anthropic acknowledges some limitations:
4
Related Stories
This development is significant for several reasons:
Some critics argue that Anthropic is essentially crowdsourcing its security work without adequate compensation. Others worry about the potential dual-use nature of such research, as it could inadvertently provide insights for creating more sophisticated jailbreaking techniques
5
.As AI technology continues to advance, the development of robust safety measures like Constitutional Classifiers will likely play a crucial role in ensuring responsible AI deployment and mitigating potential risks associated with large language models.
Summarized by
Navi
[3]
1
Technology

2
Policy and Regulation

3
Technology
