2 Sources
[1]
New AI Jailbreak Method 'Bad Likert Judge' Boosts Attack Success Rates by Over 60%
Cybersecurity researchers have shed light on a new jailbreak technique that could be used to get past a large language model's (LLM) safety guardrails and produce potentially harmful or malicious responses. The multi-turn (aka many-shot) attack strategy has been codenamed Bad Likert Judge by Palo
[2]
Unit 42 Warns Developers of Technique That Bypasses LLM Guardrails | PYMNTS.com
Unit 42, a cybersecurity-focused unit of Palo Alto Networks, has warned developers of text-generation large language models (LLMs) of a potential threat that could bypass guardrails designed to prevent LLMs from delivering harmful and malicious requests. Dubbed "Bad Likert Judge," this technique
Share
Copy Link
Cybersecurity researchers unveil a new AI jailbreak method called 'Bad Likert Judge' that significantly increases the success rate of bypassing large language model safety measures, raising concerns about potential misuse of AI systems.

Cybersecurity researchers from Palo Alto Networks' Unit 42 have discovered a new jailbreak technique called 'Bad Likert Judge' that could potentially bypass the safety guardrails of large language models (LLMs). This multi-turn attack strategy significantly increases the likelihood of generating harmful or malicious responses from AI systems
1
.The technique exploits the LLM's own understanding of harmful content by:
This method leverages the model's long context window and attention mechanisms, gradually nudging it towards producing malicious responses without triggering internal protections
1
.Unit 42's research revealed alarming results:
2
.The discovery of 'Bad Likert Judge' highlights the ongoing challenges in AI security:
2
.Related Stories
This new jailbreak method adds to the growing list of AI security concerns:
1
.2
.As AI systems become more prevalent, the need for robust security measures intensifies:
2
.Summarized by
Navi
1
Technology

2
Policy and Regulation

3
Technology
