4 Sources
[1]
AI failed to properly patch software flaws 74% of the time, 1Password's study warns
Follow ZDNET: Add us as a preferred source on Google. ZDNET's key takeaways * A new study shows the effectiveness of AI-generated patches. * Only 26% of the patches generated were actually usable. * AI isn't ready to patch the planet, but it can be used in cyberdefense. New research has
[2]
AI struggles to patch vulns without adult supervision
AI models may not be that good at fixing security flaws. Researchers at 1Password's Off-by-1 Labs analyzed security patches generated by two frontier models - ChatGPT 5.5 at "medium" effort and Claude Opus 4.8 at "high" effort - and found that autonomous patches cleanly fixed vulnerabilities only
[3]
Shock horror -- AI-generated security patches fall short of actually solving all the problems they were meant to fix
AI without oversight creates patches that rarely fix the issue entirely * Researchers tested AI-generated patches on six CVEs with poor success rates * Many fixes failed, altered behavior, or introduced new vulnerabilities * Guidance improved outcomes, leading to FLAWED evaluation harness
[4]
1Password study: only 26% of AI security patches fully fix flaws
AI1Password study: only 26% of AI security patches fully fix flaws Across 6,080 patches, most left attack paths or regressions 1Password's Off-by-1 Labs has published a new study on AI-generated security patches, examining 6,080 fixes for six recently disclosed vulnerabilities. By the lab's
Share
Copy Link
1Password's Off-by-1 Labs tested 6,080 AI-generated patches across six CVEs using ChatGPT and Claude Opus models. Only 26% fully fixed software vulnerabilities without issues. Nearly half failed to close exploit paths, while others introduced new bugs or altered application behavior, highlighting critical limitations of AI in security remediation.
1Password's newly formed security research team, Off-by-1 Labs, has released findings that challenge the readiness of AI-generated patches for software vulnerabilities
1
. The 1Password study examined 6,080 AI-generated patches across six recently disclosed CVEs, revealing that only 26% of patches fully resolved vulnerabilities without changing application behavior2
3
. The research tested two frontier models—ChatGPT 5.5 at medium effort and Claude Opus 4.8 at high effort—against vulnerabilities unlikely to exist in their training data. These included CVE-2026-31431 (Linux privilege escalation), CVE-2026-34197 (ActiveMQ Remote Code Execution), CVE-2026-8512 (Chrome use-after-free), CVE-2026-45185 (EXIM Remote Code Execution), CVE-2026-22738 (SpringAI Remote Code Execution), and GHSA-wpqr-6v78-jr5g (Gemini CLI Remote Code Execution)1
.The Off-by-1 Labs research documented troubling failure patterns in AI-generated patches for software vulnerabilities. A substantial 49.3% of patches failed to close at least one existing exploit path, leaving systems vulnerable to attack
2
4
. Another 20.1% fixed the original vulnerability but altered application behavior in ways that could disrupt normal operations3
. More concerning, 2.3% of patches introduced entirely new security issues while attempting to resolve the original flaw1
. The research team found that 2.2% of patches both failed to fix the vulnerability and simultaneously introduced new exploit paths2
. Keith Hoodlet, director of security research at 1Password, emphasized that these results demonstrate AI struggles to patch vulnerabilities without human supervision2
.
Source: ZDNet
Researchers coined the term FLAWED—Fix-Like Artifacts With Embedded Defects—to describe the superficial nature of many AI security patches
1
4
. These patches often appear functional on the surface but contain fundamental problems underneath. Among patches initially deemed successful, more than one-third were classified as fragile fixes—meaning they blocked specific proof-of-concept exploits without addressing root causes4
. The limitations of AI in security remediation become especially apparent when models receive incorrect guidance. With proper initial direction, AI models achieved a 65% success rate in vulnerability triage, but incorrect guidance dropped effectiveness to just 15.2%2
3
. Human developers typically catch misleading information during code review, while AI models lack this critical reasoning capability.
Source: The Register
Related Stories
The 1Password study concluded that autonomous patching presents net-negative outcomes for organizations in cybersecurity
2
. Authors Axel Mierczuk, Spencer Michaels, and Keith Hoodlet warn that cognitive surrender to processes with only 26% success rates poses significant long-term risks2
. While the average successful, clean patch cost just $6.74 including failed attempts, the cognitive load of reviewing mountains of mostly-incorrect, similar-yet-subtly-different patches likely exceeds the effort required for human engineers to patch vulnerabilities themselves2
. Hoodlet told ZDNET that human supervision remains paramount, with AI better suited for vulnerability discovery and triage rather than autonomous remediation1
. The research team released a patch evaluation harness, also named FLAWED, on GitHub to help organizations assess where AI-generated patches might produce better or worse outcomes1
3
. This tool enables companies to determine where human experts remain most needed in the patching process, ensuring informed decisions about business risks associated with each security choice.Summarized by
Navi
[2]
[3]
10 Jun 2026•Technology

05 Sept 2025•Technology

29 Jul 2025•Technology

1
Science and Research

2
Technology

3
Policy and Regulation
