2 Sources
[1]
TrendAI Ranks No. 1 on CyberGym Agentic AI Security Benchmark
Sharda Tickoo, Country Manager for India and SAARC at TrendAI: "CyberGym's validation changes what's possible for enterprises adopting AI. The race isn't to find first. It's to close the exposure window. Cybersecurity has been defined by the scramble from discovery to exploitation; our job is to make that gap meaningless by compressing exposure time towards zero. Our advantage is turning validated vulnerability intelligence into effective risk reduction faster." TrendAI's 97% CyberGym success rate ranks first among all current entrants on the benchmark, more than 12 points ahead of both GPT-5.6 Sol and Claude Mythos 5 The score is nearly 4 percentage points higher than the previous leader's 93.2% CyberGym tests against 1,507 confirmed vulnerabilities from 188 large open-source projects AESIR draws on seven AI models across four providers for detection and remediation, using Claude Opus 4.6 as its primary engine TrendAI's team attributes the lead to system architecture and engineering rather than raw model access, noting that top CyberGym competitors typically use three to seven models in distinct roles since no single model excels across all vulnerability types The system runs on a persistent vulnerability ontology holding over 12,500 episodic memories and over 15,500 exploit seeds across more than 180 projects, built from TrendAI Research and ZDI's 20-plus years of vulnerability data Nearly a third of exploits were solved without AI involvement, and classical fuzzing outperformed AI reasoning on roughly a quarter of tasks, informing graduated pipelines that route simpler tasks to deterministic methods first
[2]
TrendAI™ Ranks First on CyberGym Agentic AI Security Benchmark
The 97% success rate places TrendAI™ ahead of all other entrants on the AI security benchmark The TrendAI agentic exploit-remediation engine, codename AESIR, achieved the top position on CyberGym, an independent evaluation benchmark for AI-driven security tooling. The 97% success rate from TrendAI™ surpassed all other entrants currently on the leaderboard and is nearly 4 percentage points higher than the former leader, who posted a 93.2% score on Aug. 8, 2026. CyberGym, a University of California, Berkeley benchmark, evaluates AI security tooling against 1,507 confirmed vulnerabilities drawn from 188 large open-source software projects that run in every enterprise. TrendAI™ leverages seven AI models across four providers (Anthropic, DeepSeek, Google, and OpenAI) to maximize detection and remediation accuracy, and it powers a system built around one objective: close the exposure window faster than attackers can exploit it. It begins with discovery, powered by AI-scale identification of weak spots that static approaches miss. Each finding is then validated and turned into a working proof of concept, so teams understand exploitability and true business risk before the noise drowns out the signal. Only then does TrendAI rapidly reduce the risk, translating validated vulnerability intelligence into detection and virtual patches without waiting for the normal patching cycle. That last step is what makes our approach and usage of agentic security different. This capability is built on more than 20 years of vulnerability research and exploit intelligence from TrendAI™ ZDI, giving the system a deep foundation of real-world knowledge about how vulnerabilities are discovered, validated and exploited. AI is compressing the time from discovery to exploitation, so risk reduction must move faster than the patch cycle, not wait on it. By converting validated intelligence into virtual patching, TrendAI™ shrinks the exposure window from weeks to as close to zero as possible, turning vulnerability management from reactive patching into proactive defense. The TrendAI™ agentic exploit-remediation engine also powers dynamic discovery of threat actors actively exploiting vulnerabilities in the wild, giving security teams visibility into live exploitation in real time. Our real-world threat-intelligence layer supports the same goal: reduce exposure before attackers can exploit it. Sharda Tickoo, Country Manager for India and SAARC at TrendAI: "CyberGym's validation changes what's possible for enterprises adopting AI. The race isn't to find first. It's to close the exposure window. Cybersecurity has been defined by the scramble from discovery to exploitation; our job is to make that gap meaningless by compressing exposure time towards zero. Our advantage is turning validated vulnerability intelligence into effective risk reduction faster." Key findings include: * TrendAI's 97% CyberGym success rate ranks first among all current entrants on the benchmark, more than 12 points ahead of both GPT-5.6 Sol and Claude Mythos 5 * The score is nearly 4 percentage points higher than the previous leader's 93.2% * CyberGym tests against 1,507 confirmed vulnerabilities from 188 large open-source projects * AESIR draws on seven AI models across four providers for detection and remediation, using Claude Opus 4.6 as its primary engine * TrendAI's team attributes the lead to system architecture and engineering rather than raw model access, noting that top CyberGym competitors typically use three to seven models in distinct roles since no single model excels across all vulnerability types * The system runs on a persistent vulnerability ontology holding over 12,500 episodic memories and over 15,500 exploit seeds across more than 180 projects, built from TrendAI Research and ZDI's 20-plus years of vulnerability data
Share
Copy Link
TrendAI's agentic exploit-remediation engine AESIR has secured first place on the CyberGym Agentic AI Security Benchmark with a 97% success rate, surpassing all competitors by more than 12 percentage points. The system leverages seven AI models and over 20 years of vulnerability intelligence to compress exposure windows.
TrendAI has secured first place on the CyberGym Agentic AI Security Benchmark, achieving a 97% success rate that places it ahead of every other entrant currently on the leaderboard
1
2
. The score represents nearly a 4 percentage point improvement over the previous leader's 93.2% result posted on August 8, 2026, and establishes a commanding lead of more than 12 points ahead of both GPT-5.6 Sol and Claude Mythos 51
. CyberGym, developed by the University of California, Berkeley, evaluates AI security tooling against 1,507 confirmed vulnerabilities drawn from 188 large open-source software projects that run in enterprise environments2
.The company's agentic exploit-remediation engine, codenamed AESIR, draws on seven AI models across four providers—Anthropic, DeepSeek, Google, and OpenAI—for AI-driven vulnerability detection and remediation
2
. Claude Opus 4.6 serves as the primary engine within this multi-model framework1
. TrendAI's team attributes the benchmark lead to system architecture and engineering decisions rather than raw model access alone, noting that top CyberGym competitors typically deploy three to seven models in distinct roles since no single model excels across all vulnerability types1
. The system operates on a persistent vulnerability ontology holding over 12,500 episodic memories and over 15,500 exploit seeds across more than 180 projects, built from TrendAI Research and ZDI's 20-plus years of vulnerability data2
.TrendAI employs a graduated pipeline approach that routes tasks strategically between AI reasoning and deterministic methods based on complexity. Nearly a third of exploits were solved without AI involvement, and classical fuzzing outperformed AI reasoning on roughly a quarter of tasks
1
. This engineering choice reflects a pragmatic recognition that simpler vulnerability detection and remediation tasks benefit from deterministic methods first, reserving AI capabilities for more complex scenarios where pattern recognition and reasoning provide clear advantages.Related Stories
Sharda Tickoo, Country Manager for India and SAARC at TrendAI, emphasized the strategic shift from discovery speed to exposure compression: "CyberGym's validation changes what's possible for enterprises adopting AI. The race isn't to find first. It's to close the exposure window. Cybersecurity has been defined by the scramble from discovery to exploitation; our job is to make that gap meaningless by compressing exposure time towards zero"
1
2
. The system translates validated vulnerability intelligence into virtual patches without waiting for normal patching cycles, shrinking exposure windows from weeks to near-zero timeframes2
. This capability transforms vulnerability management from reactive patching into proactive defense in cybersecurity, addressing the accelerating timeline between discovery and exploitation that AI has introduced across the threat landscape2
.Summarized by
Navi
02 May 2025•Technology

21 Mar 2025•Technology

13 May 2026•Technology

1
Technology

2
Policy and Regulation

3
Health