TrendAI Achieves 97% Success Rate on CyberGym AI Security Benchmark, Outpacing All Competitors

2 Sources

Share

TrendAI's agentic exploit-remediation engine AESIR has secured first place on the CyberGym Agentic AI Security Benchmark with a 97% success rate, surpassing all competitors by more than 12 percentage points. The system leverages seven AI models and over 20 years of vulnerability intelligence to compress exposure windows.

TrendAI Claims Top Position with Record-Breaking Performance

TrendAI has secured first place on the CyberGym Agentic AI Security Benchmark, achieving a 97% success rate that places it ahead of every other entrant currently on the leaderboard

1

2

. The score represents nearly a 4 percentage point improvement over the previous leader's 93.2% result posted on August 8, 2026, and establishes a commanding lead of more than 12 points ahead of both GPT-5.6 Sol and Claude Mythos 5

1

. CyberGym, developed by the University of California, Berkeley, evaluates AI security tooling against 1,507 confirmed vulnerabilities drawn from 188 large open-source software projects that run in enterprise environments

2

.

AESIR Engine Powers Multi-Model Architecture

The company's agentic exploit-remediation engine, codenamed AESIR, draws on seven AI models across four providers—Anthropic, DeepSeek, Google, and OpenAI—for AI-driven vulnerability detection and remediation

2

. Claude Opus 4.6 serves as the primary engine within this multi-model framework

1

. TrendAI's team attributes the benchmark lead to system architecture and engineering decisions rather than raw model access alone, noting that top CyberGym competitors typically deploy three to seven models in distinct roles since no single model excels across all vulnerability types

1

. The system operates on a persistent vulnerability ontology holding over 12,500 episodic memories and over 15,500 exploit seeds across more than 180 projects, built from TrendAI Research and ZDI's 20-plus years of vulnerability data

2

.

Graduated Pipeline Approach Balances AI and Classical Methods

TrendAI employs a graduated pipeline approach that routes tasks strategically between AI reasoning and deterministic methods based on complexity. Nearly a third of exploits were solved without AI involvement, and classical fuzzing outperformed AI reasoning on roughly a quarter of tasks

1

. This engineering choice reflects a pragmatic recognition that simpler vulnerability detection and remediation tasks benefit from deterministic methods first, reserving AI capabilities for more complex scenarios where pattern recognition and reasoning provide clear advantages.

Closing the Exposure Window Through Virtual Patches

Sharda Tickoo, Country Manager for India and SAARC at TrendAI, emphasized the strategic shift from discovery speed to exposure compression: "CyberGym's validation changes what's possible for enterprises adopting AI. The race isn't to find first. It's to close the exposure window. Cybersecurity has been defined by the scramble from discovery to exploitation; our job is to make that gap meaningless by compressing exposure time towards zero"

1

2

. The system translates validated vulnerability intelligence into virtual patches without waiting for normal patching cycles, shrinking exposure windows from weeks to near-zero timeframes

2

. This capability transforms vulnerability management from reactive patching into proactive defense in cybersecurity, addressing the accelerating timeline between discovery and exploitation that AI has introduced across the threat landscape

2

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved