Microsoft launches Project Perception with AI agents that outperform Mythos at half the cost

Reviewed byNidhi Govil

9 Sources

Share

Microsoft unveiled Project Perception, an agentic cybersecurity system coordinating red, blue, and green AI agents for autonomous defense. Its MAI-Cyber-1-Flash model scored 96% on the CyberGym benchmark—12 points above Anthropic's Mythos—while delivering 50% cost savings. The system enters public preview August 3, integrated directly into Microsoft Defender.

Microsoft Cybersecurity Gets Agentic AI Upgrade with Project Perception

Microsoft announced Project Perception on Monday at a San Francisco event, marking a significant shift in how enterprises defend against AI-powered cyber threats

1

. The agentic cybersecurity system coordinates three classes of AI agents in a continuous loop: red team agents that simulate potential attacks, blue team agents that investigate and prioritize risks, and green team agents that implement fixes across the environment

5

. This represents a different level of coordination than previous Microsoft security tools, addressing the entire security lifecycle from identifying attack paths to implementing fixes that harden environments

3

.

Source: GeekWire

Source: GeekWire

The system enters public preview on August 3, built directly into Microsoft Defender, with a gradual rollout to all Microsoft Security products planned

2

. Hayete Gallot, Microsoft's vice president for security, described Project Perception as enabling enterprise defenders to "defend against AI with AI at the scale and speed that the attackers have"

1

.

MAI-Cyber-1-Flash Delivers Superior Performance on CyberGym Benchmark

Microsoft also launched its first cybersecurity-specialized model, MAI-Cyber-1-Flash, built specifically to find challenging vulnerabilities in complex codebases

1

. When combined with MDASH—Microsoft's harness dedicated to software vulnerability identification and remediation—and OpenAI GPT-5.4, the configuration achieved a 96% success rate on CyberGym, the industry's primary benchmark for vulnerability assessment

2

.

This performance significantly outpaces competitors: Anthropic Mythos 5 scored 84%, OpenAI GPT-5.5 Cyber achieved 85.6%, GPT-5.6 Sol reached 83.6%, and Google's Gemini 3.5 Flash Cyber in CodeMender scored 83.2%

4

. Mustafa Suleyman, CEO of Microsoft AI and co-founder of DeepMind, called the result "quite remarkable" during Monday's announcement

4

.

Source: TechCrunch

Source: TechCrunch

Multi-Model Architecture Delivers Cost Efficiency and Performance

The system's efficiency comes from intelligent model orchestration. MAI-Cyber-1-Flash handles approximately 90% of queries, detecting vulnerabilities, patching them, and validating fixes

2

. The remaining 10% of more complex tasks are handed off to GPT-5.4, which is approximately 10 times larger

2

. This multi-model approach delivers nearly 50% cost savings compared to the current MDASH configuration on the market, with consumption-based pricing measured by security compute units

2

.

Project Perception includes an orchestration layer that functions as a multiplexer for models, choosing the best option to manage quality, reliability, latency, and cost for each specific task

3

. This represents a best practice that Forrester recommends for organizations building AI security capabilities.

Red, Blue, and Green AI Agents Coordinate Autonomous Defense

The red team agents provide detailed attack simulations, offering context about potential threat actors and likely vulnerabilities they might exploit

1

. Blue team agents pull in and evaluate threat intelligence information, investigate potential attacks, prioritize them, and build new detections to identify attacker behavior in the future

3

. Green team agents take corrective actions against bugs, building potential fixes for vulnerabilities and even connecting to GitHub to propose fixes and open pull requests

3

.

Source: The Register

Source: The Register

Dave Weston, the lead engineer for Project Perception, described the platform as a massive efficiency upgrade: "We've gone from this taking hours and hours of manual work from multiple specialized folks across the security organization—appsec hunters, remediation engineers, you name it—and in minutes, we have a fix for all of this. Not only do we discover the issues and prioritize them, but we have detection, posture fixing, and even a code fix"

1

.

Microsoft Expands AI Security Research with New Initiatives

Alongside Project Perception, Microsoft introduced Microsoft Security FORGE Labs—a new AI security research arm led by VP of Security Research Taesoo Kim—and the External Red Team Alliance (EXTRA)

4

. EXTRA includes unrestricted funding to 18 university labs across six continents to support AI safety research, and will build a distributed network of specialists for red teaming across specific domains

4

.

Suleyman emphasized Microsoft's competitive advantage: "Microsoft has long been the trusted steward of some of the most valuable, important government and enterprise data in the world over many decades. We've accrued a phenomenal amount of data in that time. It's that data combined with the expertise that we have from world-class cybersecurity experts in the company that we've been able to really drive this combined model"

2

.

What This Means for Enterprise Security

Project Perception enters an increasingly crowded field of AI security solutions. Anthropic launched Mythos earlier this year through its Glasswing program, while OpenAI released its own security solution in May through Day Break

1

. Microsoft's announcement comes as the White House launched the Gold Eagle initiative this month to coordinate AI-powered cyber defense, positioning Project Perception as Microsoft's bid to become the platform that defense runs on

5

.

For organizations considering autonomous cybersecurity, Forrester recommends requiring observability by default to monitor and detect failures, and ensuring least privilege access to prevent incidents similar to recent OpenAI security events

3

. The public preview beginning August 3 will test whether the 96% CyberGym benchmark performance holds against real-world attacks rather than synthetic evaluations

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved