AI Agents from Microsoft and Wiz Achieve Over 90% Bug Detection Rate
Microsoft and Wiz's respective AI security agents, MDASH and Project Atlas, have demonstrated remarkable success in identifying software vulnerabilities, both exceeding 90% accuracy on the CyberGym benchmark.

In a significant leap forward for AI-driven cybersecurity, both Microsoft and Wiz have reported impressive results from their respective AI agent systems designed to detect and remediate software vulnerabilities. Wiz's Project Atlas achieved a 90.9% success rate on the CyberGym benchmark, a standard for evaluating AI's ability to find real-world flaws. Simultaneously, Microsoft's MDASH harness reported an even higher score of 95.95% on the same benchmark, outperforming several leading AI models.
These advanced systems leverage a sophisticated multi-agent, multi-model approach. Instead of relying on a single AI model, they route specific security tasks to specialized AI models best suited for that particular job. Wiz's Project Atlas, for instance, combines Claude Opus 4.6 with GPT-5.5, with plans to integrate Gemini. This strategy allows for more nuanced and accurate analysis, as different models excel at different aspects of vulnerability research, such as complex exploit chain reasoning versus precise triage.
Microsoft's MDASH employs a similar philosophy, combining internally developed agents like MAI-Cyber-1-Flash, which handles the majority of tasks, with the more powerful GPT-5.4 for complex issues. This tiered approach not only enhances detection and remediation capabilities but also contributes to cost efficiency. According to Mustafa Suleyman, CEO of Microsoft AI, this combination can reduce costs by up to 50% compared to using a single, large frontier model for all tasks.
Beyond benchmark performance, Wiz reported that Project Atlas has already uncovered over 200 zero-day vulnerabilities in widely used open-source code. This highlights the practical impact of these AI agents in proactively identifying critical flaws before they can be exploited by malicious actors. The ability of these systems to continuously scan and analyze codebases is crucial, as code changes rapidly and point-in-time analyses quickly become stale.
Both companies emphasize that the future of code security lies not in finding the single 'best' AI model, but in building architectures that can dynamically leverage the strengths of various models, adapt to new advancements, and ensure continuous, economical coverage. This approach ensures that security tooling remains effective as AI technology evolves.
While Project Atlas is currently used internally by Wiz, its success demonstrates the potential for AI agents to revolutionize vulnerability management. Microsoft's MDASH also showcases the company's commitment to integrating advanced AI into its security offerings, aiming to provide more robust and efficient protection against an ever-growing threat landscape.
The performance of these AI agents on the CyberGym benchmark, which includes models like OpenAI's GPT-5.5 Cyber and Google's Gemini 3.5 Flash Cyber, sets a new standard for AI-driven vulnerability discovery. The results suggest that a collaborative, specialized approach among AI models offers a significant advantage over monolithic solutions.
Ultimately, the development of systems like MDASH and Project Atlas signifies a paradigm shift in how software vulnerabilities will be discovered and addressed. By harnessing the power of multiple specialized AI agents, organizations can expect more comprehensive, efficient, and cost-effective security solutions in the near future.