Microsoft says its new cyber model helped the MDASH vulnerability-testing system score 95.95% on CyberGym while reducing operating costs by nearly 50%.
Microsoft will open Project Perception to public preview on Aug. 3, expanding its AI security offerings with coordinated agents that can identify weaknesses, investigate threats and make security fixes across a company’s systems.
The company also introduced MAI-Cyber-1-Flash, its first cybersecurity-specific AI model. Microsoft says adding the model to MDASH, its system for finding and fixing software vulnerabilities, helped the full system score 95.95% on CyberGym while reducing operating costs by nearly 50% compared with its current configuration.
The launch moves Microsoft beyond Security Copilot, which assists security teams through a chat interface. Project Perception’s agents can act across Microsoft security products, although the company says people will retain control over high-impact decisions.
Three teams of AI agents
Project Perception coordinates three types of agents.
Red agents look for possible ways an attacker could compromise a system. Blue agents investigate activity and decide which threats present meaningful risks. Green agents make corrections and strengthen protections.
The agents share information as they work, allowing a weakness identified by one agent to move through investigation and correction steps without requiring a person to transfer it between separate tools.
At launch, Project Perception will bring this coordinated system into Microsoft Defender. Microsoft said it plans to extend the system across additional security products over time.
Microsoft pairs its model with OpenAI
The 95.95% CyberGym result does not measure MAI-Cyber-1-Flash by itself. The score is for MDASH, which paired Microsoft’s MAI-Cyber-1-Flash with OpenAI’s GPT-5.4.
Microsoft said MAI-Cyber-1-Flash can handle up to 90% of the work, while MDASH sends the hardest 10% to the larger, more expensive OpenAI model. The company attributes its cost reduction to that division of work.
CyberGym tests whether AI systems can work through software vulnerabilities. In Microsoft’s benchmark chart, the combined MDASH system scored 95.95%, compared with 83.6% for GPT-5.6 Sol and 83.8% for Mythos 5.
Those are Microsoft’s own test results, and the chart compares different combinations of models and AI agents. It does not specify how the products would perform across all cybersecurity tasks.
AI security competition grows
Microsoft is expanding its security portfolio as other major AI companies introduce systems that find and correct software weaknesses.
Google recently added Gemini 3.5 Flash Cyber to its CodeMender security agent. OpenAI offers Codex Security and its broader Daybreak cyber program. Anthropic has introduced Claude Security and Project Glasswing.
Human approval for high-impact actions
Giving AI agents the ability to make security changes also poses the risk that an incorrect decision could disrupt systems or unnecessarily alter protections.
Microsoft says customers will set the agents’ objectives and limits, including what information and systems they may access. Every high-impact action will require human approval.
The company also says MDASH keeps activity isolated between customers, records actions for review, and runs model testing in protected environments without internet access.
Project Perception enters public preview on Aug. 3. Microsoft has not said when it will become generally available.

