Microsoft AI (via Future Tools)

Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AI

Brief

MAI-Cyber-1-Flash inside MDASH is a Microsoft security-focused LLM built from the MAI-Thinking-1 lineage and integrated into a multi-agent vulnerability identification/remediation harness. Microsoft reports 96% on CyberGym—12 points above Mythos—and says the compact model can handle roughly 90% of tasks, reserving GPT-5.4 for the hardest ~10%, producing an overall 50% cost saving versus their previous MDASH configuration (GPT-5.4 + 5.4 mini + 5.3 codex). MDASH comprises 100+ expert-tuned agents and feeds Project Perception, a new agentic security system for continuous monitoring and patching. The release emphasizes safety: security-first training, AI Red Team and third-party assessments, and enterprise controls (RBAC, tenant isolation, encryption, auditability, sandboxed no-internet execution). Microsoft cites a live reinforcement loop informed by more than 100 trillion daily signals and 1.6 million customers to continuously improve the models.

Why it matters

MAI-Cyber-1-Flash, a compact code-heavy model from the MAI-Thinking-1 lineage, is integrated into MDASH and achieves 96% on the CyberGym benchmark—+12 points over Mythos—and reportedly outperforms Mythos, Gemini, and GPT.

Key details

  • The combined MDASH+MAI-Cyber-1-Flash system handles up to 90% of vulnerability tasks, delegating the ~10% hardest cases to larger models (e.g., GPT-5.4), which enables a claimed 50% cost reduction versus the prior MDASH stack (GPT-5.4 + 5.4 mini + 5.3 codex).
  • MDASH is a multi-agent harness with 100+ expert-tuned agents; Microsoft is also launching Perception, an agentic security system to continuously monitor, patch, and close threat vectors and to expand use of MAI-Cyber-1-Flash.
  • Security and trust controls include security-first model calibration, evaluation by Microsoft’s AI Red Team and a third party, enterprise features (RBAC, tenant isolation, encryption, audit logs, sandboxed no-internet execution), and reinforcement from Microsoft’s operational data (over 100 trillion daily security signals and 1.6 million customers).
Cleaned source text

Introducing MAI-Cyber-1-Flash inside MDASH

Today we’re announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness. Together they deliver world-class performance at 50% of the cost of leading models.

Progress in AI has been startling and so has the new generation of cyber threats it’s unleashing. Attackers now wield increasingly powerful capabilities, probing an ever-growing mountain of code for just a single weakness that lets them in.

As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete. If we’re to unlock the true benefits of AI, we must first build outstanding cyber models that help all of us harden the software the world runs on.

That’s the motivation behind MAI-Cyber-1-Flash, which has been built to find challenging vulnerabilities in complex codebases. It’s been deeply integrated into MDASH, honed by the best cybersecurity experts in the industry and hardened across the largest security estate on the planet.

This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.

Picking the right model for the task

Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. MAI-Cyber-1-Flash was designed to efficiently handle up to 90% of all tasks, enabling MDASH to use the larger and most costly models in our fleet (in this case GPT-5.4) for the 10% of exceptionally hard tasks that truly need them.

The result is that the unified system of MDASH with MAI-Cyber-1-Flash delivers 96% on CyberGym (+12 pt above Mythos).

This combination delivers a 50% cost saving when compared against our best offering in MDASH today (GPT 5.4 + 5.4 mini + 5.3 codex). That’s the power of a well-tuned, multi-model system with access to uniquely rich historical training data. It ensures you always have the best model at the best price for every task.

In this new environment, being able to go from identifying a new vulnerability to addressing it in real-time is critical. And while AI remediation of software vulnerabilities is now a key security workflow, there are many jobs to be done by Security practitioners themselves.

That’s why today we’re also launching Perception, our agentic security systems, that provides teams of agents for a variety of security workflows in MDASH, to continuously monitor, patch, and close new threat vectors. Perception will also soon use MAI-Cyber-1-Flash for many more security workflows, beyond the software vulnerability work.

Three things matter today: Model. Data. Harness.

We have jointly optimized our world-class models, our unmatched historic data, and our expert-tuned harness to ensure that our customers have a uniquely powerful security offering.

Model. MAI-Cyber-1-Flash is a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data. Details in our technical report.

Data. Our deepest advantage. Decades of building world-class security systems now give us trillions of daily signals across identity, endpoint, cloud, and network, and an unmatched record of real exploits and remediations. No one can manufacture this history.

Harness. MDASH, our multi-agent vulnerability identification and remediation harness, is tuned by the best security experts in the industry, who have created 100+ agents using multiple leading models to find, validate, and remediate vulnerabilities. Agentic code scanning is a critical function in the Security Operating Center and feeds Project Perception, our new agentic security system.

Built with safety first

Because MAI-Cyber-1-Flash is Microsoft’s first cyber model, we built trust into every layer of the system, from model training to customer deployment. The model was developed with a security-first calibration, rigorously evaluated by Microsoft’s AI Red Team, tested through automated and expert-led adversarial exercises, and independently assessed by a third party.

Trust extends beyond the model itself. Through MDASH, customers get enterprise-grade controls including Role-Based Controls, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access. The result is a cyber model that delivers powerful capabilities to defenders while maintaining the governance, security, and control enterprises expect from Microsoft.

Our hill-climbing machine

Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop. Every day, defenders investigate threats, triage alerts, hunt adversaries, remediate vulnerabilities, deploy protections, and learn from the outcome.

Microsoft sees that loop end to end: vulnerabilities through Microsoft Security Response Center; attacks and defenses across identity, endpoint, cloud, data, browser, and applications; more than 100 trillion security signals every day; and operational insight from 1.6 million customers. Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data.

Our MAI reinforcement learning loop gives us the foundation to build cyber models that improve continuously and become expert cyber defenders. That’ll remain our commitment to our customers for years to come.

Build the Future With Us

We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way to do it at all. If our mission resonates with you, we’d love to talk.

Explore all jobs