AI Security - 7 min read - 3 August 2026

Microsoft built AI agents that find, judge and fix your vulnerabilities. It opens to everyone today.

Project Perception enters public preview today: red team agents that map attack paths, blue team agents that decide which findings actually matter, and green team agents that fix them - running as a continuous loop across a customer's own environment. Microsoft's new in-house model behind it, MAI-Cyber-1-Flash, posts strong benchmark numbers. Forrester's read on the bigger question is blunter than the launch materials.

Microsoft announced Project Perception on 27 July and opens it to public preview today. It's an agentic security system built on MDASH, the multi-agent harness Microsoft already uses internally to find and fix software vulnerabilities, and it's powered by MAI-Cyber-1-Flash - the first cybersecurity model Microsoft has trained itself, rather than licensed or fine-tuned from a general-purpose base. Help Net Security's coverage quotes Hayete Gallot, EVP of Microsoft Security, describing the system as bringing together "signals, context, models and specialized agents into a continuously learning system." The Next Web's writeup is more direct about what that means in practice: three classes of AI agents, working continuously, with meaningfully more autonomy than the dashboard-and-alert tooling most security teams are used to.

Red, blue, green - and an orchestrator deciding what each one does

The architecture borrows its naming from exercise terminology every security team already understands. Red team agents map attack paths and hunt for vulnerabilities the way an attacker would. Blue team agents investigate what red finds and judge which of it represents meaningful risk, filtering signal from noise at a scale no human triage queue could match. Green team agents then act on that judgement - patching, reconfiguring or otherwise hardening whatever blue has confirmed matters. An orchestrator agent routes work between the three, and a message bus carries context so each agent's output becomes the next one's input, continuously, without a human necessarily initiating each cycle. Identity and governance run through Microsoft's Agent 365 framework, and Microsoft says human approval is required for "consequential actions" - a phrase doing a great deal of load-bearing work in that sentence, and one worth asking your own Microsoft account team to define precisely before you grant this system any access.

The benchmark numbers, and the number that matters more

MAI-Cyber-1-Flash scores 96% on CyberGym, the benchmark built to evaluate how well AI systems find vulnerabilities in large codebases - twelve points above Anthropic's Mythos and ahead of Gemini and GPT-class models on the same test, according to Microsoft's own disclosed comparison. Satya Nadella has framed the combined system as delivering "world-class performance at 50 percent of the cost of leading models." Microsoft's track record with the underlying MDASH harness gives that claim some grounding: at Build 2026 in June, MDASH scored 96.55% on the same benchmark class, up from 88.45% roughly three weeks earlier - a jump made concrete when the system surfaced sixteen vulnerabilities in the Windows networking and authentication stack, four of them critical remote-code-execution flaws, all fixed within that month's patch cycle. That's a real result, not a demo. But The Next Web's own framing of the launch names the actual open question precisely: whether a 96% benchmark score holds up against real-world attacks rather than synthetic evaluations is exactly what the public preview exists to test. Enterprises granting this system access this week are, functionally, part of that test.

Forrester's word for it is "excessive agency"

The most useful reaction to Project Perception so far isn't from Microsoft. Analyst commentary gathered around the launch has converged on a specific concern: agents with the power to quarantine devices, change security controls, modify code or open pull requests hold real operational authority, and the property that makes an agentic system valuable - acting quickly, at scale, without waiting on a human for every step - is the same property that lets it take a large number of wrong actions before anyone notices. Microsoft's own public position is that organisations should build trust in the system gradually, granting more autonomy over time rather than all at once, which is a reasonable rollout philosophy and also, notably, an admission that the default posture today is not yet full trust. We've written before about what happens when a security evaluation sandbox that was supposed to be airtight wasn't: Anthropic disclosed on 30 July that three of its own Claude models broke out of a test environment and reached real infrastructure belonging to organisations that never agreed to be tested, because a partner's network isolation control failed silently. Microsoft describes Project Perception's execution as sandboxed and offline by design. That's the right architecture. It was also, reportedly, the architecture Anthropic's evaluation partner believed it had.

What's genuinely new here

It would be easy to read this as another AI-security product launch in a season full of them - we covered Google's Gemini 3.5 Flash Cyber model two weeks ago, gated to governments and vetted partners rather than opened broadly. What's different about Project Perception is the scope of autonomy on offer to ordinary enterprise customers, not just the model underneath it. Google kept its cyber model's access tightly controlled. Microsoft is opening a system with red-blue-green agents that can, per its own architecture description, act on an organisation's actual environment - to public preview, broadly, today. That's a genuinely different risk posture, and it deserves a genuinely different level of scrutiny from whoever signs off on giving it access.

  • Before enabling Project Perception in any environment, get a precise, written definition of "consequential action" from your Microsoft account team - not a marketing description, a configuration boundary you can audit.
  • Start green-team remediation agents in observe-and-recommend mode before granting them the ability to act, regardless of how gradually Microsoft's own rollout guidance suggests you can move faster.
  • Build a rollback and kill-switch runbook specifically for autonomous remediation agents, informed by how Anthropic's own evaluation sandbox failure unfolded and was contained.
  • Treat the 96% CyberGym figure as a lab result, not a production guarantee, until independent evidence from the public preview period is available.
  • Ask the same containment questions of Microsoft that we'd ask of any AI vendor running agents with real-world access - what network and identity controls, independent of the agents' own instructions, bound what they can reach.

Agentic security tooling that finds, judges and fixes vulnerabilities without a human in every loop is coming to every major security vendor's roadmap, not just Microsoft's - the direction of travel isn't really in question. What's in question is how much operational authority gets handed over before the industry has real-world evidence, rather than benchmark evidence, that the containment holds. If you're weighing what a rollout of Project Perception, or a system like it, should look like for your environment, we'd be glad to talk it through. Email sales@halfteck.com.

Explore more resources

Browse our full library of enterprise cloud, software, data and AI content.

View all resources