AI Security - 6 min read - 23 July 2026

Google built a bug-hunting AI that beats bigger models for less money. You still can't have it.

Gemini 3.5 Flash Cyber is small, cheap to run, and - on Google's own numbers - competitive with far larger frontier models at finding and fixing vulnerabilities. It's also not available to you. Google is routing it exclusively through CodeMender to governments and a short list of trusted partners, a deliberate gatekeeping decision that says as much about where enterprise AI security is headed as the model itself.

Google DeepMind used this week's model refresh to slip in something unusual alongside the expected consumer-facing updates. Alongside Gemini 3.6 Flash and 3.5 Flash-Lite, the company introduced Gemini 3.5 Flash Cyber, a variant fine-tuned specifically to find, validate and patch software vulnerabilities. According to Google's own announcement, the model runs inside CodeMender, its automated code-security agent, where multiple Flash Cyber instances work in parallel across a codebase before compiling their findings into a single report. On the CyberGym benchmark, Google says the model "reaches competitive performance at the frontier" against much larger, more expensive models, while costing a fraction as much to run per query. The Hacker News' coverage and Help Net Security's writeup both zero in on the same demonstration result: in one internal test, the CodeMender pipeline uncovered a remote-code-execution flaw in a public-facing API and a memory-corruption bug in a production service in around two hours, then produced a working exploit reliable enough to bypass standard mitigations.

The interesting decision isn't the model, it's the access list

A cheap, fast, benchmark-competitive vulnerability hunter would ordinarily be exactly the kind of capability Google ships broadly through its API within weeks of announcement. Instead, per Google's own post, Flash Cyber is going out through a "limited-access pilot programme" restricted to governments and vetted partners, with no stated timeline for wider availability. GCN's reporting frames this explicitly as a dual-use call: a model this good at finding exploitable bugs is, by the same token, a model this good at finding bugs to exploit, and Google appears to have concluded that broad release would hand attackers a capability upgrade at least as large as the one it hands defenders.

That's a defensible position, and arguably the responsible one. It's also a preview of a pattern enterprise security and procurement teams should get used to: the most capable frontier security tooling increasingly won't be something you can simply buy a seat for. It'll be gated behind government relationships, partner tiers, or vetting processes that most enterprises, however large, don't automatically qualify for. Budget cycles and vendor shortlists built around "wait for it to become generally available" are going to keep running into capabilities that are deliberately never generally available.

What this actually changes for a security programme today

Nothing you're running in production gets safer or riskier because Flash Cyber exists; it's not in your stack and, for most organisations, won't be soon. What does change is the baseline attackers are working from. If a model this size can find a reliable RCE exploit in two hours inside a controlled pilot, it's a reasonable planning assumption that comparably capable tooling - whether from a frontier lab's less carefully gated release, an open-weights model tuned the same way, or a criminal group's own effort - reaches attackers on a similar timeline to defenders, possibly faster, since attackers don't need Google's permission to fine-tune something similar on stolen or public training data.

The practical response isn't to wait for Flash Cyber-class tools to arrive on your side of the fence. It's to assume the attacker side of the equation is already moving faster than most patch cycles, and to close the gap that actually matters: the time between a vulnerability existing in your code and it being found and fixed, by anyone. We've written before about AI model evaluation and guardrails as a discipline that has to keep pace with what models can do, not what they could do six months ago; a benchmark-competitive vulnerability hunter landing in a single announcement is exactly the kind of capability jump that makes a stale evaluation framework dangerous. It's also a strong argument for treating AI vendor due diligence as an ongoing exercise rather than a one-off checklist completed at contract signature - the labs you're already trusting with production workloads are the same labs building this generation of dual-use tooling, on timelines your due diligence cadence may not track.

  • Ask your AI vendors directly whether they operate anything comparable to CodeMender internally, and if so, whether it's used to harden the products you buy from them before those flaws reach customers.
  • Treat "time to patch" as a metric that now has to compete against AI-assisted discovery timelines measured in hours, not the multi-week cycles most vulnerability management programmes are still built around.
  • Review whether your own application security tooling has any AI-assisted fuzzing or exploit-generation capability, and if not, budget for it as a near-term gap rather than a future nice-to-have.
  • Map which of your vendors or partners might plausibly qualify for gated pilot access to frontier security tooling like this, since an indirect route through a partner's stack may arrive well before general availability does.

Google drawing a hard line around who gets Flash Cyber is a sensible call on Google's part. It doesn't do anything to slow down whoever is building the version that isn't gated. If you'd like help stress-testing whether your vulnerability management cadence can survive an AI-accelerated discovery timeline, email sales@halfteck.com.

Explore more resources

Browse our full library of enterprise cloud, software, data and AI content.

View all resources