OpenAI's Preparedness Framework has four cybersecurity tiers, and until this week every model the company had shipped or tested sat at "High" or below - including, per prior reporting, GPT-5.6-Sol. Astra is the first to clear the "Critical" bar, which OpenAI defines as a model that can identify and develop functional zero-day exploits across all severity levels in hardened real-world systems without human intervention, or independently devise and execute end-to-end attack strategies against a hardened target given only a high-level goal. According to OpenAI's own announcement and coverage from Security Boulevard, Astra met that bar in internal red-teaming, not in the wild - but the framework exists precisely so that "met it in testing" triggers a response before "met it in the wild" becomes possible.
What actually tripped the threshold
- A 100% score on ExploitBench, OpenAI's internal benchmark for exploit-development capability.
- Independent discovery of two previously unknown zero-day vulnerabilities during evaluation, out of a pool of roughly twenty high-severity targets.
- Execution of full browser-compromise chains and sandbox escapes without a human operator walking it through intermediate steps.
- Demonstrated privilege escalation from a standard foothold to root/administrator access.
None of that happened against a live production target - it happened in the kind of controlled evaluation environment OpenAI and outside safety researchers use specifically to find this out before a model ships. That's the system working as designed. It's also, per Forbes' reporting from earlier in August, the reason OpenAI paused Astra's development track weeks before this week's formal announcement rather than waiting for the classification to become official first.
Gating access instead of gating capability
OpenAI isn't undoing what Astra can do - that would mean retraining a materially weaker model. Instead it's restricting who can reach the capability and building monitoring around how it's used: chain-of-thought monitoring intended to catch and interrupt unauthorized action sequences, and safety training that the company says lifted Astra's refusal rate on cyber-jailbreak attempts from 59% to 91.5%. Initial access is limited to select internal testers and participants in OpenAI's Daybreak Blue early-access tier - a narrower, seemingly more restricted programme than the Daybreak Red tier we covered when GPT-5.6-Cyber launched to vetted defenders in August. A model good enough to be classified Critical apparently needs a tighter circle than one merely good enough to find real zero-days for a bug bounty programme.
A 91.5% refusal rate is not a control
The improved refusal rate is a real mitigation and a genuine engineering achievement, but it shouldn't be mistaken for the reason access is restricted. A model that refuses 91.5% of jailbreak attempts still complies with roughly one in eleven, and at Critical-tier capability, one compliance is the only one that matters to whoever's on the receiving end of the exploit it produces. That arithmetic is presumably why the access controls - vetting, tiering, monitoring - are doing the actual containment work here, with the refusal-rate improvement functioning as a second layer rather than the primary one.
The pattern this extends
Astra's classification lands in a year that's already seen Anthropic, Google and OpenAI itself publish escalating capability disclosures for cyber-relevant model behaviour - we've tracked several in Anthropic's agentic test-environment breakout and Google's Gemini vulnerability-finding tool. What's different about Astra is that it's the first to trip a frontier lab's own "Critical" line rather than a "notably capable" one, which moves the conversation from "should defenders get this" to "who, specifically, and under what monitoring." Security teams evaluating any frontier-model vendor relationship should expect that question to arrive with increasing frequency, not decreasing.
- Ask AI vendors directly which Preparedness/responsible-scaling tier their current and next-generation models sit at for cyber capability, not just whether they have a framework.
- Treat "gated access" announcements as an access-control decision to evaluate, not a capability limit to assume - the underlying model's ability hasn't changed.
- Factor frontier-model cyber capability growth into your own red-team and detection engineering roadmap now, rather than after the first Critical-tier model reaches general availability.
- Watch for the gap between a lab's internal refusal-rate metrics and real-world jailbreak resilience - the two are related but not identical.
If your organisation is trying to work out what escalating frontier-model capability means for your own AI governance and vendor risk posture, email sales@halfteck.com.