OpenAI Rated Its Own Model "Critical" for Cyber. Google and Anthropic Moved the Same Week.

On September 1 and 2, 2026, three frontier labs shipped their most capable cybersecurity models and put every one of them behind a trusted-access gate. The capability is the headline. The access list is the actual safeguard, and that is the part worth reading carefully.

What Actually Happened, and When

On September 1, 2026, OpenAI classified its unreleased Astra model at the Critical cybersecurity capability level of its Preparedness Framework. No model from any lab had been placed in that category before.

The following day, Google announced Gemini 3.8 Flash Cyber, which it describes as its most capable cybersecurity model, released through a trusted-defender programme called Fairwind. Anthropic had shipped Claude Fable 5.1 and Claude Mythos 5.1 on September 1, with Mythos restricted to vetted organisations working in cybersecurity and the life sciences.

Three frontier labs, two days, one shape: the most capable cyber model each of them has built, and a gate in front of it.

100% Astra's reported score on ExploitBench, which measures turning a known vulnerability into a working exploit
2 Zero-day vulnerabilities Astra discovered independently during evaluation
91.5% Cyber jailbreak attempts Astra refused, against 59% for GPT-5.6 Sol
650+ Partners in Google's Fairwind trusted-defender programme

What "Critical" Means in This Framework

The Preparedness Framework is OpenAI's internal risk classification, first published in December 2023. It scores frontier models across domains including cybersecurity and biological threats, and it is the mechanism that decides whether a model ships freely, ships with restrictions, or does not ship at all.

Critical is the top of the cyber scale. OpenAI defines it as a model that can independently discover and exploit previously unknown vulnerabilities across numerous well-defended systems, or plan and execute a complete cyberattack against a hardened target from a high-level instruction.

Read that definition twice, because the load-bearing word is independently. Every capability described in the classification is one the model performs without a human directing each step.

The Evaluation Results

According to OpenAI's own reporting on the model, Astra:

On the safety side, Astra refused 91.5% of attempts in a cyber jailbreak evaluation, against 59% for GPT-5.6 Sol, and showed materially reduced tendencies to circumvent restrictions or take the bait on deliberately planted honeypot targets.

Every one of those numbers is self-reported. They come from OpenAI's measurement of its own unreleased model. There is no third-party confirmation, and the system card will not be published until launch. OpenAI has said it will request evaluations from government agencies and independent AI safety organisations before release. Until those exist, treat the figures as the vendor's claim about the vendor's product, which is exactly how you would treat any other security vendor's internal benchmark.
01
The Structural Read

The Safeguard Is No Longer the Model. It Is the Access List.

Alignment training moved Astra's cyber refusal rate from 59% to 91.5%, which is a real engineering result. It also was not enough to ship the model openly. What actually keeps the strongest capability away from the wrong hands is not a property of the weights, it is a credential, issued to a named organisation and revocable by the vendor. That is a meaningful shift. For most of the modern AI era, safety was framed as something you train into a model. As of this week, at the frontier of cyber capability, safety is a distribution decision. And distribution decisions fail in ways training does not: leaked keys, insider access, a partner with weaker controls, an approved account used for unapproved work, and eventually an open-weights model that reaches the same capability with no gate at all.

Three Labs, Three Gates

The interesting thing about this week is not that one lab drew a line. It is that all three drew roughly the same line, independently, within forty-eight hours.

OpenAI: Daybreak

Daybreak is OpenAI's Trusted Access for Cyber programme. It has two published tiers. Daybreak Blue provides access to GPT-5.6 Sol without its system-level cyber guardrails, for defensive work. Daybreak Red provides GPT-5.6-Cyber for exploit validation and advanced vulnerability research.

Astra's strongest cyber functions go first to a small alpha group, described as organisations responsible for defending critical digital infrastructure, including the U.S. government. Broader defensive access follows through Daybreak Blue, whose partners include Cisco, Cloudflare and Palo Alto Networks. OpenAI has not named the alpha testers, published selection criteria, or given a launch date beyond "soon".

Google: Fairwind

Gemini 3.8 Flash Cyber is available through the Fairwind Program to trusted defenders including governments, healthcare providers and telecommunications operators. Google reports over 650 partners globally, naming CrowdStrike, Datadog, Menlo Security, Palo Alto Networks and Snowflake among them.

Google's framing is the most explicitly defensive of the three: it says it prioritised vulnerability fixing from the start, and prioritised it over offensive capabilities like exploitation. That is a training-objective claim, not just a policy one, and it is the sort of claim that becomes checkable once independent evaluators get access.

Anthropic: Project Glasswing

Claude Fable 5.1 now permits software vulnerability identification, a category that was previously restricted. Claude Mythos 5.1 sits behind trusted access for cybersecurity and life sciences work, and certain tasks, including penetration testing, exploit generation and binary vulnerability scanning, are redirected to Opus models rather than served directly.

Anthropic also introduced Enterprise Frontier Safeguards, pairing zero data retention with misuse detection, and reports sandbox escape detection plus changes to reward metrics aimed at alignment failures. It describes Mythos 5.1 as its most robust model to date on an external prompt injection benchmark, which is one of the few figures in this entire week measured against something the lab did not build itself.

Why simultaneity matters more than any individual announcement. Three competitors with different safety philosophies, different commercial incentives and no obligation to coordinate arrived at the same architecture in the same week: build the capability, measure it, refuse to ship it openly, and hand it to a vetted list of defenders first. When rivals converge on a control they all pay a commercial price for, that is usually evidence the underlying risk assessment is real rather than performative.

The 8.5 Percent

Astra refuses 91.5% of cyber jailbreak attempts. The complement is 8.5%, and complements are where security lives.

A refusal rate is a probability, and probabilities do not survive contact with volume. At one attempt per day, 8.5% is a rounding error. At ten thousand automated attempts, trivial to generate, cheap to run, and precisely the kind of thing an adversary automates first, it is 850 successes. Anyone who has watched a credential-stuffing campaign against a login page already understands this arithmetic: the per-attempt success rate stops mattering once attempts are free.

This is not an argument that 91.5% is bad. Moving from 59% is a substantial improvement, and no realistic target is 100%. It is an argument that a refusal rate is a filter, not a wall, and that the security model has to assume the filter gets crossed rather than assume it holds.

02
The Operational Read

Exploitation Just Stopped Being a Scarce Skill

For thirty years, defensive prioritisation has rested on a quiet assumption: that turning a disclosed vulnerability into a working, chained, reliable exploit requires scarce human expertise, and that the scarcity buys defenders time. Severity scores, patch windows, "low severity, deprioritise", all of it is downstream of the belief that most vulnerabilities will never be weaponised because nobody competent enough will bother. A model that chains OS flaws to root without supervision does not eliminate that assumption today, because the capability is gated. It establishes that the assumption has an expiry date, and that the date is now a function of access policy rather than of human talent supply.

What Would Actually Change the Picture

Three developments would move this from a policy story to an operational one, and they are worth watching in this order.

Independent evaluation that confirms or contradicts the numbers. Right now the entire capability claim rests on vendor self-measurement of an unreleased model. When government agencies and independent safety organisations publish their own results, the story either hardens or deflates. Nothing else in this space matters as much as that first external data point.

An open-weights model reaching comparable capability. Gating works because the capability is scarce and centrally controlled. Open-weights models have historically trailed the frontier by roughly twelve to twenty-four months on most capability axes. Whenever this particular gap closes, every access-list control described in this article stops functioning at the same moment, and there is no version of the policy that survives it.

Documented misuse from inside a trusted-access programme. The gates are new and the credentials are held by hundreds of organisations. The first confirmed case of an approved account being used for unapproved work, whether through compromise, insider action, or a partner's weaker controls, will determine whether trusted access is treated as a durable control or as a speed bump that bought eighteen months.

What This Means If You Defend a Network

1. Your Patch Window Was Sized for Human Attackers

Most remediation SLAs encode an implicit exploitation delay: the gap between public disclosure and a reliable weaponised exploit circulating. That gap is a human-labour estimate. If the labour becomes automated and continuous, the gap compresses toward zero and the SLA stops meaning what it meant when it was written. This is not a reason to panic-patch everything. It is a reason to re-derive your windows from an explicit assumption about exploitation speed, and to write that assumption down where it can be revised when the evidence changes.

2. "Low Severity" Was Always a Bet on Effort

Deprioritisation logic assumes an attacker weighs effort against reward. Chained privilege escalation through several individually unremarkable flaws is exactly the work that was previously too tedious to be worth doing at scale. If your triage model treats low-severity findings as permanently safe rather than as cheap to chain, the chaining capability demonstrated here is aimed directly at that assumption.

3. Trusted Access Is a Control You Can Ask Vendors About

If your security stack includes CrowdStrike, Palo Alto Networks, Cloudflare, Cisco, Datadog or Snowflake, your vendors are inside one or more of these programmes. That is a reasonable question for a vendor review: which frontier cyber models do you have access to, what do you use them for, how is that access controlled on your side, and what is your exposure if the credential is misused. It is the same question you would ask about any other privileged integration.

4. Defensive Capability Arrives Before Offensive Parity, So Use the Window

The entire design of these programmes is to put the capability in defenders' hands first. That is a genuine, deliberate advantage, and it is temporary by construction. Organisations that spend the next twelve months getting AI-assisted vulnerability discovery into their actual remediation pipeline will be in a different position from organisations that read the announcement and filed it. The gap between those two groups is the whole point of the head start.

5. Expect Friction, and Plan for It

OpenAI has said it expects its safeguards to over-trigger at launch: blocking legitimate security work, prompting review on flagged actions in ChatGPT, and interrupting tasks through the API. If you are building security tooling on these models, assume false positives in the safety layer are part of your operating envelope rather than an anomaly to be surprised by in production.

The thing to watch is not the next model. It is the first independent evaluation of this one. Every number in this story, the 100%, the two zero-days, the 91.5%, was produced by the company that built the model, about a model nobody outside can test. The classification may be exactly right. It may also be conservative, or generous, and there is currently no way for anyone outside these labs to tell the difference. The credibility of the entire trusted-access framework rests on that gap being closed.

Frequently Asked Questions

What is OpenAI's Astra model?
Astra is an OpenAI frontier model that, as of September 1, 2026, has been classified at the Critical cybersecurity capability level of the company's Preparedness Framework, the first model from any lab to reach that tier. It has not been publicly released. OpenAI has said it is coming soon, with the strongest cyber functions restricted to a small alpha group at launch and broader defensive access following through the Daybreak Blue programme.
What does the "Critical" cybersecurity threshold actually mean?
In OpenAI's Preparedness Framework, Critical is the top cyber tier. It applies to a model that can independently discover and exploit previously unknown vulnerabilities across numerous well-defended systems, or plan and execute a complete cyberattack against a hardened target from a high-level instruction. The defining element is autonomy: the model performs the work without a human directing each step. Reaching the tier mandates additional safeguards before any release.
Can Astra really find zero-day vulnerabilities on its own?
According to OpenAI's own evaluations, yes. The model discovered two previously unknown vulnerabilities independently in tests built around recently disclosed flaws, scored 100% on ExploitBench, escaped a browser sandbox to execute commands on the host, and chained multiple operating system flaws to reach root. All of these results are self-reported on an unreleased model. No third party has verified them, and the system card is not scheduled for publication until launch.
What are Daybreak Blue, Fairwind and Project Glasswing?
They are the three labs' trusted-access programmes for cybersecurity work. OpenAI's Daybreak has a Blue tier for defensive work, providing GPT-5.6 Sol without system-level cyber guardrails, and a Red tier providing GPT-5.6-Cyber for exploit validation and advanced vulnerability research, with partners including Cisco, Cloudflare and Palo Alto Networks. Google's Fairwind Program serves over 650 partners including CrowdStrike, Datadog, Menlo Security, Palo Alto Networks and Snowflake. Anthropic's Project Glasswing gates Claude Mythos 5.1 for vetted cybersecurity and life sciences organisations.
Is a 91.5% jailbreak refusal rate good enough?
It is a large improvement over the 59% recorded for GPT-5.6 Sol, and no realistic system reaches 100%. The operational problem is that a refusal rate is a probability, and attack attempts are cheap to automate. At ten thousand attempts, an 8.5% compliance rate is 850 successes. The rate should be read as a filter that reduces volume, not as a boundary that holds, which is why access control rather than refusal training is doing the real work in these releases.
Should security teams change anything right now?
Nothing has to change today, because the capability is gated and unreleased. What is worth doing now is re-deriving remediation timelines from an explicit, written assumption about how fast a disclosed vulnerability becomes a working exploit, revisiting triage rules that treat low-severity findings as permanently safe rather than cheap to chain, asking security vendors which frontier cyber programmes they participate in and how that access is controlled, and getting AI-assisted vulnerability discovery into the remediation pipeline while defenders still have first access.

My Take

The capability result is not what stayed with me. Autonomous vulnerability discovery has been visible on the trend line for two years, and anyone tracking it expected a model to cross this line in 2026. It arrived roughly on schedule.

What stayed with me is that alignment training ran out of road, and everybody involved said so through their actions rather than in a blog post. Astra refuses 91.5% of cyber jailbreaks and still could not ship openly. Google trained fixing over exploitation and still routed the model through a partner list. Anthropic redirects exploit generation to a different model family and still gates Mythos behind vetted access. Three labs, three approaches to making the model behave, three identical conclusions about whether that was sufficient on its own.

So the real control is now a list of names, and lists of names have a well-documented failure mode. Every organisation reading this manages privileged access for a living and knows exactly how that story ends: not with a dramatic breach, but with a credential in the wrong place for a reason nobody thought to prohibit.

The part I am least comfortable with is the evidence base. We are being asked to accept a Critical classification, a set of benchmark results and a safeguard efficacy figure, all measured by the vendor, on a model no outsider can test, with the system card deferred to launch. That is not a reason to disbelieve OpenAI. It is a reason to notice that the most consequential safety claim of the year currently carries the same evidentiary standard as a product datasheet, and that the labs themselves have said external evaluation is coming.

The genuinely good news is the sequencing. Defenders got access first, deliberately, and that ordering was a choice these companies made at real commercial cost. It is the best-designed part of the whole week. It is also a head start, and head starts are only worth what you do with them.

Your patch window assumes exploitation takes a human weeks. What is it built on if it takes an agent an afternoon?

Related Articles:

Kodjo Apedoh

About the Author

Kodjo Apedoh

Network Engineer & AI Entrepreneur

Founder of TechVernia & SankaraShield. Certified Network Security Engineer with 4+ years of experience specializing in network automation (Python), AI tools research, and advanced security implementations. Also builds iOS and Android applications. Holds certifications from Palo Alto Networks, Fortinet, and Cisco. Based in Arlington, Virginia.

Connect on LinkedIn →