Astra Hits OpenAI’s Critical Cyber Threshold—Most Exploit Power Stays Gated

OpenAI concluded on September 1, 2026, that Astra meets its Critical cybersecurity capability threshold, making it the company’s first model to receive that designation. The finding means that, with suitable tools and access, Astra can discover unknown vulnerabilities and develop exploits across many well-protected systems without step-by-step human direction, as detailed in OpenAI’s September 1 assessment.
The designation does not give every user those capabilities. Axios independently reported that the most advanced cyber functions would initially be limited to a small group of testers—and that the safeguards could also interrupt legitimate work by businesses, government agencies and other defenders.
What the Critical threshold actually means

Critical is a capability classification under OpenAI’s Preparedness Framework, not a description of what Astra will do in every product session. The cyber threshold covers the ability to identify and develop functional zero-day exploits across multiple hardened, real-world systems, or to execute a novel end-to-end attack against a hardened target from a high-level objective.
In controlled evaluations, Astra found previously unknown flaws in a hardened browser and operating system, converted them into working exploit chains, escaped a browser sandbox and built a local privilege-escalation chain. It also discovered and used two zero-day vulnerabilities in an internal benchmark; disclosure to the affected maintainers was still in progress when the assessment appeared.
The tested configuration is an important qualification. The most advanced results came from Astra with Daybreak Blue access, not its default production configuration. The Critical label therefore describes capability that can be elicited under sufficiently equipped conditions, while deployment controls determine which users can invoke it.
Astra’s access matrix

The dividing line is based on access level, account risk and authorization rather than a simple distinction between “defensive” and “offensive” prompts. Vulnerability validation and proof-of-concept development are dual-use activities: they can establish the urgency of a patch or provide a route into a target.
- General users: Code review, vulnerability analysis and patching remain available, but standard safeguards can refuse advanced proof-of-concept creation, chained exploitation and other high-risk assistance.
- Higher-risk accounts: A more conservative behavior boundary can block a broader range of dual-use requests. The full set of signals used to assign that treatment is not public.
- Selected testers: A small alpha group receives early access to advanced cyber workflows under tighter controls, allowing safeguards and misuse risks to be evaluated before wider expansion.
- Verified defenders: Organizations may apply for Daybreak access, while individuals may verify their identity and request trusted access. Intended work includes understanding unfamiliar code, validating vulnerabilities, conducting authorized red-team exercises, and developing or testing patches.
- Prohibited workflows: Trusted status is not blanket permission to develop novel attacks against hardened targets or act outside an authorized scope. Monitoring, policy restrictions and enforcement continue to apply.
The difference between configurations is measurable. The GPT-6 Astra system card, published September 3, shows Daybreak Blue raising Astra’s completion rate for proof-of-concept exploit creation from 2.4% to 92% and its cyber red-teaming completion rate from 7.4% to 76.9%; even with that access, Astra fully completed only 3.5% of requests in a separate evaluation of arbitrary advanced cyber tasks.
Those figures clarify what “gated” means. Standard access suppresses much of the functionality needed to turn a suspected flaw into a working exploit, while trusted access restores substantially more capability for defined defensive uses without opening every cyber workflow.
The final designation follows an earlier slowdown in Astra’s development, when preliminary evidence prompted additional evaluation and stronger controls. The later decision changed the status from a risk that could not be ruled out to a capability the company formally determined Astra possessed.
Why legitimate security work can be blocked

Safeguards cannot reliably infer authorization from technical content alone. A penetration tester validating a client’s exposure and an attacker preparing an intrusion may examine the same function, issue similar commands and request the same exploit technique. When ownership or scope is unclear, a conservative boundary can stop both.
That creates friction at pivotal stages of a defensive investigation. General access may help locate and explain a vulnerability but refuse the proof of concept needed to demonstrate exploitability, rank remediation work or confirm that a patch closes the real attack path. Smaller teams that do not obtain trusted access may feel this gap most directly.
Long-running agent work faces a separate constraint. Monitoring can inspect an agent’s reasoning and actions for indications that it is operating outside the user’s authorization; a flagged task may be paused or ended even when its broader purpose is legitimate. Work that is not obviously cybersecurity-related can also be caught if an intermediate action resembles credential access, unauthorized data movement or circumvention of a control.
The effect varies by product surface. A paused task in ChatGPT or Codex may be presented for user review before it continues, whereas some stopped API tasks cannot be resumed. For an automated security pipeline, that can mean lost execution context and an incomplete investigation rather than a simple refusal to answer one prompt.
What the safeguards do—and what remains unknown
Astra’s controls operate at several layers: model refusals, system classifiers, account-level risk measures, offline detection and monitoring of extended agent activity. They address both malicious use and the possibility that an agent takes an unauthorized action while pursuing an otherwise permitted objective.
These controls do not remove the underlying Critical-level capability. They regulate when it can be reached and can intervene when behavior crosses a risk threshold. Conservative settings can therefore reduce severe misuse while still producing false positives that slow benign work.
The current position is neither unrestricted access nor a complete cyber lockout. Astra is broadly deployed, but advanced exploit-oriented functionality remains divided among standard access, stricter treatment for higher-risk accounts, selected testing and the expanding Daybreak program.
No public timetable establishes when trusted access will become broadly available, how many smaller teams will qualify or how often legitimate workflows will be interrupted outside evaluations. Those deployment outcomes will determine whether Astra’s access model preserves enough utility for defenders while keeping its most consequential exploit capabilities meaningfully restricted.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.