OpenAI (via Future Tools)

Responding to the next frontier of critical cyber capabilities

Brief

Astra, an upcoming model, has shown recent advances in agentic coding and cybersecurity such that the developer states it cannot rule out that Astra has reached the Preparedness Framework’s Critical cyber capability level. The Framework defines Critical as the ability to discover and develop functional zero‑day exploits across many hardened real systems without human intervention, or to design and execute novel end‑to‑end cyberattacks from high‑level goals. In response, the company has scaled robustness testing and imposed stricter security controls—isolated testbeds, restricted network/tool access, model weight encryption, sandboxed execution—and paused internal work that doesn’t meet these protections. They implemented universal monitoring (including Chain‑of‑Thought detection that can trigger security interrupts), will coordinate with government and safety organizations, and will provide guidance to third‑party testers. The notice references prior use of the Framework (Dec 2023; actions in June 2025 for biology) and specifies Astra was not involved in the Hugging Face exploit.

Why it matters

Internal evaluations in the days before publication (article dated 2026-08-07) found Astra shows significant agentic coding and cybersecurity advancements; the company concluded “last night” it cannot rule out that Astra meets the Preparedness Framework’s Critical cyber capability threshold.

Key details

  • The Preparedness Framework’s Critical threshold is defined as either (a) identifying and developing functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or (b) devising and executing end-to-end novel cyberattack strategies against hardened targets given only a high-level goal.
  • Immediate mitigation steps include stricter security controls (isolated testing environments, restricted network/tool access, model-weight protections and encryption, enhanced monitoring and sandboxed execution), pausing internal Astra activities that don’t meet these controls, and universal monitoring of agentic applications (including Chain-of-Thought triggers) to interrupt risky behavior; the company says Astra was not involved in exploiting Hugging Face.
  • Context: the Preparedness Framework was first published December 2023; prior models such as GPT‑5.6‑Sol were assessed at the High (not Critical) level, and the same framework guided June 2025 bio-capability precautions.
Source evidence

Responding to the next frontier of critical cyber capabilities

Cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyberdefenses and enable attacks at unprecedented speed and scale.

Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework(opens in a new window).

We are sharing this because we believe it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.

We first published our Preparedness Framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level. We created it to give us a guide for identifying progress in capability and then planning what our company would do as those capabilities emerge. Previous models, including GPT‑5.6‑Sol, have been evaluated for frontier cyber capabilities and assessed at the High (rather than Critical) threshold.

Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal.

While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time. Astra is an upcoming model, and was not involved in exploiting Hugging Face.

Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities. Internally, we have also taken the following steps so that further development of this model happens safely and securely:

  • We are implementing stricter security controls for higher-capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution.
  • We are pausing internal activities involving Astra that do not yet meet these strengthened security control requirements.
  • We have implemented universal monitoring for risky actions and misalignment across all agentic applications of Astra, including training and evaluation. Monitors evaluate the model's Chain of Thought and trigger a security response to review and interrupt high risk activity.
  • We will work with relevant government agencies and select AI safety organizations to test the capabilities for this model.
  • We will be providing recommended security controls to third-party testing partners for running higher risk evaluations and workloads safely.

The framework has already guided us through other capability transitions. In June 2025, as our models approached the high capability threshold for biology under the Preparedness Framework, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We are applying the same principle here.

We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do. We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity.

Author

Keep reading

View all