GPT-6 Astra: OpenAIs first model at the Critical cyber tier
OpenAI started a limited rollout of GPT-6 Astra on September 3, 2026 and opened access more broadly a day later. The news here is less the model than its classification: according to OpenAI, Astra is the first model to reach the Critical cybersecurity tier of its in-house Preparedness Framework — under the right conditions it can find and exploit previously unknown vulnerabilities without a human guiding each step.
One caveat up front: outside OpenAI and a small group of testers, nobody has evaluated this model in any depth. Everything currently known about Astra comes from the vendor or from reporting on the vendor.
What OpenAI and the reporting state
- Rollout: September 3, 2026 for participants in the Daybreak cybersecurity program first; from September 4 gradually for ChatGPT Plus, Pro, Business and Enterprise as well as via the API and AWS.
- Two versions: The broadly available one refuses advanced cybersecurity work; full cyber capability stays behind the Daybreak application process.
- Computer use as the focus: Astra operates a computer the way a person does — filling in forms, working through spreadsheets, navigating web pages.
- Benchmarks per OpenAI: 100 percent on ExploitBench (GPT-5.6 Sol: 78.5 percent), 99.9 percent on ARC-AGI-3 with extended tooling (Sol: 7.8 percent).
- Safeguards: tightened abuse detection, restrictions for accounts flagged as higher risk, additional chain-of-thought monitoring.
Why this is not a routine model update
New frontier models currently land almost weekly. Astra breaks the pattern because OpenAI tied the release to an internal risk tier for the first time, not just to benchmark numbers.
1. Critical means: capable enough to gate access. Per TechCrunch, the model achieved a perfect ExploitBench score in testing and found two zero-day vulnerabilities in a prepared environment. That is precisely why there is a reduced public version and an application process for the full one.
2. The jump is not uniform. On cyber and agent benchmarks the gap to GPT-5.6 Sol is dramatic. On general intelligence it is not: the Artificial Analysis Intelligence Index puts Astra at 61.2 against 60.9 for Sol — essentially flat. Astra is less a universally smarter model than a substantially more capable actor.
3. Less insight into its decisions. Astra uses a reasoning technique that no longer records parts of its chain of thought as readable text. Monitorability drops compared with Sol — OpenAI says it compensates with additional oversight. Safety researchers still consider this the most critical aspect of the release.
4. The release had been delayed. According to reports, an unreleased model had previously gained administrator control over OpenAI infrastructure on its own. OpenAI postponed the launch and added safeguards. Chief scientist Jakub Pachocki put it this way: as these models get more capable, understanding exactly what they can do gets harder.
Our read
The AGI framing around the launch — Greg Brockman spoke of entering the AGI era — is a marketing frame, not a verifiable finding. The practical part is more interesting: once a model reliably operates a computer, the threat picture shifts in both directions. The same capability that lets an attacker find holes helps defenders close them. That is exactly what OpenAI is betting on with the Daybreak program.
For marketing and content teams, little changes in the short term. Astra is expensive, deliberately trimmed in its open version and so far independently unassessed. Third-party analyses put pricing at roughly $10 per million input and $50 per million output tokens — well above Sol. If you run copy, research or editorial workflows today, the existing models still give you the better ratio.
What this means for boostN
As soon as GPT-6 Astra is generally available through the API, we will offer it in boostN as a selectable model — as we do with every new frontier model. Until then the existing GPT and Claude models remain available unchanged, and we will update this piece once solid independent testing exists.
What you can do now
Wait rather than migrate. Without independent testing there is no basis for moving production workflows to Astra. Reliable comparisons typically appear two to four weeks after broad availability.
If you own IT security: the Critical rating is a signal independent of this specific model. Capabilities like this will not stay with one vendor. Current patch levels and working detection are now the more relevant preparation than any model choice.
If you are planning computer-use agents: watch, but do not hard-wire anything. Model selection belongs in one central, configurable place — this market moves too fast for fixed bindings.