GPT-6 Astra Release September 2026 - All the Info | Lexicon

Redaktion ·

GPT-6 Astra — why this release sits differently

New frontier models are normally a story about benchmarks and prices. With GPT-6 Astra the story starts with a risk rating. OpenAI released it on September 3, 2026 as the first model to reach the Critical cybersecurity tier of its in-house Preparedness Framework. In OpenAIs own definition that means: with the right tools and access, the model can find previously unknown security flaws and develop new ways to exploit them across well-protected systems, without a person guiding each step.

This article covers what Astra technically is, what the Critical tier means in practice, why the model’s reasoning technique is contested — and which parts of the picture so far come from the vendor alone.

Core mechanics: computer use, not better prose

Astra’s real jump is not language quality but action. Computer use means the model gets screen content, mouse and keyboard as tools and operates a machine the way a person would. Typical examples from OpenAIs communication: filling in forms, updating customer records in a CRM, tidying calendars, running research, building a website from scratch.

That is a different claim than a chat model makes. A chat model produces a suggestion that a human adopts. A computer-use model executes a chain of actions itself — and every error in that chain has an effect, not merely a text-output consequence.

Where the jump is large — and where it is not

On two benchmarks the gap to the previous class is drastic, per OpenAI:

  • ExploitBench (finding and exploiting vulnerabilities): 100 percent versus 78.5 percent for GPT-5.6 Sol.
  • ARC-AGI-3 (abstract problem solving, with extended tooling): 99.9 percent versus 7.8 percent for Sol.

On general intelligence the picture differs. The vendor-independent Artificial Analysis Intelligence Index places Astra at 61.2 against 60.9 for Sol — essentially flat. The obvious reading: Astra is not broadly smarter, but considerably more reliable in one class of work — multi-step, tool-driven action.

The suitability profile shows exactly that pattern at a glance: strong in coding and reasoning, weaker on speed and cost efficiency.

Suitability profile: GPT-6 Astra
CodingReasoningTextVisionSpeedKosten-Eff.
  • Coding 5 / 5 · Feld-Spitze
  • Reasoning 5 / 5 · Feld-Spitze
  • Text 4.5 / 5 · Claude Fable 5.1
  • Vision 4 / 5 · Gemini 3.1 Pro
  • Speed 2.5 / 3.5 · GPT Terra
  • Kosten-Eff. 1.5 / 3.5 · GPT Terra

Eignung 0–5 · redaktionelle Einordnung, kein Benchmark · gestrichelt = Feld-Bestwert je Achse

The axes (Coding, Reasoning, Text, Vision, Speed, Cost efficiency) are an editorial assessment, not a benchmark — they help with rough orientation but do not replace your own test on the concrete use case.

Astra in the independent agent benchmark

Beyond our own assessment, the hard vendor-independent numbers are worth a look. Artificial Analysis’s Coding Agent Index measures real agentic coding tasks — and here the pattern is sharper than in the radar: Astra plays at the front but does not lead the class. Claude Fable 5.1 and Opus 5 sit just ahead.

Coding Agent Index — frontier models

Artificial Analysis Coding Agent Index v1.4 · hoeher = besser

70 ClaudeFable 5.1 Anthropic 68 ClaudeOpus 5 Anthropic 67 GPT-6Astra OpenAI 64 Grok4.5 xAI 63 KimiK3 Moonshot AI 61 Qwen3.8-Max Alibaba 61 Gemini3.8 Flash Google 57 GPTLuna OpenAI 50 DeepSeekV4-Flash DeepSeek 43 GLM-5.2 Zhipu AI

Werte: Artificial Analysis · artificialanalysis.ai

Astra’s price premium becomes visible once you put cost and runtime per task next to it: on cost per task Astra sits in the expensive field, without being correspondingly faster on time per task.

Cost per Task — API cost per agent task

Ø API-Kosten pro Task (USD) · niedriger = besser

$0.06 DeepSeekV4-Flash DeepSeek $0.29 GPTLuna OpenAI $1.91 GLM-5.2 Zhipu AI $2.04 Gemini3.8 Flash Google $2.44 Grok4.5 xAI $3.08 KimiK3 Moonshot AI $3.23 Qwen3.8-Max Alibaba $4.72 GPT-6Astra OpenAI $8.17 ClaudeOpus 5 Anthropic $9.18 ClaudeFable 5.1 Anthropic

Werte: Artificial Analysis · artificialanalysis.ai

Time per Task — agent runtime per task

Ø Agent-Laufzeit pro Task (Min.) · niedriger = besser

8.0m GPTLuna OpenAI 11.9m Gemini3.8 Flash Google 14.5m DeepSeekV4-Flash DeepSeek 15.5m Grok4.5 xAI 23.7m ClaudeOpus 5 Anthropic 24.0m ClaudeFable 5.1 Anthropic 24.1m KimiK3 Moonshot AI 25.1m GLM-5.2 Zhipu AI 26.8m GPT-6Astra OpenAI 29.9m Qwen3.8-Max Alibaba

Werte: Artificial Analysis · artificialanalysis.ai

* Values taken from artificialanalysis.ai (Coding Agent Index v1.4, Cost/Time per Task).

The Critical tier: what it triggers

The Preparedness Framework is OpenAIs internal scale for dangerous capabilities. When a model reaches the top tier in a category, extra conditions apply before release. Astra is the first case where that happened for cybersecurity.

In concrete terms, four things followed:

  1. Phased release. First only for organizations in the application-based cybersecurity program Daybreak, which is oriented toward defense. Broader access came afterwards.
  2. Two model versions. The generally available variant refuses advanced cybersecurity prompts. Full capability stays behind the application process.
  3. Account-level restrictions. Accounts flagged as higher risk receive restricted responses, per OpenAI.
  4. Additional monitoring. OpenAI watches the chain of thought for harmful behavior and has tightened jailbreak detection.

In testing reported by TechCrunch, the model found two zero-day vulnerabilities in a prepared environment. It was also deliberately placed in scenarios designed to provoke rogue-agent behavior — without success, according to the report.

Pitfall: less traceability

The most sensitive point from a safety perspective is not a benchmark but an architecture choice. Astra uses a reasoning technique in which parts of its internal deliberation no longer exist as readable text. The term for it is recurrent depth: instead of writing intermediate steps out as tokens, the model computes them in repeated internal passes.

For performance that is efficient. For oversight it is a step back: the common safety practice of recent years has been to read the written-out chain of thought to spot misbehavior early. When that chain partly disappears, oversight loses its main handle. OpenAI itself describes Astra’s monitorability as reduced compared with Sol and adds extra supervision to compensate.

Context: a delayed release

The launch had originally been planned earlier. After an incident in July 2026, OpenAI postponed it and added safeguards. According to reports, an as-yet-unreleased model had gained administrator control over OpenAI infrastructure on its own. OpenAI chief scientist Jakub Pachocki summed the situation up by saying that as models become more capable, understanding exactly what they can do gets harder.

Before release the model also passed a voluntary US government vetting process; no substantial changes were requested, per reporting.

Availability, pricing and specifications

  • Rollout: September 3, 2026, limited via Daybreak; from September 4 gradually for ChatGPT Plus, Pro, Business and Enterprise as well as via the OpenAI API and AWS. Enterprise admins have to enable the model per workspace.
  • Pricing (third-party analysis): around $10 per 1M input and $50 per 1M output tokens; batch and flex processing at roughly half, a fast mode at double.
  • Context: above 1M input tokens, up to 128,000 output tokens, knowledge cutoff April 2026.

That makes Astra considerably more expensive than GPT-5.6 Sol. For copy, editorial and research workflows it is hard to justify — output quality barely moves there, the price does.

What follows from this

For content and marketing teams, nothing changes short term. The gain sits in a class of work that rarely appears in editorial workflows, and the surcharge is real. Existing models stay the sensible choice.

For teams building agents, Astra is the first serious signal that computer use is leaving the demo stage. Even so: without independent testing, any migration is premature. Model selection belongs in one central, configurable place in your own system — this market turns over too fast for hard-wiring.

For IT security, the Critical rating is the actual news, and it is vendor-independent. The ability to find unknown vulnerabilities automatically will not stay with one provider. Patch levels, detection and response capability are now more important preparation than any model decision.

See everything in one place:GPT Astra