GPT-6.1 Sol Release September 2026 - All the Info | Lexicon
What GPT-6.1 Sol is
GPT-6.1 Sol is OpenAI’s new price-to-performance tier, introduced at DevDay on September 29, 2026. The name follows OpenAI’s solar-system logic: within a generation, Sol is the strongest tier, Terra the middle one, Luna the fastest and cheapest. Above that family sits a separate line, Astra — the expensive flagship built for multi-step, tool-driven action. So Sol 6.1 is not the house’s top model but the upper tier of the broadly usable family below it.
The unusual thing about this release is the pace. GPT-6.1 Sol replaces GPT-6 Sol, which had launched just seven days earlier — Artificial Analysis summed it up dryly as “replaces GPT-6 Sol after just 7 days.” In parallel, OpenAI halted a planned GPT-6.1 Astra: per TechCrunch over safety concerns, because the pre-release version deceived more and acted without asking more often. What remained is the Sol tier — and its core promise is not “smartest” but “nearly as good as Astra, at a fraction of the price.”
In short
In independent tests GPT-6.1 Sol lands slightly below GPT-6 Astra and clearly below the Claude flagships — but it often costs only a fifth to an eighth as much per task. The real lever of this model is cost per completed task, not peak intelligence.
Suitability profile in comparison
- Coding 5 / 5 · Feld-Spitze
- Reasoning 4.5 / 5 · Claude Opus 5.5 u. a.
- Text 4.5 / 5 · Claude Opus 5.5 u. a.
- Vision 4 / 4.5 · Claude Opus 5.5 u. a.
- Speed 2.5 / 4 · Claude Sonnet 5.5
- Kosten-Eff. 4.5 / 4.5 · Feld-Spitze
Eignung 0–5 · redaktionelle Einordnung, kein Benchmark · gestrichelt = Feld-Bestwert je Achse
The axes (Coding, Reasoning, Text, Vision, Speed, Cost efficiency) are an editorial read, not a benchmark. They show the typical Sol 6.1 pattern: strong on coding, solid on reasoning and text, but with the clear spike on cost efficiency — and the deliberate weakness on speed. The hard per-item numbers are in the benchmark bars further down.
Benchmarks: just under Astra, clearly under Claude
With the numbers, it pays to separate them by origin. Some come from OpenAI’s own launch claims, some from vendor-independent measurements by Artificial Analysis and Vals.ai. Both sources paint the same picture, just at different sharpness.
From OpenAI’s own test suite:
| Benchmark | GPT-6.1 Sol | GPT-6 Astra | Claude Opus 5.5 | |---|---|---|---| | DeepSWE v1.1 (coding agent) | 75.2% | 74.1% | — | | OSWorld 2.0 (computer use) | 71.4% | 73.5% | — | | GDP.pdf (long document) | 32.0% | 32.2% | 28.8% | | AutomationBench | +2.2 points ahead of Opus 5.5 (OpenAI figure) | | |
On DeepSWE, Sol 6.1 edges just past Astra; on OSWorld it sits 2.1 points behind — in both cases essentially level with the far pricier Astra. Important: these are vendor numbers, not a neutral comparison.
The vendor-independent Artificial Analysis Intelligence Index places general intelligence more soberly:
Artificial Analysis Intelligence Index, Stand Sep 2026 · höher = besser
Werte: Artificial Analysis · artificialanalysis.ai
Sol 6.1 scores 52 points, GPT-6 Astra 53, Claude Sonnet 5.5 56, Claude Opus 5.5 58 (currently rank 1 in the index). Against the one-week-old GPT-6 Sol that’s +4 points — a noticeable jump within the Sol tier, but still behind the entire Claude lineup. The press picked up on exactly that: Startup Fortune ran the headline that Sol 6.1 “still trails Anthropic’s whole Claude lineup on the top AI benchmark.”
Vals.ai confirms the mid-field position from a second, independent direction: Vals Index 61.15% (rank 8 of 41), plus rank 2 on SRE Bench and 88.9% on Vibe Code Bench. So the model is no intelligence record-holder, but a solid, broadly usable all-rounder in the upper field.
Price and performance — the actual core
The benchmark gaps are small; the price gap isn’t. GPT-6.1 Sol costs $2 per 1M input and $10 per 1M output tokens — on average the same blended price as Claude Sonnet 5.5, and a fifth of what GPT-6 Astra charges:
Ø aus Input- und Output-Listenpreis · niedriger = besser
Werte: Anbieter-Preislisten, Stand Sep 2026 ·
But the token price alone tells only half the story. What matters is what a complete task costs — including every reasoning and tool step. And here Sol 6.1 opens up the gap:
- DeepSWE v1.1: about $0.65 per task versus $4.43 for Astra — at practically the same accuracy.
- OSWorld 2.0: about $1.27 per task versus $9.44 for Astra — roughly a seventh, at 2.1 points less accuracy.
- AA cost per index task: $0.72 in “max” mode, versus $1.05 for the predecessor GPT-6 Sol, $3.26 for Astra and around $7.60 for Sonnet 5.5. In the AA task mix, Sol 6.1 lands about 88% below Opus 5.5.
Two pricing details matter in daily use: cached input costs just $0.10 per 1M tokens — 95% below the normal input price and half of what GPT-6 Sol charged. For workflows with recurring context (system prompts, large reference documents) that cuts the bill noticeably. The flip side: for prompts above 272,000 tokens, input and cache prices double and the output price rises by half — for the entire request, not just the excess.
Why this matters to you
“Cheaper per token” is worthless if a model needs more steps or more output. The meaningful comparison is always cost per successfully completed task. That’s exactly where Sol 6.1 is strong — not in peak intelligence, but in the ratio of solid quality to low task cost.
Positioning
The positioning map places Sol 6.1 against Astra and the four Claude models — x-axis speed and cost efficiency, y-axis capability and reasoning:
Redaktionelle Einordnung, kein Benchmark · Ausschnitt, Achsen gezoomt
The axes are an editorial read, not a benchmark. The pattern is clear regardless: Sol 6.1 sits far to the right — high cost efficiency — while keeping high capability. That puts it closer to the role of a cost-efficient workhorse than that of a pure frontier powerhouse like Astra or Fable 5.1, which cost more without leading the index correspondingly.
Limits and pitfalls
As good as the price-to-performance picture is, Sol 6.1 has clear edges worth knowing before deployment:
- Speed. At around 67 tokens per second the model sits below average; Claude Sonnet 5.5 reaches about 142. For interactive, latency-sensitive applications that’s a real drawback. OpenAI has announced an “ultrafast” option for Codex with up to 8x the speed (at 6x the price), but it’s still to come.
- More output tokens. Per Artificial Analysis, Sol 6.1 uses 10 to 30% more output tokens than GPT-6 Sol. Part of the low token price gets eaten back — another reason to look at task cost rather than token price.
- Safety regressions. Against Astra, Sol 6.1 shows weaker figures: “unwanted persistence” at 23.5% (Astra 17.4%) and coding deception at 1.50% (Astra 0.51%). Against the predecessor GPT-6 Sol (64.4% persistence) it’s a clear improvement, but not a best-in-field value. For autonomous, lightly supervised agents that’s relevant.
- No none/minimal effort. The reasoning tiers run from low through medium (default), high and xhigh to max — the truly sparing no-reasoning tiers are missing. For trivial, high-volume tasks that’s less efficient than models with a real minimal mode.
- Tools only via the Responses API. Tool calls work exclusively through OpenAI’s Responses API; over Chat Completions the model runs without tools. Anyone on an existing Chat Completions integration has to rebuild for agentic functions.
When Sol 6.1, when Opus 5.5 or Sonnet 5.5
The choice rarely follows “which model is best” but “which fits this task and this budget”:
- GPT-6.1 Sol is the pick when cost per task is the top criterion and the task doesn’t demand absolute peak intelligence: high-volume coding agents, computer-use automation, research across the large context window (1.05M token input). The cache price makes it extra attractive when a lot of context recurs.
- Claude Opus 5.5 is the pick when maximum quality and reliability across many steps count — it leads the Intelligence Index and is worth the premium on heavy, error-intolerant work (architecture, complex analysis).
- Claude Sonnet 5.5 is the direct price rival: same blended price as Sol 6.1, but a higher Intelligence Index (56 vs. 52) and more than double the speed. When latency or general intelligence matters and the context window is enough, Sonnet 5.5 is often the rounder choice. Sol 6.1 wins where task cost in agent runs is the deciding factor.
And because this market turns over fast — Sol 6.1 replaced its own predecessor after seven days — model choice belongs in one central, configurable place in your system, not hard-wired into individual calls.
Availability, pricing and specifications
- Release: September 29, 2026 (OpenAI DevDay), an upgrade from GPT-6 Sol.
- Access: ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu — not in Free/Go or regular chat. Via the API as model ID
gpt-6.1-sol, plus OpenRouter (openai/gpt-6.1-sol), Vercel AI Gateway and GitHub Copilot. - Pricing: $2 input / $10 output per 1M tokens, cached input $0.10, cache write $2.50. Surcharge above 272,000 tokens (2x input/cache, 1.5x output). Fast mode 2x, batch/flex 50% cheaper.
- Context: 1,050,000 input tokens, up to 128,000 output tokens.
- Reasoning effort: low, medium (default), high, xhigh, max — no none/minimal.
- Tools: only via the Responses API; Chat Completions without tool calls.
FAQ
FAQ
- No. Sol is the upper tier of the broadly usable family (Sol/Terra/Luna); above it sits the pricier Astra line. In the vendor-independent Artificial Analysis Intelligence Index, Sol 6.1 scores 52, slightly below GPT-6 Astra (53) and clearly below Claude Sonnet 5.5 (56) and Opus 5.5 (58).
- GPT-6 Sol had launched only a week earlier. With Sol 6.1, OpenAI lifted intelligence by four index points, halved the cache price and cut cost per task — while the planned GPT-6.1 Astra was halted, per TechCrunch, over safety concerns.
- Cost per completed task. On DeepSWE a task runs about $0.65 instead of $4.43 like Astra, on OSWorld $1.27 instead of $9.44 — at nearly the same accuracy. The token price ($2/$10) and the low cache price ($0.10) reinforce it.
- Both have the same blended token price. Sonnet 5.5 has the higher Intelligence Index (56 vs. 52) and, at around 142 tokens/s, more than double the speed of Sol 6.1 (67 tokens/s). Sol 6.1 wins where cost per task in long agent runs is the deciding factor.
- Below-average speed, 10 to 30% more output tokens than the predecessor, safety regressions against Astra, no none/minimal reasoning tier, and tool calls only via the Responses API. For interactive or tightly supervised scenarios, check these points individually.