GPT-6.1 Sol: Faster, Smarter, More Reliable — My First Impression
OpenAI unveiled GPT-6.1 Sol at DevDay on September 29 — as the replacement for GPT-6 Sol, which had shipped just seven days earlier. I wired the new model into boostN right after release and ran my first real tasks through it. Short version up front: it works reliably and well. Everything else I’m still checking.
Status: first impressions, test results to follow
I’m in the middle of testing. This article captures my early hands-on take — the targeted comparisons against the predecessor GPT-6 Sol and against Opus 5.5 and Sonnet 5.5 are still running. I’ll add the hard detail numbers once I’ve measured the model against concrete tasks.
How I got started
My approach was deliberately plain: I wired Sol 6.1 into boostN and ran the same kind of tasks I hand out anyway. No special prompting, no tuned settings. I wanted to see first whether the model holds up in everyday use — before pitting it against the competition.
It does: the first tasks ran reliably and delivered good results. For a model that replaces its own predecessor after a week, that’s a solid first impression. Whether it actually catches the stronger Claude models on the hard cases is a different question — and that’s exactly what I’m testing now.
Where the model sits
Before the details, the rough map. The chart below places GPT-6.1 Sol against the relevant comparison group — capability versus cost-efficiency. Both axes are an editorial assessment, not a benchmark:
Redaktionelle Einordnung, kein Benchmark · Ausschnitt, Achsen gezoomt
Sol 6.1 lands far to the right: it’s one of the most cost-efficient models in the field without dropping off sharply on capability. That’s the story of this release — and that’s where the closer look pays off.
Intelligence: just behind Astra, behind the Claude top
On the Artificial Analysis Intelligence Index, the overall measure across many tasks, Sol 6.1 sits just behind its big sibling Astra and clearly behind the current Claude top:
Artificial Analysis Intelligence Index, Stand Sep 2026 · höher = besser
Werte: Artificial Analysis · artificialanalysis.ai
52 points against 53 for Astra, 56 for Sonnet 5.5 and 58 for Opus 5.5. That’s progress over GPT-6 Sol (up four points), but it doesn’t reach the Claude line. The press put it soberly: Sol 6.1 still trails Anthropic’s whole Claude lineup on the top benchmark.
Price: the real argument
On price, the picture flips. Sol 6.1 costs 2 US dollars per million input tokens and 10 dollars per million output tokens — on par with Sonnet 5.5 and a fifth of what Astra or Fable 5.1 charge:
Ø aus Input- und Output-Listenpreis · niedriger = besser
Werte: Anbieter-Preislisten, Stand Sep 2026 ·
Cache reads even drop to 0.10 dollars, half of what the predecessor charged. For a frontier-adjacent provider, that’s an aggressive price.
What convinces me
The real lever is price-performance, not raw intelligence. Three points stand out:
Coding at Astra level, at a fraction of the cost. On DeepSWE, Sol 6.1 even edges past Astra with 75.2 percent (Astra: 74.1) — at roughly a fifth of the cost per task (0.65 versus 4.43 dollars). On OSWorld it lands 2.1 points behind Astra but costs only about a seventh per task. For agentic coding, that’s a strong starting position.
Cost per task, not per token. The number that matters in production is the price per completed task. In the AA task mix, Sol 6.1 costs about 0.72 dollars against 3.26 for Astra — roughly 88 percent below Opus 5.5. When a model solves a task reliably and cheaply, the slightly lower overall index often matters less in daily work.
Fewer hallucinations than the predecessor. The error rate at low effort dropped from 11.4 to 7.7 percent. That matches my early impression: the first tasks came back without the little fabrications you’d otherwise have to double-check.
Where I’m skeptical
Just as honestly, the flip side:
- The overall index trails Claude. For the genuinely hard, multi-layered tasks, Opus 5.5 with 58 points is still the other league. Price-performance is strong, absolute peak performance isn’t.
- Speed is below average. Around 67 tokens per second — Sonnet 5.5 delivers more than double that. You feel it in interactive work. A faster “Ultrafast” option is announced for Codex, but it costs six times as much.
- The 272K surcharge. Above 272,000 tokens per request, the input price doubles and output rises by half — for the entire request. On very long contexts, the price advantage melts.
- Safety regressions versus Astra. “Unwanted persistence” sits at 23.5 percent (Astra: 17.4), the coding deception rate at 1.50 percent (Astra: 0.51). Both are better than GPT-6 Sol, but a step back from Astra — for autonomous runs without close oversight, something I’m keeping an eye on.
What I’d use it for
This is a preliminary take — the hard comparisons are still running. With that caveat:
For high-volume, clearly scoped coding and agent tasks where cost per task matters, Sol 6.1 looks like a strong candidate on first impression. That’s exactly where it plays its price-performance advantage.
For the heavy cases — tricky architecture, long chains with many intermediate decisions, anything where a mistake gets expensive — I’m staying with Opus 5.5 for now. And where I need speed at decent quality, Sonnet 5.5 with its double throughput is still my reflex. I’ll put Sol 6.1 where volume and cost decide, not where the last bit of reasoning depth does.
Interim verdict
After the first tasks, my impression is clearly positive but incomplete. Sol 6.1 works reliably, is surprisingly cheap, and matches the far pricier Astra on coding. Whether it reaches the Claude models on the hard tasks, I’ll only know after the targeted tests — which I’ll add.