GPT-6.1 Sol review: 3 tests vs Opus 5.5 and Grok 4.7
I'll get the awkward part out of the way first. I pay for Claude Max, and I pay for ChatGPT's $20 plan. That ratio tells you roughly where I spend my days. If you read this blog regularly, you know I mostly run Claude, with Grok and DeepSeek for the cheap everyday stuff. I do like Codex, though, and I thought the GPT-6 models were great. Still, every time I try a new GPT model on a real project, I'm back in Claude a week later.
GPT-6.1 Sol came closer than any of them to breaking that habit. On hard coding it's within touching distance of Opus 5.5, at half the API price. It's also slow, it needs more hand-holding in the prompt, and in my web design test it was plain lazy. Opus 5.5 stays my default, but not by much.
Here's what I ran: the same web design brief from my Claude Sonnet 5.5 review, plus two hard algorithmic tests in Go. Every test pitted Sol against Opus 5.5 and Grok 4.7.
What is GPT-6.1 Sol?
GPT-6.1 Sol is OpenAI's mid-tier GPT-6 model, released on 29 September 2026 at DevDay. It sits between the flagship GPT-6 Astra and the budget GPT-6 Luna, and it replaced GPT-6 Sol after exactly seven days. If you're struggling to keep up with OpenAI's release pace, you're not alone.
| GPT-6.1 Sol | |
|---|---|
| API model ID | gpt-6.1-sol |
| Context window | 1,050,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | 30 April 2026 |
| API price (per 1M tokens) | $2 input, $0.10 cached input, $10 output |
| Reasoning effort | low, medium (default), high, xhigh, max |
| Available in | API, Codex, ChatGPT Work (Plus and up), GitHub Copilot, Microsoft Foundry |
Note the "ChatGPT Work" bit. At launch, Sol isn't in regular ChatGPT chat at all, and free users don't get it.
OpenAI's pitch is near-Astra performance at a fifth of the token price. On OpenAI's own numbers, Sol scores 75.2% on DeepSWE v1.1 at high effort, against 74.1% for Astra at xhigh. Cached input dropped to half of GPT-6 Sol's price, and responses with factual errors at low effort fell from 11.4% to 7.7%. One odd detail: the DeepSWE score drops to 71.9% at xhigh and max. So turning the effort dial all the way up made it worse at coding, which is not what the dial is for.
Test 1: the web design brief
I gave Sol the exact brief from my Sonnet 5.5 review, so you can compare it with every model I've run through it so far. All three results are side by side on the web design comparison page.
I've always liked GPT's taste in design. Its pages usually look clean and modern without much fuss, and Sol's page is nice too. The problem is how little it did. The design is too simple, and there are no animations at all. It felt like Sol read the brief, did the minimum that technically satisfies it, and went to lunch.
Opus 5.5 was much better. It did more with the same brief, and the result is the one I'd actually show a client. Grok 4.7 landed in the middle: okay, nothing to complain about, nothing to remember either.
What surprised me is that this isn't a capability problem. Sol is #4 on the LMArena WebDev leaderboard, so it can clearly build good frontends. My guess is it wanted a brief that said "add motion, add more sections, go further" in so many words. That turned out to be a pattern, and I'll come back to it below.
Tests 2 and 3: hard algorithmic work in Go
This is where Sol earned its spot. Both tests are logic-heavy Go tasks where a plausible-looking answer is easy to write and a correct one isn't.
The expression-language interpreter
This one is nasty on purpose. Each model had to build an interpreter for a small expression language in Go, tests first, standard library only. The prompt contradicts itself in three places. The grammar allows exactly one argument per call but demands an "arity" error. It never allows let inside an expression but asks for nested let rec. And it claims division never overflows, which MinInt64 / -1 would like a word about. On top of that, the prompt says: if a case is unspecified, return an error. Do not guess.
So the real test was how each model handles a spec that argues with itself. All three spotted all three contradictions, which is impressive on its own. Then they chose differently.
Opus kept to the grammar and returned errors wherever the prompt was silent. Sol and Grok both satisfied the prose by inventing new syntax for let expressions, and they invented two incompatible versions. Sol's version needs a special parsing rule, and that rule sets a trap: let f = fn x -> x + 1 f(let y = 2 y) returns 2 instead of 3, because it quietly gets parsed as two statements. That is exactly the guessing the prompt banned.
On the language the prompt actually defines, all three are correct. They pass all 206 hand-written conformance cases, and across 400,000 randomly generated programs no two implementations ever returned different values. The differences are at the edges:
| Opus 5.5 | GPT-6.1 Sol | Grok 4.7 | |
|---|---|---|---|
| Implementation size | 625 lines | 542 lines | 990 lines |
fib(27) runtime |
45 ms | 31 ms | 58 ms |
| Test functions | 49 | 3 (119 subtests) | 64 |
| Injected bugs caught by own tests | 95.8% | 91.1% | 95.9% |
| Unbounded recursion | returns an error | crashes the process | crashes the process |
| Stays within the stated grammar | yes | no | no |
| Tests written before implementation | yes | yes | test file edited afterwards |
Sol wrote the most compact code and the fastest interpreter, and it followed the tests-first rule properly. Then it wrote the thinnest test suite by a wide margin. In the mutation test (injecting small bugs one at a time and checking whether the model's own tests catch them), Sol's suite missed < flipped to <=, because it never compares equal values with a strict operator. It also missed most of the overflow boundaries.
The robustness gap is the one I'd worry about in production. This 33-character program, let rec f = fn n -> f(n + 1) f(0), kills any process that embeds Sol's (or Grok's) interpreter with a fatal stack overflow. Opus returns a clean error in 0.3 seconds. To be fair, Opus pays for its depth limits elsewhere: it wrongly rejects a few extreme but valid programs, like recursion 200,000 levels deep.
The Collatz worker
The second test looks easier and isn't. Each model had to write a Collatz (3n+1) checker in Go: a Trajectory function that follows one starting number down to 1, and a parallel MaxTotalStoppingTime search over a range. Same rules as before: tests first, public API only, standard library only, and no weakening tests to make the code pass.
The maths is trivial. Even numbers halve, odd numbers become 3n+1, repeat until you hit 1. The traps are everywhere else:
- 3n+1 has to stay exact above
math.MaxInt64, so everything runs onmath/big, with all the pointer-aliasing fun that brings in Go. - The parallel search has to return the same answer with one worker or sixteen, and ties go to the smaller start. Get the merge wrong and the result depends on goroutine scheduling.
- A step cap has to stop a long trajectory without flagging it as a cycle, and a cap of 0 must not count as reaching 1.
- The tests have to lock known values, like 27 taking 111 steps and peaking at 9,232, or 97 having the longest run between 1 and 100 at 118 steps.
The sneaky part is the Cycle flag. A repeat before reaching 1 would be a counterexample to the Collatz conjecture, and nobody has ever found one. So no valid input can reach that code path through the public API, and an honest test suite has to say so instead of pretending to cover it. The prompt also states the conjecture is unsolved and asks for a final line confirming no counterexample turned up. Part of the test is simply whether the model overclaims.
All three got the maths right. Checked against an independent reference on about 12,000 trajectories and 17,000 range searches, with starts up to 1,000 bits long, none of them returned a wrong step count, peak or search result. None claimed a proof or a counterexample either, so nobody failed the honesty check.
| Opus 5.5 | GPT-6.1 Sol | Grok 4.7 | |
|---|---|---|---|
| Own tests (all passing) | 42 | 39 | 31 |
| Seeded bugs its tests catch (of 36) | 36 | 30 | 35 |
| Search 1 to 10,000,000, 8 workers | 0.2 s | 100 s | 92 s |
| Implementation size | 300 lines | 163 lines | 218 lines |
| Sticks to the spec | no, adds an exported Step |
yes | no, capped StoppingTime is -1 |
This is the test where Sol's literal streak paid off. It was the only model that did exactly what the prompt asked and nothing else: the listed API, the specified behaviour, and the smallest code of the three. The prompt also wants tests for "the full sequence of 6", but Result only holds summary numbers, so no public call returns a sequence. Sol solved that neatly by running Trajectory from every suffix start and every capped prefix, with a comment explaining why.
The other two each bent the spec. Opus added an exported Step function the prompt never listed, and four of its tests call it. Grok returns -1 for StoppingTime when the step cap hits early, where the prompt says "steps taken". Grok's reading is arguably the more useful design, but it isn't what the spec says.
So why isn't Sol first? Tests, again. Its suite checks 39 trajectories and 13 ranges, against more than 20,000 trajectories and 854 ranges for Opus. Of 36 seeded bugs, its tests let six through, so they'd happily pass a search that treats hi as exclusive, or one that silently gives up on long trajectories. Its widest search range is 1 to 100.
Then there's speed. Sol and Grok detect cycles by storing every term as a decimal string in a map. That's simple and obviously correct, and it allocates 12.5 GB to search the first million numbers. Opus works in machine words with Brent's cycle detection, which stores nothing, and it searched 1 to 10 million in 0.2 seconds against 100 for Sol. The prompt never asked for speed, but it did warn that the range may be large.
Scoreboard
| Test | 1st | 2nd | 3rd |
|---|---|---|---|
| Web design | Opus 5.5 | Grok 4.7 | GPT-6.1 Sol |
| Interpreter | Opus 5.5 | Grok 4.7 (narrowly) | GPT-6.1 Sol |
| Collatz worker | Opus 5.5 | GPT-6.1 Sol | Grok 4.7 |
| Hard logic overall | Opus 5.5 | GPT-6.1 Sol | Grok 4.7 |
Opus wins outright. On the hard logic, Sol is a close second, and Opus is really only a little better. That "little" is mostly judgement: knowing when to say no to a spec instead of improvising.
Speed and prompting: where Sol annoys me
It's slow
Sol is slow, and it isn't just me being impatient. Artificial Analysis measures it at 51 tokens per second at max effort, which puts it 147th of 224 models on speed. Their own word for it is "notably slow". One developer measured 27.4 tok/s in Codex against roughly 90 for Opus 5.5, and flagged the method as rough because Claude did the measuring. I respect the honesty. Launch week also brought "model is at capacity" errors in ChatGPT and Codex.
OpenAI's answer is Ultrafast, a premium tier with up to 8x faster generation in Codex. It's live for Astra. For Sol, OpenAI still says it's coming "in the coming days", and in ChatGPT you'll only get it on the new $500-a-month Pro plan. On my $20 plan, slow is what I get.
You have to spell everything out
This is the difference I notice most day to day. With Opus, I describe what I want and it fills in the obvious gaps. With Sol, I have to say exactly what I want, including the parts I thought were obvious.
Looking back at my tests, it's the same pattern each time. The design had no animations because the brief didn't say "animate things". The test suite was the smallest one that technically covered the requirements. In the interpreter, Sol followed the prose literally even where the grammar said otherwise. It does what you write, which isn't always what you mean.
That isn't all bad. If you're feeding a model precise specs in an automated pipeline, doing exactly what it's told is a feature. For exploratory work, where I want the model to be a bit clever and occasionally push back, it gets tiring fast. Claude also feels more consistent to me from one run to the next. I can't put a number on that one. It's vibes from a lot of daily use, but they're strong vibes.
GPT-6.1 Sol benchmarks and pricing vs Opus 5.5 and Grok 4.7
The independent numbers line up with my tests. Opus is ahead, Sol sits close behind at half the API price, and Grok trails on capability but has the cheapest output.
| Model | Artificial Analysis Intelligence Index | API price, input / output per 1M tokens |
|---|---|---|
| Claude Opus 5.5 | 58 | $4 / $20 |
| Claude Sonnet 5.5 | 56 | $2 / $10 |
| GPT-6.1 Sol (max effort) | 52 | $2 / $10 |
| Grok 4.7 | 46 | $2 / $6 (xAI) |
On coding specifically, Mercor's independent DeepSWE v1.1 run has Sol and Opus 5.5 tied at 72.3%. On reasoning, ARC Prize found similar results: 94.2% on ARC-AGI-2 at $0.25 per task, close to Astra's 95.0%. So the "near the top for a fraction of the cost" story holds up outside OpenAI's own charts, which isn't always the case with launch claims.
There's one awkward row for OpenAI in that table. Sonnet 5.5 costs exactly the same as Sol and scores higher. If price is your main reason to look at Sol, run Sonnet on your workload too.
On the subscription side, my $20 Plus plan gets 15 to 160 local Codex messages per five hours with Sol, going by OpenAI's Codex pricing page. Pro 200 subscribers lose half their allowance from 30 October, which is a fun thing to announce on the same day as a $500 plan.
If you're moving an API integration over from gpt-6-sol, check these first (OpenAI's GPT-6 guide has the details):
noneandminimalreasoning effort are gone, solowis now the floor.- Tool calling only works through the Responses API. Chat Completions still works, without tools.
- Prompts over 272K tokens bill the whole request at 2x input and 1.5x output, not just the tokens past the line.
- Sol is rated Critical for cybersecurity, and the filters show. It ranks 42nd of 44 on Vals' CyberBench with 39.29%, where GPT-6 Sol scored 78.0%. Expect refusals on security work unless you're in OpenAI's Daybreak program.
Verdict: who should use GPT-6.1 Sol?
GPT-6.1 Sol makes sense if:
- you run agentic coding at volume and cost per task matters more than the last few points of quality
- you write precise specs and want a model that follows them to the letter
- you already live in Codex or GitHub Copilot
Stick with Opus 5.5 if:
- you want the model to work out your intent and fill the gaps
- you build frontends and care about polish
- consistency across runs matters more to you than the API bill
And skip Sol for security tooling for now, unless you're in Daybreak.
For me, Opus 5.5 stays my favourite. It's more confident, sometimes cleverer, and its results are more reliable. When I pick models for client agent builds, that reliability is worth more than the price difference. But Sol is the closest an OpenAI model has come to changing my mind. If Ultrafast lands on the $20 plan, or the speed improves some other way, I'll run these tests again.
Related Articles
I Tested GPT 5.4 Against Every Rival — Here's My Honest Review
I tested GPT 5.4 head-to-head against Claude, Gemini, and MiniMax on a real coding task. Here's what the benchmarks don't tell you.
10 min read
AI NewsClaude Sonnet 5.5 Review: I Tested It Against Opus 5.5
Claude Sonnet 5.5 review: I tested it against Opus 5.5 and Sonnet 5 on a web design and a Go chess engine. Real costs, results, and who should use it.
14 min read
AI NewsClaude Fable 5.1 review: I built an image codec with it
I used Claude Fable 5.1 to build an image codec that beats WebP and a full agency-grade site. Here's what changed from Fable 5, and what it cost.
11 min read