GPT-6 Sol vs. Claude Opus 5.5: Is cheaper important when results aren’t consistent?

Sep 28, 2026 494 views
Abstract digital glitch art with neon pink and blue halftone waveforms, symbolizing AI data fragmentation.

OpenAI launched GPT-6 Sol on September 22, claiming it scores higher than Claude Opus 5 on business workflow tasks at 9% of the cost per task. OpenAI’s marketing claims Sol scores 33.2% on Zapier’s AutomationBench, compared with Opus 5’s 26.9%, at a fraction of the per-task cost. Even though OpenAI compared Sol with Opus 5, Anthropic released Opus 5.5 the same day, so I tested against the newer model.

As for pricing, OpenAI’s launch page puts Sol at $2 per million input tokens and $10 per million output tokens, half the price of GPT-5.6 Sol. Opus 5.5 costs $4 and $20. So even if Sol uses more tokens, it could still come in cheaper. 

As these models improve, it’s getting harder to separate them on accuracy alone. Lately I’ve been testing new criteria and running more trials per model to find where they differ. For GPT-6 Sol vs. Opus 5.5, both models came in dead even on standard developer tasks, so I rebuilt the tests to measure consistency. My thinking is that a score means more if a model can repeat it. My hypothesis was correct. This is the first test set in a while where the models gave me different results.

I usually share my prompts, but I can’t for these tests because each test depends on large files.

The tests

My first round covered the kind of work developers hand these models every day. Both fixed three planted bugs in a Python billing service, triaged 40 failed CI jobs, and answered 20 questions about a fake SDK without inventing a single method (not again… eye roll). So then I built harder versions of these tests, like a 3,664-line outage postmortem and a spec implementation graded by 120 hidden tests. Both models scored perfectly on the harder tests, too. It was time to try something different. And that’s when I decided to judge based on consistency.

I called both models through their APIs with identical prompts, at the highest effort setting each offers: max for both. OpenAI’s AutomationBench comparison ran Sol at xhigh against Opus 5 at max. I ran each test five times to see whether the same answers came back consistently.

  • CI triage – The model reads 40 failed CI job logs and a runbook with conditional rules, then decides whether to retry, block, or page for each one. This is the closest match to OpenAI’s AutomationBench claim, though AutomationBench tests agents on end-to-end workflows using 47 tools, while this is a single call.
  • Incident logs – The model reads 3,664 lines of logs from five services during a two-hour outage and answers seven postmortem questions. The answers depend on exact counting, a timezone conversion, and a red herring.
  • Resolver spec – The model writes a dependency resolver for a fictional package manager from a two-page spec, without running any code. A hidden suite of 120 tests grades it.

I logged tokens, cost at list price, and time for every call.

CI triage

Both models went 5 for 5, with all 40 calls right every time. Opus 5.5 averaged 1 minute 27 seconds, 11,127 output tokens, and $0.24 per run. Sol averaged 18 seconds, 1,143 output tokens, and $0.02.

On CI triage, the test closest to OpenAI’s claim, Sol cost 8% of what Opus 5.5 did. That’s right in line with the 9% OpenAI advertises, even though Opus 5.5 is cheaper per token than Opus 5. The reason is that Opus 5.5 used almost four times as many output tokens as Opus 5 did when I ran the same test.

Incident logs

Across five runs, Opus 5.5 was perfect every time. Sol was only perfect twice. Two of its runs missed the same customer, whose original charge was confirmed 52 seconds after the retry had already succeeded. Another run counted 27 failed checkouts instead of 28.

Sol beat Opus 5.5 on price, speed, and token usage… but does that matter if it’s not as accurate? Opus 5.5 averaged 6 minutes 44 seconds, 53,308 output tokens, and $1.68 per run. Sol averaged 1 minute 41 seconds, 7,073 output tokens, and $0.30. Sol also read the same log file in 25% fewer input tokens, 113,966 against 152,345.

Resolver spec

Similar to incident logs, Opus 5.5 passed every test every time. Sol did better than it did with incident logs, with four out of five runs passing every test. On the fifth, it left a stray closing parenthesis on line 78, so the module crashed on import and failed all 120 tests.

Sol won on cost and speed here too. Opus 5.5 averaged 9 minutes 40 seconds, 70,687 output tokens, and $1.42 per run. Sol averaged 6 minutes 28 seconds, 21,435 output tokens, and $0.22.

Results

Test (5 runs each)Opus 5.5GPT-6 Sol
CI triage5/5 perfect, 1:27, 5,447 in / 11,127 out, $0.245/5 perfect, 0:18, 3,867 in / 1,143 out, $0.02
Incident logs5/5 perfect, 6:44, 152,345 in / 53,308 out, $1.682/5 perfect, 1:41, 113,966 in / 7,073 out, $0.30
Resolver spec5/5 perfect, 9:40, 2,430 in / 70,687 out, $1.424/5 perfect, 6:28, 1,687 in / 21,435 out, $0.22
Perfect runs15 of 1512 of 15
Total cost, 15 runs$16.72$2.68

Opus 5.5 scored perfectly on all 15 runs across CI triage, incident logs, and the resolver spec. GPT-6 Sol scored perfectly on 12 of 15, missing runs on the two hardest tests. Sol was faster and cheaper on every test, but Opus 5.5 was consistently more accurate. Sol’s 15 runs cost $2.68 in total, about 16% of Opus 5.5’s $16.72.

What do I think?

OpenAI says Sol beats Opus 5 on business workflows. It didn’t beat Opus 5.5 in my tests. On a single attempt, the two models tied on every task, and across repeat runs, Opus 5.5 came out ahead. It was perfect on all 15 runs, while Sol missed three.

Sol wins on price and speed. It cost about a sixth of what Opus 5.5 did overall, and 8% on CI triage, the test closest to OpenAI’s claim. It was also faster on every test, up to nearly five times faster on CI triage. But Sol’s misses were the kind that slip past review, like a customer left off a refund list or a file that won’t import.

Use Sol for high-volume work where a person or a test suite checks the output, such as triage, drafting, and code that runs through CI. At those prices, you can afford to run it twice and compare. Use Opus 5.5 when a wrong answer is expensive and nobody is checking, such as financial reconciliation or incident reports that go straight to customers.

The post GPT-6 Sol vs. Claude Opus 5.5: Is cheaper important when results aren’t consistent? appeared first on The New Stack.

Comments

Sign in to comment.
No comments yet. Be the first to comment.

Related Articles

GPT-6 Sol vs. Claude Opus 5.5: Is cheaper important when ...