Anthropic's Sonnet 5.5 Outperforms Opus on Most Tasks Despite Lower Per-Token Cost
In head-to-head testing, Claude Sonnet 5.5 delivered perfect results across three complex coding challenges while costing 42% less than Opus 5.5, though the cheaper pricing model does not always translate to proportional savings.

Anthropic released Claude Sonnet 5.5 just six days after unveiling Opus 5.5, positioning the newer model as a more economical alternative in its frontier-class lineup. The company reports that Sonnet 5.5 achieves a 70.6% score on Terminal-Bench 4.0, surpassing Opus 5.5's 66.4% at xhigh effort levels. Additionally, Sonnet 5.5 generates output more than 30% faster than its predecessor, Sonnet 5, while consuming fewer tokens per task.
Pricing represents the most obvious distinction between the two models. Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, exactly half the rate of Opus 5.5's $4 and $20. However, lower per-token pricing does not automatically translate to lower overall costs. When a model requires additional tokens to complete a task, the savings diminish accordingly. Independent analysis from Artificial Analysis discovered this dynamic at maximum effort settings, with Sonnet 5.5 costing $7.67 per task compared to Opus 5.5's $5.98.
Half the price per token doesn't guarantee half the bill, though.
Testing methodology
Both models were invoked through the Anthropic API using identical prompts, adaptive thinking enabled, and maximum effort settings. Each test ran five times per model to assess consistency, with every execution graded against a hidden test suite the models had never encountered. Measurements included token usage, cost at list pricing, and execution time for each run.
- Agentic bug fix: Models received a small Python order-pricing repository containing four intentional bugs and one flaky test, along with tools to read files, write files, and execute tests. A hidden suite of 12 tests validated the fixes.
- Resolver spec: Models wrote a dependency resolver for a fictional package manager based on a two-page specification without executing any code. A hidden suite of 120 tests evaluated the implementation.
- Concurrency bugs: Models addressed an asyncio job queue containing three race conditions and received an incident report describing double charges and unexecuted jobs. Eight hidden tests verified each fix.
Results by task
Agentic bug fix
Both models successfully fixed all four bugs across all five runs and passed all 12 hidden tests without modifying test files. Both identified the flaky test correctly. Sonnet 5.5 averaged 5 minutes 8 seconds, 29 tool calls, 42,608 output tokens, and $0.70 per run. Opus 5.5 averaged 3 minutes 21 seconds, 25 tool calls, 20,625 output tokens, and $0.75.
Sonnet 5.5 encountered a significant obstacle during initial testing. The model operates in steps, each constrained by a 32,000-token limit. On the first attempt, Sonnet 5.5 deliberated so extensively within a single step that it exceeded this limit in four of five runs, halting before completion. Opus 5.5, operating under the same constraint, never approached the threshold. Raising the limit to 128,000 tokens and rerunning the four failed attempts allowed all to pass. The failed attempts incurred approximately $1.40 in costs, not reflected in the baseline totals. Including these redone runs, Sonnet 5.5 averaged roughly $0.98 per run, exceeding Opus 5.5's $0.75.
Sonnet 5.5 thought so long in a single step that it hit that limit in four of five runs…
Opus 5.5 won this test. Although both models fixed every bug, Opus 5.5 completed the task approximately 35% faster and, accounting for Sonnet 5.5's failed runs, cost less per execution.
Resolver spec
Sonnet 5.5 prevailed on this task. Both models passed all 120 hidden tests across all five runs, but Sonnet 5.5 achieved this faster and at lower cost. Sonnet 5.5 averaged 8 minutes 58 seconds, 81,097 output tokens, and $0.82 per run. Opus 5.5 averaged 9 minutes 40 seconds, 70,687 output tokens, and $1.42. Despite consuming 15% more tokens, Sonnet 5.5 cost 42% less and finished slightly quicker.
Concurrency bugs
This test produced the most striking results. Sonnet 5.5 resolved all three race conditions and passed all eight hidden tests on every run. Opus 5.5 matched this performance on three runs but exhausted its 128,000-token output limit on the other two without producing a solution.
Sonnet 5.5 averaged 12 minutes 13 seconds, 101,788 output tokens, and $1.02 per run. Opus 5.5 averaged 16 minutes 43 seconds, 111,428 output tokens, and $2.24 per run, including failures. Sonnet 5.5 dominated this test decisively, delivering flawless results every time while running faster and costing less than half as much.
Overall performance
Across all 15 runs, Sonnet 5.5 achieved perfect results on every execution, while Opus 5.5 succeeded on 13, with both failures occurring on the concurrency test. Sonnet 5.5 accumulated a total cost of $12.69 compared to Opus 5.5's $22.07, representing a 42% savings. When including the four redone runs, Sonnet 5.5's cost rose to $14.09, still approximately 36% less than Opus 5.5. Across all 15 runs, Sonnet 5.5 also completed tasks 11% faster overall. Sonnet 5.5 won the resolver and concurrency tests, while Opus 5.5 prevailed on the agentic test, finishing roughly 35% faster despite Sonnet 5.5's doubled token consumption.
Analysis
Artificial Analysis reported that Sonnet 5.5 costs more per task than Opus 5.5 at maximum effort while scoring marginally lower on its Intelligence Index. This pattern did not emerge in these tests. Sonnet 5.5 demonstrated superior accuracy, and even accounting for the four redone runs, it cost approximately 36% less overall than Opus 5.5.
The per-token discount did not translate directly to proportional bill reductions. Sonnet 5.5 frequently required additional tokens to complete tasks, resulting in approximately 36% lower costs overall rather than the theoretical 50% reduction.
For complex coding tasks, Sonnet 5.5 should serve as the default choice. It delivered flawless performance on every run once its output limit was increased, making it advisable to set this parameter high since the model requires more deliberation per step than Opus 5.5. For agent loops, Opus 5.5 remains the better option. It completed the agentic test approximately 35% faster and, once accounting for Sonnet 5.5's failed runs, incurred lower costs.