3-Line TL;DR
- Sonnet 5.5 costs half of Opus 5.5 per token, but Artificial Analysis measured the heaviest token use it has seen, so at Max a task costs $7.60 against Opus's $5.98
- Matched by score instead of effort label, Opus can win, since Opus 5.5 at Xhigh scores 56 like Sonnet at Max for $3.46, while Artificial Analysis calls Sonnet's High (47 for $1.08) its most competitive setting
- I would still start Sonnet at High and measure cost per accepted task on my own work, because a per-token price says little about a per-task bill
Same $2 and $10 as Sonnet 5, half of Opus 5.5
- Anthropic released Claude Sonnet 5.5 on September 28, calling it the second model in the Claude 5.5 family (Anthropic)
- Haiku 5.5 is still listed as coming "in the coming weeks"
- The price per token did not move, so any saving has to come from fewer tokens per task
| Price per 1M tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Input | $2 | $4 |
| Output | $10 | $20 |
| Cache read | $0.20 | $0.20 |
| Cache write | $2.50 | $5 |
Source: Anthropic, Cost and speed table, checked September 29, 2026
- Anthropic says Sonnet 5.5 costs up to 30% less per task than Sonnet 5 in its own testing
- Artificial Analysis says its $7.60 per task at Max is about 50% higher than Sonnet 5's (Artificial Analysis)
- The two are different setups, and I have not seen a like-for-like reconciliation
- Customers quoted on Anthropic's page report anything from 12% fewer total tokens (Box) to about 121k against 497k tokens per answer (Balyasny, private finance set)
- Batch processing takes 50% off and caching up to 90%, while US-only inference costs 1.1× (Anthropic)
- Context is 1M tokens and max output 128K, the same numbers the pre-launch leaks reprinted from Sonnet 5 (models overview, my pre-release post)
Half the price per token, 60% more tokens
- Artificial Analysis put Sonnet 5.5 at Max at 56 on its Intelligence Index against 58 for Opus 5.5 at Max
- It used about 193k output tokens per index task, the most Artificial Analysis has measured, around 60% more than Opus 5.5 at Max and about 7× GPT-6 Astra at Max
- That is how a model with half the token price ended up costing $7.60 per task against Opus's $5.98
- The discount mostly went to the extra tokens
| Effort | Sonnet 5.5 score / cost per task | Opus 5.5 score / cost per task |
|---|---|---|
| Low | 36 / $0.41 | 42 / $0.55 |
| Medium | 41 / $0.59 | 51 / $1.34 |
| High | 47 / $1.08 | 54 / $1.82 |
| Xhigh | 52 / $2.74 | 56 / $3.46 |
| Max | 56 / $7.60 | 58 / $5.98 |
Source: Artificial Analysis Intelligence Index v4.3.2, default fallback on, Sonnet 5.5 and Opus 5.5 release pages, checked September 29, 2026. Cost is the weighted average over all ten index evaluations, not coding alone
- At the same effort label Sonnet is cheaper on every row except Max, which flatters it
- Matched by score, Opus at Medium (51 for $1.34) sits next to Sonnet at Xhigh (52 for $2.74), and Opus at Xhigh ties Sonnet at Max at 56 for $3.46 against $7.60
- Those two comparisons are our arithmetic from the published figures: Opus costs about 49% and 46% of Sonnet's in them
- Artificial Analysis says Sonnet 5.5 sits off its score-versus-cost frontier, and that High is its most competitive setting, narrowly behind GPT-6 Sol at about the same cost per task
- It tested a pre-release build that Anthropic says had a structured-output bug, and says it will re-run the evaluations
The papers show tokens, not retries
- My gripe is that the scores felt bought with unlimited retries, because on Artificial Analysis's benchmarks Sonnet's cost per task ends up above Opus's despite the half-price tokens
- I found no source describing unlimited retries: Anthropic averages five trials at Max effort, and Artificial Analysis averages pass@1 across three attempts on its coding agent benchmarks
- What the sources do document is token volume, which produces the same bill without any retry
- Anthropic's headline table is Max effort, the setting that costs the most per task
| Benchmark (Sonnet 5.5 at Max) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% (Xhigh) |
| FrontierCode 1.1 Main | 46.2% | 42.4% | 54.4% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% |
| GDPval-AA v2.1, Elo | 1844 | 1449 | 1846 |
Source: Anthropic and the system card. Cursor ran CursorBench, Artificial Analysis ran GDPval-AA, Cognition ran FrontierCode; Sonnet 5 and Opus 5.5 settings as reported by Anthropic
- Artificial Analysis's own Terminal-Bench 4.0 run gives Sonnet 5.5 at Max 64% against 60% for Opus 5.5, lower than Anthropic's 70.6% for the same model
- Anthropic's 4.2-point Terminal-Bench 4.0 lead over Opus 5.5 carries a standard error of about ±2.5 points for Sonnet and ±2.6 for Opus, so it is not settled
- Sonnet 5's 10.3% is Anthropic's figure, and the system card's Terminal-Bench section does not list Sonnet 5 at all
- I could not read Artificial Analysis's Coding Agent Index numbers for Sonnet 5.5, so the cost table above is the whole index, not coding alone
Max lost to Xhigh, and High is 7.7 points under Max

Source: Artificial Analysis release pages and Anthropic's system card, section 8.8 (CursorBench, run by Cursor). Chart by HSL
- On FrontierCode, Sonnet 5.5 scores 52.1% at Xhigh but only 46.2% at Max
- Anthropic's footnote says Max more often ran Claude Code's code-review skill, which splits the review across many subagents, and in two cases Cognition examined that caused a timeout or edits outside the task
- The top setting lost to the one below it because it called in too many reviewers
- On CursorBench 4.0, the scores run 39.2% at Medium, 47.8% at High, 53.1% at Xhigh and 55.5% at Max, so High is 7.7 points below Max
- That still clears Sonnet 5's best reported 34.1% by a wide margin
My reason for High is a feeling, not a benchmark
- With reasoning set to around High, I found the price-performance good
- Having used it, I am personally satisfied
- That is my impression from using it, not a measured result
- Artificial Analysis's numbers do not contradict it, since it calls High Sonnet's most competitive setting, though Opus at Medium scores 4 points higher (51 against 47) for $0.26 more per index task
- Anthropic's docs also point to High for the API, where it is the default, and to Medium for well-specified agentic coding, which is the default in Claude Code and the Claude apps (migration guide)
- Anthropic positions Sonnet 5.5 for well-scoped everyday tasks, bug fixes and documents, and Opus 5.5 for complex work needing judgment
Settings that can break the switch
- The migration guide lists five settings that return a 400 error on Sonnet 5.5: thinking budgets, sampling parameters, assistant prefill, forced tool choice and
thinking: {"type": "disabled"} - To turn up-front thinking off, use
between_tools, which is accepted at Low, Medium and High effort and returns a 400 at Xhigh or Max - Higher-risk cybersecurity tasks visibly fall back to Sonnet 5, and in Anthropic's Terminal-Bench run 1.2% of requests were answered by a fallback model
- Preserved thinking is now tied to the account that created it, which matters if you switch accounts mid-session in Claude Code
- Anthropic says its effort levels were recalibrated, so a level does not mean the same thinking as on Sonnet 5
Test it on your own tasks
| Effort | Where I would start | Basis |
|---|---|---|
| Medium | Well-specified agentic coding, latency-sensitive chat | Default in Claude Code and the apps |
| High | Harder or longer tasks, general API use | Default on the Claude API; Artificial Analysis's most competitive Sonnet setting |
| Xhigh or Max | Only if the extra passes pay for the tokens | Max cost $7.60 per index task and lost to Xhigh on FrontierCode |
- Pick 20 real tasks you already know the right answer to
- Run them on Sonnet at Medium and High, and on Opus at Medium as the comparison
- Record how many results you accepted and how many output tokens each run used
- Divide the bill by accepted tasks, using each model's per-token prices
- Keep the model and setting where that number is lowest, not the one the leaderboard shows at Max
Community Reactions
- In r/ClaudeCode, one commenter objected to Sonnet's higher Max cost per task, while another urged comparing Medium and High before writing off the model (Reddit)
- An r/ClaudeAI tester's ten coding tasks, each repeated three times at High, yielded 30/30 passes for both models in Claude Code, with Sonnet averaging ~$0.15 versus Opus ~$0.36 in API cost per attempt; the author calls the suite saturated (Reddit)
- Another r/ClaudeAI poster, explicitly without running their own tests, pointed to Anthropic's note that extra review subagents made Max lose to Xhigh on FrontierCode (Reddit)
Q&A (Field Notes)
- Q. Is Sonnet 5.5 cheaper than Opus 5.5 per task?
- At the same effort label mostly yes, but Artificial Analysis measured $7.60 against $5.98 at Max, and Opus at Xhigh matches Sonnet at Max for $3.46
- The gap comes from Sonnet using about 60% more output tokens at Max
- Q. Were the benchmark scores pushed up by retries?
- I found no source saying so: Anthropic averages five trials and Artificial Analysis averages pass@1 over three attempts on its coding agent benchmarks
- I did not read the harness details of Cursor or Cognition, so this covers only what Anthropic and Artificial Analysis published
- Q. Does Sonnet 5.5 replace Opus 5.5?
- Anthropic says Opus 5.5 remains clearly stronger on complex, open-ended work
- Sonnet 5.5 at Max nearly ties it on GDPval-AA (1844 against 1846) but trails on FrontierCode at 46.2% against 54.4%
- Q. Which effort should I use in Claude Code?
- Anthropic's guidance is Medium for well-specified tasks and High for harder or longer ones, and Medium is the default there
- Q. Will my Sonnet 5 code run unchanged?
- Not if it sends
thinking: {"type": "disabled"}or non-default sampling values, since both return a 400 error
- Not if it sends
