3-Line TL;DR
- Opus 5.5 still feels stronger to me on intelligence, while GPT-6.1 Sol interests me as an everyday model inside the Codex environment I prefer
- Sol and Sonnet 5.5 both charge $2 input and $10 output per million tokens, half Opus's ordinary rates, so Sol cannot win this three-way comparison on price alone
- I would try Sol for the Codex workflow, keep Sonnet in contention for well-scoped Claude Code tasks, and pay for Opus when better judgment saves corrections
Opus still has my confidence; Sol has another chance
- In my earlier Astra versus Opus notes, I preferred Opus 5.5's finished output and value
- My current impression still puts Opus ahead on intelligence
- GPT-6 Sol was disappointing enough that a cheap replacement alone would not interest me
- What makes 6.1 interesting is the prospect of useful everyday performance in a working environment I already like
- OpenAI reports a 6.4-percentage-point DeepSWE 1.1 improvement over the previous Sol's best result at lower effort and cost, which is a reason to test it again (OpenAI)
- That is a maker-reported result, and this article is my current buying judgment rather than a fresh controlled three-model test
- I do not need the smartest model to rename a file, but I do need it to rename the right file
Sol and Sonnet share the same ordinary rates

| Model | Input | Cached input / cache read | Output |
|---|---|---|---|
| GPT-6.1 Sol | $2 | $0.10 | $10 |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 |
| Claude Opus 5.5 | $4 | $0.20 | $20 |
Sources: OpenAI model docs, Sonnet announcement, Opus model docs; USD per million text tokens, standard API processing, checked September 29, 2026
- Sol's rates here apply to requests with at most 272K input tokens; beyond that the full request uses double input/cache rates and 1.5 times output rates
- Tools, cache writes and processing premiums are excluded, and subscription allowances are a separate comparison
- For an illustrative 100K uncached input and 20K output tokens, the bill is $0.40 on Sol or Sonnet and $0.80 on Opus
- Those are calculations from fixed token counts, not measured costs for an actual job
- Two identical $0.40 attempts already match one $0.80 attempt, before accounting for my review time
- The headline discount is attractive, although buying the same correction twice is a familiar subscription benefit
- Sol's lower cache-read rate helps with eligible repeated input, but it does not make the whole bill half Sonnet's
Sonnet makes this a harder choice
- My Sonnet 5.5 experience was positive at High for routine work, with value slipping on harder tasks
- That gives Sonnet a real place in this comparison instead of treating every Claude job as an Opus bill
- Anthropic positions Sonnet for well-scoped tasks and Opus for complex, open-ended work, while Claude Code defaults Sonnet to Medium (Official positioning)
- Turning effort up is not a guaranteed fix: Anthropic reports FrontierCode 1.1 Main at 52.1% for Sonnet Xhigh and 46.2% for Max
- Its footnote describes extra code-review subagents contributing to timeouts or out-of-scope changes in two examined cases
- A small fix can apparently acquire a surprisingly large review committee
- Those results do not measure GPT-6.1 Sol, and the older GPT-6 Sol column in that launch table cannot stand in for it
- I would compare accepted results at sensible settings before deciding that any model is cheaper per task
Codex is why I lean toward Sol for daily work
- As a working environment, Codex currently feels more capable and convenient to me than Claude Code
- The pace of integration matters, as does moving between writing, code and images without rebuilding the workflow
- Built-in image generation is a concrete example, using
gpt-image-2within Codex usage limits (Image generation docs) - Claude Code also supports MCP, skills, command execution and cloud workflows, so feature checkboxes alone do not explain my preference (Claude Code overview)
- Sol's value proposition looks particularly strong to me when I count that whole setup, while a proven Sol-versus-Sonnet saving still needs comparable jobs
| What I need | Where I would start | What would change my choice |
|---|---|---|
| Everyday work across code, text and images in Codex | GPT-6.1 Sol | Repeated corrections erase the low rates |
| A defined task in an established Claude Code setup | Sonnet 5.5 at Medium or High | The task needs repeated judgment calls |
| Complex planning or difficult open-ended work | Opus 5.5 | A cheaper model produces equally acceptable results |
- For a fair trial I would keep the task, available tools and acceptance criteria fixed, then record each model's effort setting
- I would divide all attempt costs by accepted tasks and record review minutes separately
- A workflow preference can settle my choice even when the token-price table ends in a tie
Community Reactions
- In r/codex, one commenter asks which effort levels the unlabeled points represent before treating High as Sol's sweet spot (Reddit)
- In r/ClaudeCode, one reader criticizes Sonnet's Max cost while another argues that Medium and High deserve a separate comparison (Reddit)
Q&A (Field Notes)
- Q. Does Opus being smarter make Codex a worse choice?
- My impression of model intelligence and my preference for a working environment answer different questions, and the task decides how much each matters
- Q. Is Sol half the price of Sonnet?
- Only the listed cache-read rate is half; ordinary input and output rates are the same, and actual usage can differ
- Q. Should I move every Sonnet task to Max?
- I would first check whether Medium or High already meets the acceptance criteria, then compare Opus before paying for more effort
- Q. Have these three models been tested here on the same jobs?
- No, this combines my stated preferences, earlier experience, official rates and attributed maker evaluations; the proposed trial is still a next step
