3-Line TL;DR

  • Opus 5.5 still feels stronger to me on intelligence, while GPT-6.1 Sol interests me as an everyday model inside the Codex environment I prefer
  • Sol and Sonnet 5.5 both charge $2 input and $10 output per million tokens, half Opus's ordinary rates, so Sol cannot win this three-way comparison on price alone
  • I would try Sol for the Codex workflow, keep Sonnet in contention for well-scoped Claude Code tasks, and pay for Opus when better judgment saves corrections

Opus still has my confidence; Sol has another chance

  • In my earlier Astra versus Opus notes, I preferred Opus 5.5's finished output and value
  • My current impression still puts Opus ahead on intelligence
  • GPT-6 Sol was disappointing enough that a cheap replacement alone would not interest me
  • What makes 6.1 interesting is the prospect of useful everyday performance in a working environment I already like
  • OpenAI reports a 6.4-percentage-point DeepSWE 1.1 improvement over the previous Sol's best result at lower effort and cost, which is a reason to test it again (OpenAI)
  • That is a maker-reported result, and this article is my current buying judgment rather than a fresh controlled three-model test
  • I do not need the smartest model to rename a file, but I do need it to rename the right file

Sol and Sonnet share the same ordinary rates

Standard input and output prices for GPT-6.1 Sol, Sonnet 5.5 and Opus 5.5, with cached-input rates

Model Input Cached input / cache read Output
GPT-6.1 Sol $2 $0.10 $10
Claude Sonnet 5.5 $2 $0.20 $10
Claude Opus 5.5 $4 $0.20 $20

Sources: OpenAI model docs, Sonnet announcement, Opus model docs; USD per million text tokens, standard API processing, checked September 29, 2026

  • Sol's rates here apply to requests with at most 272K input tokens; beyond that the full request uses double input/cache rates and 1.5 times output rates
  • Tools, cache writes and processing premiums are excluded, and subscription allowances are a separate comparison
  • For an illustrative 100K uncached input and 20K output tokens, the bill is $0.40 on Sol or Sonnet and $0.80 on Opus
  • Those are calculations from fixed token counts, not measured costs for an actual job
  • Two identical $0.40 attempts already match one $0.80 attempt, before accounting for my review time
  • The headline discount is attractive, although buying the same correction twice is a familiar subscription benefit
  • Sol's lower cache-read rate helps with eligible repeated input, but it does not make the whole bill half Sonnet's

Sonnet makes this a harder choice

  • My Sonnet 5.5 experience was positive at High for routine work, with value slipping on harder tasks
  • That gives Sonnet a real place in this comparison instead of treating every Claude job as an Opus bill
  • Anthropic positions Sonnet for well-scoped tasks and Opus for complex, open-ended work, while Claude Code defaults Sonnet to Medium (Official positioning)
  • Turning effort up is not a guaranteed fix: Anthropic reports FrontierCode 1.1 Main at 52.1% for Sonnet Xhigh and 46.2% for Max
  • Its footnote describes extra code-review subagents contributing to timeouts or out-of-scope changes in two examined cases
  • A small fix can apparently acquire a surprisingly large review committee
  • Those results do not measure GPT-6.1 Sol, and the older GPT-6 Sol column in that launch table cannot stand in for it
  • I would compare accepted results at sensible settings before deciding that any model is cheaper per task

Codex is why I lean toward Sol for daily work

  • As a working environment, Codex currently feels more capable and convenient to me than Claude Code
  • The pace of integration matters, as does moving between writing, code and images without rebuilding the workflow
  • Built-in image generation is a concrete example, using gpt-image-2 within Codex usage limits (Image generation docs)
  • Claude Code also supports MCP, skills, command execution and cloud workflows, so feature checkboxes alone do not explain my preference (Claude Code overview)
  • Sol's value proposition looks particularly strong to me when I count that whole setup, while a proven Sol-versus-Sonnet saving still needs comparable jobs
What I need Where I would start What would change my choice
Everyday work across code, text and images in Codex GPT-6.1 Sol Repeated corrections erase the low rates
A defined task in an established Claude Code setup Sonnet 5.5 at Medium or High The task needs repeated judgment calls
Complex planning or difficult open-ended work Opus 5.5 A cheaper model produces equally acceptable results
  • For a fair trial I would keep the task, available tools and acceptance criteria fixed, then record each model's effort setting
  • I would divide all attempt costs by accepted tasks and record review minutes separately
  • A workflow preference can settle my choice even when the token-price table ends in a tie

Community Reactions

  • In r/codex, one commenter asks which effort levels the unlabeled points represent before treating High as Sol's sweet spot (Reddit)
  • In r/ClaudeCode, one reader criticizes Sonnet's Max cost while another argues that Medium and High deserve a separate comparison (Reddit)

Q&A (Field Notes)

  • Q. Does Opus being smarter make Codex a worse choice?
    • My impression of model intelligence and my preference for a working environment answer different questions, and the task decides how much each matters
  • Q. Is Sol half the price of Sonnet?
    • Only the listed cache-read rate is half; ordinary input and output rates are the same, and actual usage can differ
  • Q. Should I move every Sonnet task to Max?
    • I would first check whether Medium or High already meets the acceptance criteria, then compare Opus before paying for more effort
  • Q. Have these three models been tested here on the same jobs?
    • No, this combines my stated preferences, earlier experience, official rates and attributed maker evaluations; the proposed trial is still a next step