Within one week, both of Claude Code's default models changed. Opus 5.5 landed on September 22, Sonnet 5.5 followed on September 28, and suddenly "which one do I pick?" has a new answer. Search for opus 5.5 vs sonnet 5.5 and you get benchmark pages. Useful, but none of them tell you what to type in Claude Code on a Tuesday morning with a PRD to review and a bug to fix.
This is that guide. It covers which model to pick, which effort level to run it at, and whether turning on ultracode changes the answer. The short version: Sonnet 5.5 reaches near parity with Opus 5.5 at half the token price, but "cheaper per token" is not "cheaper per task". That gap is the whole story.
Key Takeaways
- Sonnet 5.5 reaches near parity with Opus 5.5 at half the token price, even winning Terminal-Bench 4.0 (70.6% vs 66.4%), with real gaps only on FrontierCode and Humanity's Last Exam.
- Cheaper per token is not cheaper per task: Artificial Analysis measured Sonnet 5.5 at about $7.60 per task at max effort, its heaviest token use yet.
- Treat effort as a cost lever: both models now default to medium effort, so check `/effort` after upgrading, since it is saved per model.
- Default to Sonnet 5.5 at medium or high for well-scoped work, and switch to Opus 5.5 when judgment matters more than speed, such as ambiguous bugs and architecture calls.
- Ultracode is a separate toggle for breadth, not difficulty: use it when work splits into many independent pieces, and raise effort for one hard single-threaded problem.
- For PMs, Sonnet 5.5 at medium covers PRDs, specs and prototypes, while Opus 5.5 with ultracode suits competitive teardowns and multi-angle spec reviews.
Learn this hands-on
Go from idea to a live product with real users. Build and ship your first SaaS with Claude Code in a live cohort, with auth, payments and deployment done by the end. Join the Ship Your First SaaS with Claude Code cohort.
What changed in one week
Two defaults flipped, and it is easy to miss the second one.
- September 22: Opus 5.5 became Claude Code's default Opus (version 2.1.280). The same day, Pro and Team Standard plans switched their default model from Sonnet to Opus. Max, Team Premium and Enterprise already defaulted to Opus.
- September 28: Sonnet 5.5 became the default Sonnet (version 2.1.284). In the same release, ultracode became its own toggle instead of being welded to a single effort level.
Both models also share a default effort of medium in Claude Code. Opus 5 defaulted to high, so if you upgraded without touching settings, you are now running a lighter reasoning budget than before. The opus alias now resolves to Opus 5.5 and sonnet to Sonnet 5.5 on the Anthropic API.
Opus 5.5 vs Sonnet 5.5: price and benchmarks
Here are the numbers from Anthropic's own pages for Opus 5.5 and Sonnet 5.5, plus the models overview.
| Opus 5.5 | Sonnet 5.5 | |
|---|---|---|
| Model id | claude-opus-5-5 | claude-sonnet-5-5 |
| Input / output per MTok | $4 / $20 | $2 / $10 |
| Cache reads per MTok | $0.20 | $0.20 |
| Cache writes per MTok | $5 | $2.50 |
| Context / max output | 1M / 128K | 1M / 128K |
| Terminal-Bench 4.0 | 66.4% | 70.6% |
| CursorBench 4.0 | 57.8% | 55.5% |
| OSWorld 2.1 | 81.8% | 80.1% |
| Humanity's Last Exam | 67.7% | 64.5% |
| FrontierCode 1.1 | 54.4% | 46.2% (at high) |
| GDPval-AA v2.1 (Elo) | 1846 | 1844 |
Two things jump out. First, Sonnet 5.5 wins Terminal-Bench 4.0 outright and ties on GDPval. Second, the only real gaps are on the hardest reasoning-heavy tests: FrontierCode and Humanity's Last Exam. That matches how Anthropic positions the pair: "Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets."
Opus 5.5 also got cheaper than its predecessor. Opus 5 was $5 input and $25 output, and Anthropic says a typical workload costs about 40% less on Opus 5.5, with output more than 30% faster. If you are comparing against the previous generation, our Claude Opus 5 breakdown has the context.
The per-task trap
Here is where the "half the price" pitch gets complicated. A model that costs half as much per token only saves money if it uses a similar number of tokens to finish the job.
Artificial Analysis, an independent benchmarking group, measured exactly this on September 29:
- Sonnet 5.5 scores 56 on their Intelligence Index, just 2 points behind Opus 5.5 at max effort.
- On Terminal-Bench 4.0 they measured 64% for Sonnet 5.5 versus 60% for Opus 5.5, so the independent run also puts Sonnet ahead.
- But they call Sonnet 5.5 "the heaviest token use we have measured": about $7.60 per task at max effort, roughly 50% higher cost per task than Sonnet 5, because it emits far more output tokens to get there.
So parity comes with heavy token usage. Anthropic's own claim is more modest: up to 30% lower cost per task than Sonnet 5, which is consistent with a model that is cheaper at normal effort but burns tokens when you push it to the ceiling.
The practical lesson is that effort level is a cost lever as much as a quality lever. Sonnet 5.5 at medium is a very different bill from Sonnet 5.5 at max.
There is one more detail worth knowing. Anthropic's usage data shows the input-to-output token ratio in Claude Code moved from 189:1 to 324:1 over six months, with sessions 3.3 times longer per prompt (Anthropic, September 24). In other words, agentic coding is dominated by reading context, not writing output. Cache reads cost $0.20 per MTok on both models, so on long, cache-heavy sessions the per-token gap between them matters less than the headline $4 versus $2 suggests. Our Claude API pricing guide walks through how to model this for your own workload.
Which model, which effort, ultracode or not
Both models support five effort levels in Claude Code: low, medium, high, xhigh and max. Ultracode is a separate toggle on top. When it is on, Claude orchestrates dynamic workflows for substantive tasks, spinning up parallel subagents from a script instead of working through the job in one thread. It is available on any model that supports xhigh, so both Opus 5.5 and Sonnet 5.5 qualify.
This table is my recommendation, built from the positioning and data above. It is a starting point to test on your own work, not a law.
| Task | Model | Effort | Ultracode |
|---|---|---|---|
| Fix a bug, small well-scoped change | Sonnet 5.5 | medium | off |
| Feature across several files | Sonnet 5.5 | high | off |
| Hard bug, unclear root cause | Opus 5.5 | high | off |
| Audit or review of a large area of code | Opus 5.5 | xhigh | on |
| Docs, specs, slides, spreadsheets | Sonnet 5.5 | medium | off |
| Research across many sources | Opus 5.5 | high | on |
| Architecture or high-judgment decision | Opus 5.5 | xhigh | off |
The reasoning behind the pattern:
- Default to Sonnet 5.5 at medium or high. Most daily work is well-scoped, which is exactly where Anthropic says Sonnet is strongest and where the benchmark gap is smallest.
- Move to Opus 5.5 when the cost of a wrong answer is high. Ambiguous bugs, architecture trade-offs and anything that needs judgment rather than execution.
- Turn ultracode on for breadth, not difficulty. It shines when the task splits into many independent pieces (auditing dozens of files, comparing many sources). It does not make a single hard decision smarter.
- Avoid max effort on Sonnet 5.5 by default. That is where the Artificial Analysis numbers show the token bill climbing.
Does ultracode change the answer?
Mostly it changes how much work happens in parallel, not which model is right. Each request uses more tokens with ultracode on, and the large-workflow warning (25 agents or 1.5M tokens) is suppressed, so you will not get a speed bump telling you the run is getting expensive. Anthropic's workflows documentation gives size guidelines: small is fewer than 5 agents (the default on Pro), medium fewer than 10, large fewer than 50, with a cap of 1,000 agents per run and 16 running concurrently by default.
The question "opus 5.5 max vs ultracode" comes up a lot. They are different dials. Max is how hard one model thinks on one thread. Ultracode is how many threads run at once. For a deep single-threaded problem, raise effort. For a wide problem, turn ultracode on. Our ultracode explainer covers the mechanics in detail.
And "sonnet ultracode" is a perfectly valid combination. If the job is wide but each piece is simple, Sonnet 5.5 with ultracode on is a sensible way to keep the per-agent cost down.
How to set it in Claude Code
The settings live in a few places, and one gotcha is worth flagging first.
- Pick the model. Run
/model, or set it at launch. The aliasesopusandsonnetnow point at the 5.5 releases. The full reference is in the model configuration docs. - Set effort per model. Run
/effort. Effort is saved per model now, so a saved effort level from before per-model effort does not carry over to Opus 5.5. It starts at medium. If you remember tuning it to high on Opus 5, check it again. - Toggle ultracode. In
/effort, press Tab, or run/effort ultracode onoroff. It no longer forces xhigh and stays on at any effort level. Launching withclaude --effort ultracodestill sets xhigh plus ultracode on. - Try opusplan. The
opusplanalias runs Opus in plan mode and Sonnet for execution. It maps neatly onto the table above: expensive judgment for the plan, cheaper execution for the build. - Consider the advisor. The /advisor tool (see our Claude Code advisor guide) lets a Sonnet 5.5 main model consult an Opus 5.5 advisor at decision points, so you keep Sonnet's pace and still get a second opinion on the hard calls.
One small quality-of-life change: switching effort mid-session no longer rebuilds the prompt cache on Opus 5.5, so you can drop to medium for a routine step and go back up without paying for a cold cache.
Opus 5.5 vs Sonnet 5.5 for product managers
If you are a PM using Claude Code, most of your work is not the hardest coding problem on earth. It is documents, synthesis and prototypes, which lands squarely in Sonnet 5.5 territory.
- PRDs and specs: Sonnet 5.5 at medium. Anthropic specifically calls out polished documents, slides and spreadsheets as a strength. Switch to Opus 5.5 for the review pass, where you want a skeptical reader rather than a fast writer.
- Discovery synthesis: Sonnet 5.5 for a handful of interviews. For a large pile of transcripts and tickets, Opus 5.5 with ultracode on lets the work split across many parallel readers.
- Competitive teardowns: Opus 5.5 with ultracode, one agent per competitor, then a synthesis. This is the wide-not-hard pattern.
- Prototypes: Sonnet 5.5 at high. It is a well-scoped build, and you will iterate on it many times, so cost per attempt matters.
- Multi-angle spec reviews: Opus 5.5 with ultracode, with separate reviewers for engineering, design and legal angles.
If you are new to the tool, our Claude Code mods guide is a good next step, and the Master Claude Code course takes you from first prompt to shipping agents. The habit to build is simple: start cheap, escalate when the output disappoints, and watch effort as closely as you watch model choice.
Product manager and want to work like this? This is exactly what we teach in Claude Code for PMs, our live cohort for product teams: 3 live sessions of 90 minutes over 2 weeks. Every PM ships a real feature, builds their own agent, and gets personalized written feedback.
The bottom line
Pick Sonnet 5.5 by default, at medium or high effort, for well-scoped work. Reach for Opus 5.5 when judgment matters more than speed. Use ultracode when the task is wide, not when it is merely hard. And never judge a model by its price per token alone: the number that matters is what a finished task costs you.
The two defaults will keep moving. Anthropic is shipping model and Claude Code changes weekly, so check your /model and /effort settings after each release rather than assuming last month's configuration still holds. If you are weighing these two against an older default, the Fable 5.1 versus Sonnet 5 comparison covers the same kind of default-model decision for an earlier lineup.