Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

THE HEAD-TO-HEAD

Claude Opus 5.5
vs Claude Sonnet 5

Two models. The same task mix. Clear tradeoffs.
Configurations: max / xhigh.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Customize this comparison ↗
Our take

Claude Opus 5.5 leads by 6.1 points on our General agent profile. Claude Sonnet 5 has the lower baseline API estimate at $40 per 1,000 calls. Test both on your workflow before choosing.

ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5Claude Sonnet 5
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxClaude Sonnet 5xhigh
Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights.79.473.3
Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved.61.061.8
Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market.Insufficient evidenceInsufficient evidence
Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.$80$40
Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits.1.00M1.00M
Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly.SupportedSupported
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.92.288.7
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.89.380.7
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.71.759.4
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.97.192.9
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.80.371.7
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.86.375.0
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.65.763.9
Input / output per 1M$4 / $20$2 / $10

Where their benchmark results differ

Reasoning

Claude Opus 5.5 leads by 3.5 points.

Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.

Coding

Claude Opus 5.5 leads by 8.6 points.

Can it write or complete code that passes tests? These are contained programming problems, not entire software projects.

Agentic coding

Claude Opus 5.5 leads by 12.3 points.

Can it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.

Mathematics

Claude Opus 5.5 leads by 4.1 points.

Can it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.

Data analysis

Claude Opus 5.5 leads by 8.6 points.

Can it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.

Language

Claude Opus 5.5 leads by 11.3 points.

Can it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.

Instruction following

Claude Opus 5.5 leads by 1.9 points.

Can it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.

How to choose between them

Start with your task, not the overall winner. Coding agents need a strong edit-and-test loop; research agents need structured information handling; support agents need instructions followed consistently. Our default weights are only one interpretation of those needs.

Both models are compared on LiveBench’s 25 June 2026 question set. The tested reasoning configurations differ where shown. Higher thinking budgets may improve results while changing token use and latency. No speed advantage is inferred from these scores.

Cost assumptions

All prices are USD. The baseline uses 1,000 calls with 10,000 input tokens and 2,000 billed output tokens per call. Include reasoning tokens; caching, tools, retries and provider premiums are excluded. Change the workload and weights to see whether the tradeoff changes.

Evidence and uncertainty

Consensus is available only with an explicitly matched Arena configuration. A missing score is not evidence that a model is worse. Small differences can be sensitive to the tasks, configuration and sampling. Our aggregate is not a statistical confidence statement.

The cost-capability tradeoff

Efficient frontierOther modelsHigher + further left = more for less
020406080100$10$100Agent Fit / 100 ↑Estimated workload cost · USD · logarithmic scaleClaude Opus 5.5: 79.4 · $80Claude Sonnet 5: 73.3 · $40
Claude Opus 5.579.4 Agent Fit$80

Open the complete evidence

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗