Reasoning
GPT-6 Astra leads by 0.5 points.
Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.
THE HEAD-TO-HEAD
Two models. The same task mix. Clear tradeoffs.
Configurations: max / max.
Snapshot · LiveBench question set: 25 June 2026 · View every source ↗
Customize this comparison ↗Claude Opus 5.5 leads by 0.8 points on our General agent profile. Claude Opus 5.5 has the lower baseline API estimate at $80 per 1,000 calls. Test both on your workflow before choosing.
| Measure | Claude Opus 5.5max | GPT-6 Astramax |
|---|---|---|
| Agent Fit | 79.4 | 78.6 |
| Practical Value | 61.0 | 55.6 |
| Evidence Consensus | Insufficient evidence | 53.1 |
| Workload cost | $80 | $200 |
| Context window | 1.00M | 1.05M |
| Tool calling | Supported | Supported |
| Reasoning | 92.2 | 92.7 |
| Coding | 89.3 | 80.4 |
| Agentic coding | 71.7 | 57.3 |
| Mathematics | 97.1 | 96.8 |
| Data analysis | 80.3 | 83.0 |
| Language | 86.3 | 89.4 |
| Instruction following | 65.7 | 75.6 |
| Input / output per 1M | $4 / $20 | $10 / $50 |
GPT-6 Astra leads by 0.5 points.
Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.
Claude Opus 5.5 leads by 8.9 points.
Can it write or complete code that passes tests? These are contained programming problems, not entire software projects.
Claude Opus 5.5 leads by 14.4 points.
Can it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.
Claude Opus 5.5 leads by 0.3 points.
Can it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.
GPT-6 Astra leads by 2.7 points.
Can it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.
GPT-6 Astra leads by 3.2 points.
Can it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.
GPT-6 Astra leads by 9.8 points.
Can it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.
Start with your task, not the overall winner. Coding agents need a strong edit-and-test loop; research agents need structured information handling; support agents need instructions followed consistently. Our default weights are only one interpretation of those needs.
Both models are compared on LiveBench’s 25 June 2026 question set. The tested reasoning configurations differ where shown. Higher thinking budgets may improve results while changing token use and latency. No speed advantage is inferred from these scores.
All prices are USD. The baseline uses 1,000 calls with 10,000 input tokens and 2,000 billed output tokens per call. Include reasoning tokens; caching, tools, retries and provider premiums are excluded. Change the workload and weights to see whether the tradeoff changes.
Consensus is available only with an explicitly matched Arena configuration. A missing score is not evidence that a model is worse. Small differences can be sensitive to the tasks, configuration and sampling. Our aggregate is not a statistical confidence statement.
Add more models for a meaningful zone. Numbers mark efficient models in or nearest this zone. This shortlist changes with the metric, workload and filters.