Reasoning
GPT-6 Sol leads by 6.9 points.
Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.
THE HEAD-TO-HEAD
Two models. The same task mix. Clear tradeoffs.
Configurations: max / max.
Snapshot · LiveBench question set: 25 June 2026 · View every source ↗
Customize this comparison ↗GPT-6 Sol leads by 7.0 points on our General agent profile. GPT-6 Luna has the lower baseline API estimate at $2 per 1,000 calls. Test both on your workflow before choosing.
| Measure | GPT-6 Solmax | GPT-6 Lunamax |
|---|---|---|
| Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights. | 74.7 | 67.7 |
| Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. | 62.7 | 76.2 |
| Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market. | Insufficient evidence | Insufficient evidence |
| Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums. | $40 | $2 |
| Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits. | 1.05M | 1.05M |
| Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly. | Supported | Supported |
| ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents. | 88.7 | 81.8 |
| CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects. | 81.8 | 79.0 |
| Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model. | 52.9 | 51.2 |
| MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use. | 96.4 | 89.1 |
| Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information. | 81.2 | 73.4 |
| LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage. | 85.3 | 73.8 |
| Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format. | 68.6 | 55.9 |
| Input / output per 1M | $2 / $10 | $0.1 / $0.5 |
GPT-6 Sol leads by 6.9 points.
Can it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.
GPT-6 Sol leads by 2.8 points.
Can it write or complete code that passes tests? These are contained programming problems, not entire software projects.
GPT-6 Sol leads by 1.7 points.
Can it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.
GPT-6 Sol leads by 7.2 points.
Can it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.
GPT-6 Sol leads by 7.8 points.
Can it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.
GPT-6 Sol leads by 11.5 points.
Can it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.
GPT-6 Sol leads by 12.6 points.
Can it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.
Start with your task, not the overall winner. Coding agents need a strong edit-and-test loop; research agents need structured information handling; support agents need instructions followed consistently. Our default weights are only one interpretation of those needs.
Both models are compared on LiveBench’s 25 June 2026 question set. The tested reasoning configurations differ where shown. Higher thinking budgets may improve results while changing token use and latency. No speed advantage is inferred from these scores.
All prices are USD. The baseline uses 1,000 calls with 10,000 input tokens and 2,000 billed output tokens per call. Include reasoning tokens; caching, tools, retries and provider premiums are excluded. Change the workload and weights to see whether the tradeoff changes.
Consensus is available only with an explicitly matched Arena configuration. A missing score is not evidence that a model is worse. Small differences can be sensitive to the tasks, configuration and sampling. Our aggregate is not a statistical confidence statement.