Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

OPENAI · PROPRIETARY

GPT-5.6 Sol

Benchmarks, pricing and tradeoffs for agent builders.
Tested configuration: source ID max; display xhigh.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Add to a custom comparison ↗
Agent Fit77.1General agent profile
Practical Value59.370% capability · 30% cost
Estimated monthly API cost$801,000 calls · 10K input + 2K output
Catalog context window1.05MAdvertised capacity

Capability profile

Explain these tests ↗
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents.91.7 / 100
CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects.83.9 / 100
Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model.56.2 / 100
MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use.96.2 / 100
Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information.79.8 / 100
LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage.87.7 / 100
Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format.71.8 / 100

What the evidence says

GPT-5.6 Sol scores 81.1 on the equal-weight LiveBench category average. Its highest measured category is mathematics (96.2); its lowest is agentic coding (56.2). Categories differ in difficulty, so these are test results, not universal strengths or weaknesses.

Our General agent profile produces an Agent Fit score of 77.1 by emphasizing reasoning, instruction following, agentic coding, coding and data analysis. See the exact weights or adjust the profile in the interactive comparison.

Independent preference evidence

A matching Arena configuration is not verified in this snapshot. Evidence Consensus is withheld. A result for another reasoning setting or version is not substituted.

The LiveBench row ID and display label disagree about reasoning effort. Both labels are preserved above; this model is excluded from Consensus until that mismatch is resolved.

API features to verify

The catalog lists tool-calling support and image input. These are declared features, not evaluations of tool accuracy or visual understanding.

This is a proprietary model accessed through a provider. Availability, rate limits, data handling and service tiers depend on the endpoint and account.

GPT-5.6 Sol API pricing

Input / 1M tokens$4
Output / 1M tokens$20

At the baseline of 1,000 calls, 10,000 input tokens and 2,000 billed output tokens per call, the text API estimate is $80. Reasoning belongs in billed output. Caching, tools, retries, long-context premiums, service tiers, hosting and taxes are excluded.

Observed source difference: OpenRouter lists $2 input / $10 output per million tokens. Our calculator uses the direct-provider rates above. These are distinct offers, not prices to average.

Trace this model to the source

LiveBench row ID
gpt-5.6-sol-max
Question-set release
2026-06-25 · Source CSV ↗
Arena configuration
No verified matching observation
Reviewed snapshot
2026-09-23

Compare alternatives

For production decisions, test representative tasks with a fixed harness and success criteria. Neither a large context window nor a strong benchmark score guarantees reliability in your own agent.

AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗