Read input and output prices together
A model priced at $2 input and $10 output per million tokens costs $40 for 1,000 calls of 10,000 input and 2,000 output tokens each. Different tokenizers can turn the same text into different token counts.
THE MODEL SELECTION SERIES
API prices are usually quoted per million tokens, but your budget depends on how much context you send and how much output the model bills. Change the calculator to compare the same workload across providers.
Snapshot · LiveBench question set: 25 June 2026 · View every source ↗
A model priced at $2 input and $10 output per million tokens costs $40 for 1,000 calls of 10,000 input and 2,000 output tokens each. Different tokenizers can turn the same text into different token counts.
Include billed reasoning tokens in output. Cache reads may be cheaper while cache writes can cost more. Some providers charge higher rates for large prompts or premium service tiers. Our estimate excludes those adjustments; check the linked source before committing a budget.
We use verified direct prices for selected OpenAI and Anthropic models and explicitly labeled OpenRouter catalog rates for other models. A catalog entry can route across endpoints with different prices and limits. Each profile shows the exact catalog identifier and any observed price difference.
The default score combines 70% task-weighted capability with 30% affordability rank within the snapshot. Move the slider to express your tradeoff. It does not predict task success, retries, hosting charges or actual monthly savings.
YOUR WORKLOAD. YOUR PRIORITIES.
Standard text API estimate. Include reasoning tokens in output.
A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.
Practical Value Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. · General agent
70% Agent Fit + 30% affordability. Rankings use your workload; filters do not change the reference cohort.
THE EVIDENCE, SIDE BY SIDE
| Compare | Our task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights. | Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. | Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market. | Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set. | People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference. | Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums. | The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits. | Input / outputUSD / 1M tokens | |
|---|---|---|---|---|---|---|---|---|---|
01 DeepSeek V4.1 Flash ↗DeepSeek · maxOpen weights | 78.9 | 83.5 | — | 81.1 | — | $2.14 | 1.05M | $0.094 / $0.6OpenRouter catalog | |
02 DeepSeek V4 Flash Vision Exp ↗DeepSeek · experimental · source defaultOpen weightsPreview | 75.1 | 79.8 | — | 76.8 | — | $3.52 | 1.05M | $0.22 / $0.66OpenRouter catalog | |
03 DeepSeek V4 Flash 0731 ↗DeepSeek · source defaultOpen weights | 71.1 | 79.2 | — | 74.2 | — | $1.68 | 1.31M | $0.04 / $0.64OpenRouter catalog | |
04 GPT-6 Luna ↗OpenAI · max | 67.7 | 76.2 | — | 72.0 | — | $2 | 1.05M | $0.1 / $0.5Direct API | |
05 GPT-5.6 Luna ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 70.4 | 75.9 | — | 73.6 | — | $4.4 | 1.05M | $0.2 / $1.2Direct API | |
06 GLM-5.3 Flash ↗Z.ai · source defaultOpen weights | 67.2 | 74.8 | — | 71.6 | — | $2.5 | 1.31M | $0.15 / $0.5OpenRouter catalog | |
07 Qwen3.8 27B ↗Alibaba · source defaultOpen weights | 73.5 | 74.7 | — | 75.3 | — | $10.2 | 1.00M | $0.42 / $3OpenRouter catalog | |
08 DeepSeek V4 Pro 0813 ↗DeepSeek · source defaultOpen weights | 73.3 | 73.9 | — | 77.4 | — | $10.56 | 1.05M | $0.66 / $1.98OpenRouter catalog | |
09 GPT-5.4 nano ↗OpenAI · xhigh | 67.7 | 73.4 | — | 69.6 | — | $4.5 | 400K | $0.2 / $1.25Direct API | |
10 DeepSeek V4 Flash ↗DeepSeek · source defaultOpen weights | 61.6 | 73.1 | — | 65.5 | — | $1.24 | 1.05M | $0.0886 / $0.1772OpenRouter catalog | |
11 Gemini 3.7 Flash ↗Google · high | 76.1 | 71.9 | 56.32 sources | 78.8 | 1,490±8 · preliminary | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
12 GLM-5.3 ↗Z.ai · source defaultOpen weights | 73.7 | 71.4 | — | 76.1 | — | $13.68 | 1.31M | $0.84 / $2.64OpenRouter catalog | |
13 Muse Spark 1.3 ↗Meta · xhigh | 79.3 | 70.2 | — | 81.6 | — | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
14 GLM-5.2 ↗Z.ai · source defaultOpen weights | 68.5 | 70.1 | — | 73.2 | — | $10.58 | 1.05M | $0.6496 / $2.04OpenRouter catalog | |
15 Gemini 3.8 Flash ↗Google · high | 73.3 | 70.0 | 50.02 sources | 75.8 | 1,493±9 · preliminary | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
16 Qwen3.6 Plus ↗Alibaba · source default | 63.9 | 69.6 | — | 68.9 | — | $7.15 | 1.00M | $0.325 / $1.95OpenRouter catalog | |
17 MiniMax M3 ↗MiniMax · source default | 63.1 | 69.6 | — | 67.3 | — | $5.4 | 1.05M | $0.3 / $1.2OpenRouter catalog | |
18 Muse Spark 1.2 ↗Meta · xhigh | 76.3 | 68.1 | 68.82 sources | 78.0 | 1,500±11 | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
19 DeepSeek V4 Pro ↗DeepSeek · source defaultOpen weights | 67.1 | 67.9 | — | 71.6 | — | $13.37 | 1.05M | $0.9553 / $1.91OpenRouter catalog | |
20 Gemini 3.6 Flash ↗Google · high | 70.3 | 67.9 | 9.42 sources | 73.6 | 1,480±5 | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
21 Muse Spark 1.1 ↗Meta · source ID xhigh; display highConfiguration discrepancy | 74.0 | 66.5 | — | 75.3 | — | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
22 Nemotron 3 Ultra ↗NVIDIA · 550B A55B · source defaultOpen weights | 63.8 | 66.1 | — | 67.4 | — | $10.8 | 262K | $0.6 / $2.4OpenRouter catalog | |
23 Gemini 3.5 Flash-Lite ↗Google · high | 59.5 | 66.0 | — | 63.9 | — | $8 | 1.05M | $0.3 / $2.5OpenRouter catalog | |
24 Qwen3.6 27B ↗Alibaba · source defaultOpen weights | 60.0 | 65.8 | — | 64.0 | — | $8.6 | 262K | $0.32 / $2.7OpenRouter catalog | |
25 Kimi K2.7 Code ↗Moonshot AI · source defaultOpen weights | 64.8 | 65.8 | — | 68.4 | — | $13.66 | 262K | $0.7062 / $3.3OpenRouter catalog | |
26 Grok 4.6 ↗xAI · source default | 75.3 | 64.9 | — | 78.0 | — | $32 | 500K | $2 / $6OpenRouter catalog | |
27 Grok 4.7 ↗xAI · xhigh | 73.7 | 64.6 | — | 77.4 | — | $25.6 | 500K | $1.6 / $4.8OpenRouter catalog | |
28 Inkling ↗Thinking Machines · xhighOpen weights | 68.9 | 64.1 | — | 71.9 | — | $18.1 | 1.05M | $1 / $4.05OpenRouter catalog | |
29 Kimi K2.6 ↗Moonshot AI · thinkingOpen weights | 66.9 | 63.5 | — | 70.5 | — | $17.5 | 262K | $0.95 / $4OpenRouter catalog | |
30 Grok 4.5 ↗xAI · source default | 73.1 | 63.4 | — | 75.8 | — | $32 | 500K | $2 / $6OpenRouter catalog | |
31 Qwen3.7 Max ↗Alibaba · source default | 70.4 | 62.9 | — | 73.1 | — | $23.6 | 1.00M | $1.48 / $4.43OpenRouter catalog | |
32 GPT-6 Sol ↗OpenAI · max | 74.7 | 62.7 | — | 79.2 | — | $40 | 1.05M | $2 / $10Direct API | |
33 Claude Sonnet 5 ↗Anthropic · xhigh | 73.3 | 61.8 | — | 76.0 | — | $40 | 1.00M | $2 / $10Direct API | |
34 GPT-5.4 mini ↗OpenAI · xhigh | 62.5 | 61.3 | — | 66.4 | — | $16.5 | 400K | $0.75 / $4.5Direct API | |
35 GPT-5.6 Terra ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 74.1 | 61.2 | — | 77.9 | — | $44 | 1.05M | $2 / $12Direct API | |
36 Claude Opus 5.5 ↗Anthropic · max | 79.4 | 61.0 | — | 83.2 | — | $80 | 1.00M | $4 / $20Direct API | |
37 Gemini 3.5 Flash ↗Google · high | 70.8 | 60.9 | 12.52 sources | 74.6 | 1,478±4 | $33 | 1.05M | $1.5 / $9OpenRouter catalog | |
38 Kimi K3 ↗Moonshot AI · source defaultOpen weights | 77.4 | 60.7 | — | 79.2 | — | $60 | 1.05M | $3 / $15OpenRouter catalog | |
39 Gemini 3.1 Pro Preview ↗Google · highPreview | 73.2 | 60.6 | — | 77.0 | — | $44 | 1.05M | $2 / $12OpenRouter catalog | |
40 GPT-5.4 ↗OpenAI · xhigh | 74.4 | 59.4 | — | 78.0 | — | $55 | 1.05M | $2.5 / $15Direct API | |
41 GPT-5.6 Sol ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 77.1 | 59.3 | — | 81.1 | — | $80 | 1.05M | $4 / $20Direct API | |
42 GPT-5.2 Codex ↗OpenAI · source default | 69.9 | 57.1 | — | 74.0 | — | $45.5 | 400K | $1.75 / $14OpenRouter catalog | |
43 GPT-5.2 ↗OpenAI · high · 2025-12-11 | 69.8 | 57.1 | — | 74.6 | — | $45.5 | 400K | $1.75 / $14OpenRouter catalog | |
44 Claude Opus 5 ↗Anthropic · max | 75.7 | 56.4 | 56.32 sources | 80.1 | 1,487±5 | $100 | 1.00M | $5 / $25Direct API | |
45 Claude Fable 5.1 ↗Anthropic · max | 79.6 | 56.3 | 87.52 sources | 83.4 | 1,498±8 | $200 | 1.00M | $10 / $50Direct API | |
46 Grok 4.3 ↗xAI · source default | 56.0 | 55.9 | — | 62.2 | — | $17.5 | 1.00M | $1.25 / $2.5OpenRouter catalog | |
47 Claude Fable 5 ↗Anthropic · source ID max; display xhighConfiguration discrepancy | 79.0 | 55.8 | — | 83.0 | — | $200 | 1.00M | $10 / $50Direct API | |
48 GPT-6 Astra ↗OpenAI · max | 78.6 | 55.6 | 53.12 sources | 82.2 | 1,480±12 | $200 | 1.05M | $10 / $50Direct API | |
49 Claude Sonnet 4.6 ↗Anthropic · medium · adaptive thinking | 69.4 | 55.1 | — | 73.0 | — | $60 | 1.00M | $3 / $15Direct API | |
50 GPT-5.5 ↗OpenAI · xhigh | 75.8 | 54.8 | — | 80.2 | — | $110 | 1.05M | $5 / $30Direct API | |
51 Claude Opus 4.8 ↗Anthropic · max | 73.0 | 54.5 | — | 76.2 | — | $100 | 1.00M | $5 / $25Direct API | |
52 Claude Opus 4.7 ↗Anthropic · xhigh | 72.9 | 54.4 | — | 76.5 | — | $100 | 1.00M | $5 / $25Direct API | |
53 Claude Opus 4.6 ↗Anthropic · high · adaptive thinking | 70.5 | 52.8 | 56.32 sources | 74.5 | 1,505±4 | $100 | 1.00M | $5 / $25Direct API | |
54 Claude Opus 4.5 ↗Anthropic · high · 64K thinking | 66.7 | 50.1 | — | 72.6 | — | $100 | 200K | $5 / $25Direct API | |
55 Qwen3.8 Flash Next ↗Alibaba · source defaultOpen weights | 76.2 | — | — | 76.2 | — | — | Not verified | Not verified | |
56 Qwen3.8 Max ↗Alibaba · source default · version mapping unconfirmed | 77.0 | — | — | 78.5 | — | — | Not verified | Not verified |
— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.
LOOK BEYOND ONE NUMBER
| Measure | Claude Opus 5.5max | GPT-6 Solmax | DeepSeek V4.1 Flashmax |
|---|---|---|---|
| Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights. | 79.4 | 74.7 | 78.9 |
| Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. | 61.0 | 62.7 | 83.5 |
| Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market. | Insufficient evidence | Insufficient evidence | Insufficient evidence |
| Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums. | $80 | $40 | $2.14 |
| Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits. | 1.00M | 1.05M | 1.05M |
| Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly. | Supported | Supported | Supported |
| ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents. | 92.2 | 88.7 | 86.7 |
| CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects. | 89.3 | 81.8 | 80.0 |
| Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model. | 71.7 | 52.9 | 77.3 |
| MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use. | 97.1 | 96.4 | 93.3 |
| Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information. | 80.3 | 81.2 | 79.3 |
| LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage. | 86.3 | 85.3 | 81.2 |
| Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format. | 65.7 | 68.6 | 70.0 |
| Input / output per 1M | $4 / $20 | $2 / $10 | $0.094 / $0.6 |