MODEL INTELLIGENCE · SEPTEMBER 2026
AI models, compared.
Choose with confidence.
The model market, made understandable. Explore benchmarks, compare costs, and find the right intelligence for your next AI agent.
Snapshot · LiveBench question set: 25 June 2026 · View every source ↗
Claude Fable 5.1 leads the default Agent Fit mix at 79.6. Change the workload and priorities below to find where capability is worth its cost.
Explore the evidence ↓YOUR WORKLOAD. YOUR PRIORITIES.
Find your model sweet spot.
Make the numbers yours Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums.
Standard text API estimate. Include reasoning tokens in output.
A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.
More capability.
Less budget.
Practical Value Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. · General agent
70% Agent Fit + 30% affordability. Rankings use your workload; filters do not change the reference cohort.
THE EVIDENCE, SIDE BY SIDE
AI model comparison table
| Compare | Our task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights. | Our blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. | Equal-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market. | Automatically checked tasks in seven areas: reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following. Overall is the equal-weight category mean. All rows use the 2026-06-25 question set. | People compare answers without knowing which model wrote them. A higher rating means more preferred answers, not a percent correct. Published ± ranges can overlap; small score gaps may not mean a real difference. | Estimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums. | The advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits. | Input / outputUSD / 1M tokens | |
|---|---|---|---|---|---|---|---|---|---|
01 Claude Fable 5.1 ↗Anthropic · max | 79.6 | 56.3 | 87.52 sources | 83.4 | 1,498±8 | $200 | 1.00M | $10 / $50Direct API | |
02 Claude Opus 5.5 ↗Anthropic · max | 79.4 | 61.0 | — | 83.2 | — | $80 | 1.00M | $4 / $20Direct API | |
03 Muse Spark 1.3 ↗Meta · xhigh | 79.3 | 70.2 | — | 81.6 | — | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
04 Claude Fable 5 ↗Anthropic · source ID max; display xhighConfiguration discrepancy | 79.0 | 55.8 | — | 83.0 | — | $200 | 1.00M | $10 / $50Direct API | |
05 DeepSeek V4.1 Flash ↗DeepSeek · maxOpen weights | 78.9 | 83.5 | — | 81.1 | — | $2.14 | 1.05M | $0.094 / $0.6OpenRouter catalog | |
06 GPT-6 Astra ↗OpenAI · max | 78.6 | 55.6 | 53.12 sources | 82.2 | 1,480±12 | $200 | 1.05M | $10 / $50Direct API | |
07 Kimi K3 ↗Moonshot AI · source defaultOpen weights | 77.4 | 60.7 | — | 79.2 | — | $60 | 1.05M | $3 / $15OpenRouter catalog | |
08 GPT-5.6 Sol ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 77.1 | 59.3 | — | 81.1 | — | $80 | 1.05M | $4 / $20Direct API | |
09 Qwen3.8 Max ↗Alibaba · source default · version mapping unconfirmed | 77.0 | — | — | 78.5 | — | — | Not verified | Not verified | |
10 Muse Spark 1.2 ↗Meta · xhigh | 76.3 | 68.1 | 68.82 sources | 78.0 | 1,500±11 | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
11 Qwen3.8 Flash Next ↗Alibaba · source defaultOpen weights | 76.2 | — | — | 76.2 | — | — | Not verified | Not verified | |
12 Gemini 3.7 Flash ↗Google · high | 76.1 | 71.9 | 56.32 sources | 78.8 | 1,490±8 · preliminary | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
13 GPT-5.5 ↗OpenAI · xhigh | 75.8 | 54.8 | — | 80.2 | — | $110 | 1.05M | $5 / $30Direct API | |
14 Claude Opus 5 ↗Anthropic · max | 75.7 | 56.4 | 56.32 sources | 80.1 | 1,487±5 | $100 | 1.00M | $5 / $25Direct API | |
15 Grok 4.6 ↗xAI · source default | 75.3 | 64.9 | — | 78.0 | — | $32 | 500K | $2 / $6OpenRouter catalog | |
16 DeepSeek V4 Flash Vision Exp ↗DeepSeek · experimental · source defaultOpen weightsPreview | 75.1 | 79.8 | — | 76.8 | — | $3.52 | 1.05M | $0.22 / $0.66OpenRouter catalog | |
17 GPT-6 Sol ↗OpenAI · max | 74.7 | 62.7 | — | 79.2 | — | $40 | 1.05M | $2 / $10Direct API | |
18 GPT-5.4 ↗OpenAI · xhigh | 74.4 | 59.4 | — | 78.0 | — | $55 | 1.05M | $2.5 / $15Direct API | |
19 GPT-5.6 Terra ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 74.1 | 61.2 | — | 77.9 | — | $44 | 1.05M | $2 / $12Direct API | |
20 Muse Spark 1.1 ↗Meta · source ID xhigh; display highConfiguration discrepancy | 74.0 | 66.5 | — | 75.3 | — | $21 | 1.05M | $1.25 / $4.25OpenRouter catalog | |
21 GLM-5.3 ↗Z.ai · source defaultOpen weights | 73.7 | 71.4 | — | 76.1 | — | $13.68 | 1.31M | $0.84 / $2.64OpenRouter catalog | |
22 Grok 4.7 ↗xAI · xhigh | 73.7 | 64.6 | — | 77.4 | — | $25.6 | 500K | $1.6 / $4.8OpenRouter catalog | |
23 Qwen3.8 27B ↗Alibaba · source defaultOpen weights | 73.5 | 74.7 | — | 75.3 | — | $10.2 | 1.00M | $0.42 / $3OpenRouter catalog | |
24 Gemini 3.8 Flash ↗Google · high | 73.3 | 70.0 | 50.02 sources | 75.8 | 1,493±9 · preliminary | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
25 Claude Sonnet 5 ↗Anthropic · xhigh | 73.3 | 61.8 | — | 76.0 | — | $40 | 1.00M | $2 / $10Direct API | |
26 DeepSeek V4 Pro 0813 ↗DeepSeek · source defaultOpen weights | 73.3 | 73.9 | — | 77.4 | — | $10.56 | 1.05M | $0.66 / $1.98OpenRouter catalog | |
27 Gemini 3.1 Pro Preview ↗Google · highPreview | 73.2 | 60.6 | — | 77.0 | — | $44 | 1.05M | $2 / $12OpenRouter catalog | |
28 Grok 4.5 ↗xAI · source default | 73.1 | 63.4 | — | 75.8 | — | $32 | 500K | $2 / $6OpenRouter catalog | |
29 Claude Opus 4.8 ↗Anthropic · max | 73.0 | 54.5 | — | 76.2 | — | $100 | 1.00M | $5 / $25Direct API | |
30 Claude Opus 4.7 ↗Anthropic · xhigh | 72.9 | 54.4 | — | 76.5 | — | $100 | 1.00M | $5 / $25Direct API | |
31 DeepSeek V4 Flash 0731 ↗DeepSeek · source defaultOpen weights | 71.1 | 79.2 | — | 74.2 | — | $1.68 | 1.31M | $0.04 / $0.64OpenRouter catalog | |
32 Gemini 3.5 Flash ↗Google · high | 70.8 | 60.9 | 12.52 sources | 74.6 | 1,478±4 | $33 | 1.05M | $1.5 / $9OpenRouter catalog | |
33 Claude Opus 4.6 ↗Anthropic · high · adaptive thinking | 70.5 | 52.8 | 56.32 sources | 74.5 | 1,505±4 | $100 | 1.00M | $5 / $25Direct API | |
34 Qwen3.7 Max ↗Alibaba · source default | 70.4 | 62.9 | — | 73.1 | — | $23.6 | 1.00M | $1.48 / $4.43OpenRouter catalog | |
35 GPT-5.6 Luna ↗OpenAI · source ID max; display xhighConfiguration discrepancy | 70.4 | 75.9 | — | 73.6 | — | $4.4 | 1.05M | $0.2 / $1.2Direct API | |
36 Gemini 3.6 Flash ↗Google · high | 70.3 | 67.9 | 9.42 sources | 73.6 | 1,480±5 | $15 | 1.05M | $0.75 / $3.75OpenRouter catalog | |
37 GPT-5.2 Codex ↗OpenAI · source default | 69.9 | 57.1 | — | 74.0 | — | $45.5 | 400K | $1.75 / $14OpenRouter catalog | |
38 GPT-5.2 ↗OpenAI · high · 2025-12-11 | 69.8 | 57.1 | — | 74.6 | — | $45.5 | 400K | $1.75 / $14OpenRouter catalog | |
39 Claude Sonnet 4.6 ↗Anthropic · medium · adaptive thinking | 69.4 | 55.1 | — | 73.0 | — | $60 | 1.00M | $3 / $15Direct API | |
40 Inkling ↗Thinking Machines · xhighOpen weights | 68.9 | 64.1 | — | 71.9 | — | $18.1 | 1.05M | $1 / $4.05OpenRouter catalog | |
41 GLM-5.2 ↗Z.ai · source defaultOpen weights | 68.5 | 70.1 | — | 73.2 | — | $10.58 | 1.05M | $0.6496 / $2.04OpenRouter catalog | |
42 GPT-5.4 nano ↗OpenAI · xhigh | 67.7 | 73.4 | — | 69.6 | — | $4.5 | 400K | $0.2 / $1.25Direct API | |
43 GPT-6 Luna ↗OpenAI · max | 67.7 | 76.2 | — | 72.0 | — | $2 | 1.05M | $0.1 / $0.5Direct API | |
44 GLM-5.3 Flash ↗Z.ai · source defaultOpen weights | 67.2 | 74.8 | — | 71.6 | — | $2.5 | 1.31M | $0.15 / $0.5OpenRouter catalog | |
45 DeepSeek V4 Pro ↗DeepSeek · source defaultOpen weights | 67.1 | 67.9 | — | 71.6 | — | $13.37 | 1.05M | $0.9553 / $1.91OpenRouter catalog | |
46 Kimi K2.6 ↗Moonshot AI · thinkingOpen weights | 66.9 | 63.5 | — | 70.5 | — | $17.5 | 262K | $0.95 / $4OpenRouter catalog | |
47 Claude Opus 4.5 ↗Anthropic · high · 64K thinking | 66.7 | 50.1 | — | 72.6 | — | $100 | 200K | $5 / $25Direct API | |
48 Kimi K2.7 Code ↗Moonshot AI · source defaultOpen weights | 64.8 | 65.8 | — | 68.4 | — | $13.66 | 262K | $0.7062 / $3.3OpenRouter catalog | |
49 Qwen3.6 Plus ↗Alibaba · source default | 63.9 | 69.6 | — | 68.9 | — | $7.15 | 1.00M | $0.325 / $1.95OpenRouter catalog | |
50 Nemotron 3 Ultra ↗NVIDIA · 550B A55B · source defaultOpen weights | 63.8 | 66.1 | — | 67.4 | — | $10.8 | 262K | $0.6 / $2.4OpenRouter catalog | |
51 MiniMax M3 ↗MiniMax · source default | 63.1 | 69.6 | — | 67.3 | — | $5.4 | 1.05M | $0.3 / $1.2OpenRouter catalog | |
52 GPT-5.4 mini ↗OpenAI · xhigh | 62.5 | 61.3 | — | 66.4 | — | $16.5 | 400K | $0.75 / $4.5Direct API | |
53 DeepSeek V4 Flash ↗DeepSeek · source defaultOpen weights | 61.6 | 73.1 | — | 65.5 | — | $1.24 | 1.05M | $0.0886 / $0.1772OpenRouter catalog | |
54 Qwen3.6 27B ↗Alibaba · source defaultOpen weights | 60.0 | 65.8 | — | 64.0 | — | $8.6 | 262K | $0.32 / $2.7OpenRouter catalog | |
55 Gemini 3.5 Flash-Lite ↗Google · high | 59.5 | 66.0 | — | 63.9 | — | $8 | 1.05M | $0.3 / $2.5OpenRouter catalog | |
56 Grok 4.3 ↗xAI · source default | 56.0 | 55.9 | — | 62.2 | — | $17.5 | 1.00M | $1.25 / $2.5OpenRouter catalog |
— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.
LOOK BEYOND ONE NUMBER
Your shortlist, under the microscope.
| Measure | Claude Opus 5.5max | GPT-6 Solmax | DeepSeek V4.1 Flashmax |
|---|---|---|---|
| Agent FitOur task-weighted mix of LiveBench category scores. A useful shortlist signal, not a measured agent success rate. Change the task profile to change the weights. | 79.4 | 74.7 | 78.9 |
| Practical ValueOur blend of Agent Fit and affordability: 70% capability + 30% cost percentile by default. Cost uses your token workload. A relative score within this 56-model snapshot, not a claim of dollars saved. | 61.0 | 62.7 | 83.5 |
| Evidence ConsensusEqual-weight average of LiveBench and Arena percentile ranks within the same nine matched configurations. Models without both results receive no score. This small cohort does not rank the whole market. | Insufficient evidence | Insufficient evidence | Insufficient evidence |
| Workload costEstimated text API bill: requests × (input tokens × input price + output tokens × output price) ÷ 1,000,000. Include billed reasoning in output tokens. Excludes caching, tools, retries, taxes, hosting and long-context premiums. | $80 | $40 | $2.14 |
| Context windowThe advertised token capacity for the prompt, conversation and response. This is a size limit, not proof the model can reliably use every detail. Catalog endpoints may have different limits. | 1.00M | 1.05M | 1.05M |
| Tool callingThe catalog declares support for returning structured tool calls. Support does not measure whether the model picks the right tool or uses it correctly. | Supported | Supported | Supported |
| ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents. | 92.2 | 88.7 | 86.7 |
| CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects. | 89.3 | 81.8 | 80.0 |
| Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model. | 71.7 | 52.9 | 77.3 |
| MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use. | 97.1 | 96.4 | 93.3 |
| Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information. | 80.3 | 81.2 | 79.3 |
| LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage. | 86.3 | 85.3 | 81.2 |
| Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format. | 65.7 | 68.6 | 70.0 |
| Input / output per 1M | $4 / $20 | $2 / $10 | $0.094 / $0.6 |
START WITH YOUR QUESTION
Choose a model for the work you do.
Best models for AI agents
Choose an LLM for an AI agent using task-weighted benchmarks, API costs, tool calling and transparent comparison data.
Explore the guide ↗02 / DECISION GUIDEBest models for coding
Compare LLMs for coding agents with code generation, agentic coding benchmarks, reasoning and realistic API cost scenarios.
Explore the guide ↗03 / DECISION GUIDELLM API pricing calculator
Compare AI model API prices per million tokens and estimate monthly agent costs. Adjust input, output, calls and capability-versus-cost priorities.
Explore the guide ↗04 / DECISION GUIDEOpen-weight model comparison
Compare open-weight LLM capabilities and hosted API prices, and learn what to check before self-hosting a model for an agent.
Explore the guide ↗Popular model comparisons
LESS JARGON. BETTER DECISIONS.
AI model questions,
plain-English answers.
Explore the benchmark field guide ↗What is the best AI model for an AI agent?+
There is no universal winner. Compare the task profile, tool support, reasoning configuration and total workload cost. Our Agent Fit score is a task-weighted shortlist; validate it using real tasks and your own tools.
How does AI Agent Store combine model benchmarks?+
Agent Fit uses transparent weights over seven LiveBench categories. Evidence Consensus averages percentile ranks from LiveBench and Arena for nine explicitly matched configurations. Practical Value combines Agent Fit with a cost percentile derived from sourced API prices. Missing evidence is never filled with a guessed score.
Why do some models have no Consensus score?+
The snapshot has no verified matching Arena configuration for those models. Reasoning effort and version matter, so a result for a different setting is not silently reused. A missing score does not mean the model is weak.
Are these AI model rankings live?+
This is a reviewed snapshot dated 23 September 2026, not a live feed. LiveBench uses its 25 June 2026 question set with model observations collected on the snapshot date. Arena and other sources have their own update dates, shown in the methodology.
What does an AI model context window mean?+
It is the advertised token capacity available to a request and its response. A large window lets you supply more information, but does not prove the model will retrieve or reason about every detail reliably.
Does a high benchmark score guarantee a reliable agent?+
No. Reliability also depends on prompts, tool design, permissions, retrieval, retry policies and evaluation. These pages compare models and reported evidence, not complete agent systems.
TURN YOUR SHORTLIST INTO SOMETHING USEFUL