Skip to model comparison
AI Agent Store/Model intelligenceBuild an AI worker ↗

MODEL INTELLIGENCE · SEPTEMBER 2026

AI models, compared.
Choose with confidence.

The model market, made understandable. Explore benchmarks, compare costs, and find the right intelligence for your next AI agent.

Snapshot · LiveBench question set: 25 June 2026 · View every source ↗

Our snapshot take

Claude Fable 5.1 leads the default Agent Fit mix at 79.6. Change the workload and priorities below to find where capability is worth its cost.

Explore the evidence ↓

YOUR WORKLOAD. YOUR PRIORITIES.

Find your model sweet spot.

How our ratings work ↗

Make the numbers yours

Standard text API estimate. Include reasoning tokens in output.

A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.

THE MARKET MAP

Capability meets cost.

Sweet-spot zoneEfficient frontierTap a dot to inspect
SWEET SPOTCAPABILITY FIRST020406080100$1$10$100$1,000Agent Fit / 100 ↑← Lower cost · USD · log scale321SWEET SPOT020406080100$1$10$100$1,000Agent Fit / 100 ↑← Lower cost · USD · log scale321
78.9Agent Fit$2.14your workload
YOUR FAST SHORTLIST

Closest to the sweet spot

How we choose ↗
  1. 1
    DeepSeek V4.1 Flash78.9 Agent Fit · $2.14
  2. 2
    Muse Spark 1.379.3 Agent Fit · $21
  3. 3
    DeepSeek V4 Flash 073171.1 Agent Fit · $1.68

Shaded zone: 75.0+ points and $21 or less. Numbers mark efficient models in or nearest this zone. This shortlist changes with the metric, workload and filters.

THE EVIDENCE, SIDE BY SIDE

AI model comparison table

56 of 56 models · select up to 4 to comparePrices in USD · tap ? for instant explanations
01
Claude Fable 5.1

Anthropic · max

Agent Fit 79.6
Value 56.3
Your cost $200
1.00M contextTool calling
Benchmarks & API pricing +
LiveBench
83.4
Evidence Consensus
87.5
Input / output per 1M tokens
$10 / $50
Price observation
Direct API
02
Claude Opus 5.5

Anthropic · max

Agent Fit 79.4
Value 61.0
Your cost $80
1.00M contextTool calling
Benchmarks & API pricing +
LiveBench
83.2
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$4 / $20
Price observation
Direct API
03
Muse Spark 1.3

Meta · xhigh

Agent Fit 79.3
Value 70.2
Your cost $21
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
81.6
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$1.25 / $4.25
Price observation
OpenRouter catalog
04
Claude Fable 5

Anthropic · source ID max; display xhigh

Agent Fit 79.0
Value 55.8
Your cost $200
1.00M contextTool calling
Benchmarks & API pricing +
LiveBench
83.0
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$10 / $50
Price observation
Direct API
05
Agent Fit 78.9
Value 83.5
Your cost $2.14
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
81.1
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$0.094 / $0.6
Price observation
OpenRouter catalog
06
GPT-6 Astra

OpenAI · max

Agent Fit 78.6
Value 55.6
Your cost $200
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
82.2
Evidence Consensus
53.1
Input / output per 1M tokens
$10 / $50
Price observation
Direct API
07
Kimi K3

Moonshot AI · source default

Agent Fit 77.4
Value 60.7
Your cost $60
1.05M contextOpen weightsTool calling
Benchmarks & API pricing +
LiveBench
79.2
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$3 / $15
Price observation
OpenRouter catalog
08
GPT-5.6 Sol

OpenAI · source ID max; display xhigh

Agent Fit 77.1
Value 59.3
Your cost $80
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
81.1
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
$4 / $20
Price observation
Direct API
09
Qwen3.8 Max

Alibaba · source default · version mapping unconfirmed

Agent Fit 77.0
Value
Your cost
Not verified context
Benchmarks & API pricing +
LiveBench
78.5
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
Not verified
Price observation
Not verified
10
Muse Spark 1.2

Meta · xhigh

Agent Fit 76.3
Value 68.1
Your cost $21
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
78.0
Evidence Consensus
68.8
Input / output per 1M tokens
$1.25 / $4.25
Price observation
OpenRouter catalog
11
Qwen3.8 Flash Next

Alibaba · source default

Agent Fit 76.2
Value
Your cost
Not verified contextOpen weights
Benchmarks & API pricing +
LiveBench
76.2
Evidence Consensus
Not enough evidence
Input / output per 1M tokens
Not verified
Price observation
Not verified
12
Gemini 3.7 Flash

Google · high

Agent Fit 76.1
Value 71.9
Your cost $15
1.05M contextTool calling
Benchmarks & API pricing +
LiveBench
78.8
Evidence Consensus
56.3
Input / output per 1M tokens
$0.75 / $3.75
Price observation
OpenRouter catalog
AI model benchmarks, aggregate ratings, API pricing and context windows. Snapshot 2026-09-23.
CompareInput / outputUSD / 1M tokens
01
Claude Fable 5.1Anthropic · max
79.656.387.52 sources83.41,498±8$2001.00M$10 / $50Direct API
02
Claude Opus 5.5Anthropic · max
79.461.083.2$801.00M$4 / $20Direct API
03
Muse Spark 1.3Meta · xhigh
79.370.281.6$211.05M$1.25 / $4.25OpenRouter catalog
04
Claude Fable 5Anthropic · source ID max; display xhighConfiguration discrepancy
79.055.883.0$2001.00M$10 / $50Direct API
05
DeepSeek V4.1 FlashDeepSeek · maxOpen weights
78.983.581.1$2.141.05M$0.094 / $0.6OpenRouter catalog
06
GPT-6 AstraOpenAI · max
78.655.653.12 sources82.21,480±12$2001.05M$10 / $50Direct API
07
Kimi K3Moonshot AI · source defaultOpen weights
77.460.779.2$601.05M$3 / $15OpenRouter catalog
08
GPT-5.6 SolOpenAI · source ID max; display xhighConfiguration discrepancy
77.159.381.1$801.05M$4 / $20Direct API
09
Qwen3.8 MaxAlibaba · source default · version mapping unconfirmed
77.078.5Not verifiedNot verified
10
Muse Spark 1.2Meta · xhigh
76.368.168.82 sources78.01,500±11$211.05M$1.25 / $4.25OpenRouter catalog
11
Qwen3.8 Flash NextAlibaba · source defaultOpen weights
76.276.2Not verifiedNot verified
12
Gemini 3.7 FlashGoogle · high
76.171.956.32 sources78.81,490±8 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
13
GPT-5.5OpenAI · xhigh
75.854.880.2$1101.05M$5 / $30Direct API
14
Claude Opus 5Anthropic · max
75.756.456.32 sources80.11,487±5$1001.00M$5 / $25Direct API
15
Grok 4.6xAI · source default
75.364.978.0$32500K$2 / $6OpenRouter catalog
16
DeepSeek V4 Flash Vision ExpDeepSeek · experimental · source defaultOpen weightsPreview
75.179.876.8$3.521.05M$0.22 / $0.66OpenRouter catalog
17
GPT-6 SolOpenAI · max
74.762.779.2$401.05M$2 / $10Direct API
18
GPT-5.4OpenAI · xhigh
74.459.478.0$551.05M$2.5 / $15Direct API
19
GPT-5.6 TerraOpenAI · source ID max; display xhighConfiguration discrepancy
74.161.277.9$441.05M$2 / $12Direct API
20
Muse Spark 1.1Meta · source ID xhigh; display highConfiguration discrepancy
74.066.575.3$211.05M$1.25 / $4.25OpenRouter catalog
21
GLM-5.3Z.ai · source defaultOpen weights
73.771.476.1$13.681.31M$0.84 / $2.64OpenRouter catalog
22
Grok 4.7xAI · xhigh
73.764.677.4$25.6500K$1.6 / $4.8OpenRouter catalog
23
Qwen3.8 27BAlibaba · source defaultOpen weights
73.574.775.3$10.21.00M$0.42 / $3OpenRouter catalog
24
Gemini 3.8 FlashGoogle · high
73.370.050.02 sources75.81,493±9 · preliminary$151.05M$0.75 / $3.75OpenRouter catalog
25
Claude Sonnet 5Anthropic · xhigh
73.361.876.0$401.00M$2 / $10Direct API
26
DeepSeek V4 Pro 0813DeepSeek · source defaultOpen weights
73.373.977.4$10.561.05M$0.66 / $1.98OpenRouter catalog
27
Gemini 3.1 Pro PreviewGoogle · highPreview
73.260.677.0$441.05M$2 / $12OpenRouter catalog
28
Grok 4.5xAI · source default
73.163.475.8$32500K$2 / $6OpenRouter catalog
29
Claude Opus 4.8Anthropic · max
73.054.576.2$1001.00M$5 / $25Direct API
30
Claude Opus 4.7Anthropic · xhigh
72.954.476.5$1001.00M$5 / $25Direct API
31
DeepSeek V4 Flash 0731DeepSeek · source defaultOpen weights
71.179.274.2$1.681.31M$0.04 / $0.64OpenRouter catalog
32
Gemini 3.5 FlashGoogle · high
70.860.912.52 sources74.61,478±4$331.05M$1.5 / $9OpenRouter catalog
33
Claude Opus 4.6Anthropic · high · adaptive thinking
70.552.856.32 sources74.51,505±4$1001.00M$5 / $25Direct API
34
Qwen3.7 MaxAlibaba · source default
70.462.973.1$23.61.00M$1.48 / $4.43OpenRouter catalog
35
GPT-5.6 LunaOpenAI · source ID max; display xhighConfiguration discrepancy
70.475.973.6$4.41.05M$0.2 / $1.2Direct API
36
Gemini 3.6 FlashGoogle · high
70.367.99.42 sources73.61,480±5$151.05M$0.75 / $3.75OpenRouter catalog
37
GPT-5.2 CodexOpenAI · source default
69.957.174.0$45.5400K$1.75 / $14OpenRouter catalog
38
GPT-5.2OpenAI · high · 2025-12-11
69.857.174.6$45.5400K$1.75 / $14OpenRouter catalog
39
Claude Sonnet 4.6Anthropic · medium · adaptive thinking
69.455.173.0$601.00M$3 / $15Direct API
40
InklingThinking Machines · xhighOpen weights
68.964.171.9$18.11.05M$1 / $4.05OpenRouter catalog
41
GLM-5.2Z.ai · source defaultOpen weights
68.570.173.2$10.581.05M$0.6496 / $2.04OpenRouter catalog
42
GPT-5.4 nanoOpenAI · xhigh
67.773.469.6$4.5400K$0.2 / $1.25Direct API
43
GPT-6 LunaOpenAI · max
67.776.272.0$21.05M$0.1 / $0.5Direct API
44
GLM-5.3 FlashZ.ai · source defaultOpen weights
67.274.871.6$2.51.31M$0.15 / $0.5OpenRouter catalog
45
DeepSeek V4 ProDeepSeek · source defaultOpen weights
67.167.971.6$13.371.05M$0.9553 / $1.91OpenRouter catalog
46
Kimi K2.6Moonshot AI · thinkingOpen weights
66.963.570.5$17.5262K$0.95 / $4OpenRouter catalog
47
Claude Opus 4.5Anthropic · high · 64K thinking
66.750.172.6$100200K$5 / $25Direct API
48
Kimi K2.7 CodeMoonshot AI · source defaultOpen weights
64.865.868.4$13.66262K$0.7062 / $3.3OpenRouter catalog
49
Qwen3.6 PlusAlibaba · source default
63.969.668.9$7.151.00M$0.325 / $1.95OpenRouter catalog
50
Nemotron 3 UltraNVIDIA · 550B A55B · source defaultOpen weights
63.866.167.4$10.8262K$0.6 / $2.4OpenRouter catalog
51
MiniMax M3MiniMax · source default
63.169.667.3$5.41.05M$0.3 / $1.2OpenRouter catalog
52
GPT-5.4 miniOpenAI · xhigh
62.561.366.4$16.5400K$0.75 / $4.5Direct API
53
DeepSeek V4 FlashDeepSeek · source defaultOpen weights
61.673.165.5$1.241.05M$0.0886 / $0.1772OpenRouter catalog
54
Qwen3.6 27BAlibaba · source defaultOpen weights
60.065.864.0$8.6262K$0.32 / $2.7OpenRouter catalog
59.566.063.9$81.05M$0.3 / $2.5OpenRouter catalog
56
Grok 4.3xAI · source default
56.055.962.2$17.51.00M$1.25 / $2.5OpenRouter catalog

— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.

LOOK BEYOND ONE NUMBER

Your shortlist, under the microscope.

Change models ↑
ReasoningCodingAgentic codeMathematicsDataLanguageInstructions
Claude Opus 5.5GPT-6 SolDeepSeek V4.1 Flash
Selected AI models compared, with scores from the same LiveBench question set
MeasureClaude Opus 5.5maxGPT-6 SolmaxDeepSeek V4.1 Flashmax
Agent Fit79.474.778.9
Practical Value61.062.783.5
Evidence ConsensusInsufficient evidenceInsufficient evidenceInsufficient evidence
Workload cost$80$40$2.14
Context window1.00M1.05M1.05M
Tool callingSupportedSupportedSupported
Reasoning92.288.786.7
Coding89.381.880.0
Agentic coding71.752.977.3
Mathematics97.196.493.3
Data analysis80.381.279.3
Language86.385.381.2
Instruction following65.768.670.0
Input / output per 1M$4 / $20$2 / $10$0.094 / $0.6

START WITH YOUR QUESTION

Choose a model for the work you do.

Popular model comparisons

LESS JARGON. BETTER DECISIONS.

AI model questions,
plain-English answers.

Explore the benchmark field guide ↗
What is the best AI model for an AI agent?+

There is no universal winner. Compare the task profile, tool support, reasoning configuration and total workload cost. Our Agent Fit score is a task-weighted shortlist; validate it using real tasks and your own tools.

How does AI Agent Store combine model benchmarks?+

Agent Fit uses transparent weights over seven LiveBench categories. Evidence Consensus averages percentile ranks from LiveBench and Arena for nine explicitly matched configurations. Practical Value combines Agent Fit with a cost percentile derived from sourced API prices. Missing evidence is never filled with a guessed score.

Why do some models have no Consensus score?+

The snapshot has no verified matching Arena configuration for those models. Reasoning effort and version matter, so a result for a different setting is not silently reused. A missing score does not mean the model is weak.

Are these AI model rankings live?+

This is a reviewed snapshot dated 23 September 2026, not a live feed. LiveBench uses its 25 June 2026 question set with model observations collected on the snapshot date. Arena and other sources have their own update dates, shown in the methodology.

What does an AI model context window mean?+

It is the advertised token capacity available to a request and its response. A large window lets you supply more information, but does not prove the model will retrieve or reason about every detail reliably.

Does a high benchmark score guarantee a reliable agent?+

No. Reliability also depends on prompts, tool design, permissions, retrieval, retry policies and evaluation. These pages compare models and reported evidence, not complete agent systems.

TURN YOUR SHORTLIST INTO SOMETHING USEFUL

A model is the beginning.
Give it a job to do.

Build an AI worker ↗
AI Agent Store · Model intelligence · Snapshot 2026-09-23Sources, limitations & corrections ↗