Standard text API estimate. Include reasoning tokens in output.
A scenario estimate, not a measured task bill. Excludes cache discounts, tool fees, retries and provider premiums. Workloads exceeding a catalog context limit receive no cost or Value rating. Advertised context is not a guarantee of retrieval quality.
THE MARKET MAP
Capability meets cost.
Sweet-spot zoneEfficient frontierTap a dot to inspect
Shaded zone: 75.0+ points and $21 or less. Numbers mark efficient models in or nearest this zone. This shortlist changes with the metric, workload and filters.
— means missing comparable evidence, never zero. Consensus covers only nine matched configurations. Model IDs, reasoning settings, source links and pricing differences are available on every model page. Open weights does not imply unrestricted commercial use.
There is no universal winner. Compare the task profile, tool support, reasoning configuration and total workload cost. Our Agent Fit score is a task-weighted shortlist; validate it using real tasks and your own tools.
How does AI Agent Store combine model benchmarks?+
Agent Fit uses transparent weights over seven LiveBench categories. Evidence Consensus averages percentile ranks from LiveBench and Arena for nine explicitly matched configurations. Practical Value combines Agent Fit with a cost percentile derived from sourced API prices. Missing evidence is never filled with a guessed score.
Why do some models have no Consensus score?+
The snapshot has no verified matching Arena configuration for those models. Reasoning effort and version matter, so a result for a different setting is not silently reused. A missing score does not mean the model is weak.
Are these AI model rankings live?+
This is a reviewed snapshot dated 23 September 2026, not a live feed. LiveBench uses its 25 June 2026 question set with model observations collected on the snapshot date. Arena and other sources have their own update dates, shown in the methodology.
What does an AI model context window mean?+
It is the advertised token capacity available to a request and its response. A large window lets you supply more information, but does not prove the model will retrieve or reason about every detail reliably.
Does a high benchmark score guarantee a reliable agent?+
No. Reliability also depends on prompts, tool design, permissions, retrieval, retry policies and evaluation. These pages compare models and reported evidence, not complete agent systems.