Low
Tested mode-11.0 pts vs Extra high
$0.1748 / source task
98 s median duration
How much does extra reasoning help? Compare the same model at different effort settings. Spot the useful gains, the expensive plateaus and the times more effort scores lower.
Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.
Analytics in the Hex agent harness · Leaderboard updated: 2026-09-22 · Reviewed September 23, 2026Check the source ↗Lowest measured cost within 1.0 points of the best tested analytics score.
61.3 / 100 · $0.5817 / source taskSet 0 for the highest score. Increase the gap to explore cheaper or lighter options.
● Near-best pickOrange = score decreased
Scores out of 100 · Full 0–100 scale · tap a point to select mode B. Cyan band: within 1.0 points of the best tested score.
Measured average USD per benchmark task.
Within this benchmark and harness. Your task mix, tools, caching and endpoint can change the result.
Noticeable gap · B scored lower
Same benchmark workload
Measured in the source harness
Gap labels are editorial: under 1 point is small, 1–under 5 noticeable, 5+ large. They do not establish statistical significance. A tied score does not prove equivalent behavior.
Extra high → Max-4.3 pts61.3 → 57.0 on analytics score
-11.0 pts vs Extra high
$0.1748 / source task
98 s median duration
-10.0 pts vs Extra high
$0.2708 / source task
146 s median duration
-6.7 pts vs Extra high
$0.3874 / source task
238 s median duration
0.0 pts vs Extra high
$0.5817 / source task
327 s median duration
-4.3 pts vs Extra high
$0.9845 / source task
618 s median duration
Swipe inside the table for all measures. Figures from different study selections are not directly comparable.
| Effort | Analytics score | Cost / task | Median seconds | Output tokens | Exact source ID |
|---|---|---|---|---|---|
| Low | 50.3 | $0.1748 | 97.6 | — | GPT-6 Sol · Low |
| Medium | 51.3 | $0.2708 | 146.2 | — | GPT-6 Sol · Medium |
| High | 54.7 | $0.3874 | 238.0 | — | GPT-6 Sol · High |
| Extra high | 61.3 | $0.5817 | 326.7 | — | GPT-6 Sol · XHigh |
| Max | 57.0 | $0.9845 | 617.5 | — | GPT-6 Sol · Max |
What would A versus B cost if you repeated this benchmark’s average workload?
A source-workload scenario using average measured USD costs. Source prices, caching and tools are already reflected in those observed costs. Your agent’s costs may differ.
Estimate the extra bill from thinking tokens. The starting token counts are editable examples, not measurements of GPT-6 Sol. Rates start from our September 23 model snapshot.
Mode B costs $40 more per month with these assumptions.
Standard text-token arithmetic. Excludes retries, tools, caching, long-context premiums, hosting and taxes. Includes thinking in output billing; do not also include it in “visible output.” Check endpoint limits separately. Changing mode does not automatically predict tokens or quality.
Max is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 3.3 points.
Hex DataBench v1.1Anthropic · 1 sweepMax is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 10.7 points.
Hex DataBench v1.1Anthropic · 2 sweepsHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 11.9 points.
LiveBenchAnthropic · 1 sweepExtra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 8.1 points.
LiveBenchAnthropic · 3 sweepsExtra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 3.7 points.
Hex DataBench v1.1 · LiveBench · OckBenchAnthropic · 2 sweepsMax is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 4.0 points.
Hex DataBench v1.1 · OckBenchAnthropic · 2 sweepsMax is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 16.5 points.
Hex DataBench v1.1 · LiveBenchAnthropic · 1 sweepMedium is the lowest-effort option within 1 point of the best general agent fit in LiveBench. High trails the best tested score by 0.2 points.
LiveBenchAnthropic · 1 sweepMedium is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 2.0 points.
Hex DataBench v1.1DeepSeek · 1 sweepMax is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 62.5 points.
OckBenchDeepSeek · 1 sweepHigh is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 59.0 points.
OckBenchDeepSeek · 1 sweepLow is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 0.3 points.
Hex DataBench v1.1Google · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 22.1 points.
LiveBenchGoogle · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 13.3 points.
LiveBenchGoogle · 1 sweepMedium is the lowest-token option within 1 point of the best overall accuracy in OckBench. The tested score range is 8.0 points.
OckBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 21.3 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 28.7 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 15.1 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 15.3 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 9.8 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best general agent fit in LiveBench. Extra high trails the best tested score by 0.8 points.
LiveBenchOpenAI · 2 sweepsExtra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 7.0 points.
LiveBench · OckBenchOpenAI · 1 sweepExtra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 23.1 points.
LiveBenchOpenAI · 1 sweepExtra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 25.4 points.
LiveBenchOpenAI · 3 sweepsExtra high is the lowest-effort option within 1 point of the best general agent fit in LiveBench. The tested score range is 11.0 points.
LiveBench · OckBench · Stet coding case studyOpenAI · 1 sweepExtra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 16.3 points.
Hex DataBench v1.1OpenAI · 1 sweepMedium is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 10.7 points.
Hex DataBench v1.1OpenAI · 1 sweepExtra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 12.3 points.
Hex DataBench v1.1OpenAI · 1 sweepLow is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 3.0 points.
Hex DataBench v1.1OpenAI · 1 sweepMax is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 22.0 points.
Hex DataBench v1.1OpenAI · 1 sweepExtra high is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 4.3 points.
Hex DataBench v1.1OpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 1.5 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 11.6 points.
LiveBenchOpenAI · 1 sweepHigh is the lowest-effort option within 1 point of the best overall benchmark in LiveBench. The tested score range is 4.3 points.
LiveBenchZ.ai · 2 sweepsHigh is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. Max trails the best tested score by 4.3 points.
Hex DataBench v1.1 · OckBenchZ.ai · 1 sweepMax is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 7.7 points.
Hex DataBench v1.1No. Some published sweeps show a lower score at a higher effort. For example, GPT-6 Sol scores 61.3 at extra high and 57.0 at max in Hex DataBench v1.1. That is a 4.3-point observed decline in this analytics harness. It does not prove that max is worse on every task.
Choose a matching benchmark and task measure first. LiveBench offers coding, research and other task-weighted profiles; OckBench separates math, coding and science; DataBench covers analytics agents. Our near-best pick selects the lowest measured cost, otherwise lowest measured output tokens, otherwise lowest tested effort within your allowed score gap. Validate the shortlist on your own tasks.
Xhigh means extra high reasoning effort. It asks a supported model to invest more effort than high, but it is not a fixed token budget. Available settings and defaults depend on the model, provider and tool. The same label is not a comparable quantity across providers.
It depends on the benchmark, task count and repeatability. We call gaps under 1 point small, 1 to under 5 noticeable, and 5 or more large as a reading aid. Those labels and the default 1-point tolerance are editorial choices, not statistical significance tests. The sources generally do not supply per-mode confidence intervals.
Only explicitly identified, published configurations enter a sweep. Missing modes are not estimated, and unknown defaults are excluded. We show dollar costs or duration only when the same source measured them. Output tokens alone cannot determine the full bill.
Yes. The measured-cost calculator scales the average cost of a benchmark task. The separate reasoning-token calculator uses your own token assumptions and editable input/output prices. A mode label alone cannot predict its token usage, cost or score on your workload.