Low
Near-best pick0.0 pts vs Low
$0.8549 / source task
89 s median duration
Compare Low vs Medium vs High vs Extra high vs Max where tested. Find the effort setting that fits your task, and see how much extra reasoning changes the result.
Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.
Analytics in the Hex agent harness · Leaderboard updated: 2026-09-22 · Reviewed September 23, 2026Check the source ↗Lowest measured cost within 1.0 points of the best tested analytics score.
55.0 / 100 · $0.8549 / source taskSet 0 for the highest score. Increase the gap to explore cheaper or lighter options.
● Near-best pickOrange = score decreased
Scores out of 100 · Full 0–100 scale · tap a point to select mode B. Cyan band: within 1.0 points of the best tested score.
Measured average USD per benchmark task.
Within this benchmark and harness. Your task mix, tools, caching and endpoint can change the result.
Small gap · same observed score
Same benchmark workload
Measured in the source harness
Gap labels are editorial: under 1 point is small, 1–under 5 noticeable, 5+ large. They do not establish statistical significance. A tied score does not prove equivalent behavior.
Low → Medium-2.3 pts55.0 → 52.7 on analytics score
Medium → High-0.7 pts52.7 → 52.0 on analytics score
0.0 pts vs Low
$0.8549 / source task
89 s median duration
-2.3 pts vs Low
$1.24 / source task
128 s median duration
-3.0 pts vs Low
$1.62 / source task
187 s median duration
-2.7 pts vs Low
$1.75 / source task
224 s median duration
0.0 pts vs Low
$2.11 / source task
314 s median duration
Swipe inside the table for all measures. Figures from different study selections are not directly comparable.
| Effort | Analytics score | Cost / task | Median seconds | Output tokens | Exact source ID |
|---|---|---|---|---|---|
| Low | 55.0 | $0.8549 | 88.6 | — | GPT-6 Astra · Low |
| Medium | 52.7 | $1.24 | 128.3 | — | GPT-6 Astra · Medium |
| High | 52.0 | $1.62 | 186.5 | — | GPT-6 Astra · High |
| Extra high | 52.3 | $1.75 | 224.3 | — | GPT-6 Astra · XHigh |
| Max | 55.0 | $2.11 | 313.9 | — | GPT-6 Astra · Max |
What would A versus B cost if you repeated this benchmark’s average workload?
A source-workload scenario using average measured USD costs. Source prices, caching and tools are already reflected in those observed costs. Your agent’s costs may differ.
Estimate the extra bill from thinking tokens. The starting token counts are editable examples, not measurements of GPT-6 Astra. Rates start from our September 23 model snapshot.
Mode B costs $200 more per month with these assumptions.
Standard text-token arithmetic. Excludes retries, tools, caching, long-context premiums, hosting and taxes. Includes thinking in output billing; do not also include it in “visible output.” Check endpoint limits separately. Changing mode does not automatically predict tokens or quality.
Each card is its own comparable sweep. Scores from different benchmarks or question sets measure different work and are not averaged together.
Published blended DataBench score, on a 0–100 scale. Judge-based analytical task performance; not a general intelligence or guaranteed success rate.
Near-best pick: Low at 1-point tolerance. Score spread: 3.0 points.
Explore this sweep ↗Low is the lowest-cost option within 1 point of the best analytics score in Hex DataBench v1.1. The tested score range is 3.0 points. This uses a 1-point practical tolerance in the first listed study. Change the evidence, task measure and tolerance above to get a recommendation for another benchmark. It is a shortlist for testing your own workload.
Low, Medium, High, Extra high, Max appear in the collected studies. Each chart only includes modes tested together in that study. These published labels do not establish that every endpoint or application exposes every setting.
In Hex DataBench v1.1, Low to Medium lowers the observed score by 2.3 points; Medium to High lowers the observed score by 0.7 points. The overall tested range is 3.0 points. These are observed differences, not proof that extra thinking caused the decline.
Use the A/B selector and measured-cost calculator when the source reports dollar costs. For your own agent, enter average input, visible output and thinking tokens in the token calculator. Higher effort is adaptive, so the label itself cannot tell you the bill. Missing costs stay missing.