What makes an effort setting
worth choosing?
We bring published effort sweeps into one comparison tool, then calculate practical recommendations within each sweep. Every recommendation has a task, a source and an adjustable tolerance.
Reviewed September 23, 2026 · 36 model families · 45 sweeps · 156 configurations
1. Compare the same model on the same work
A sweep contains one model identity tested at two or more explicitly named effort levels in the same benchmark version or question set. We preserve exact source IDs. We do not interpolate missing levels, guess unknown default settings or substitute another model release.
The source defines the harness, prompts, sampling and effort configuration. A question-set date identifies a benchmark release, not the model’s launch date or the date it was evaluated. Older LiveBench results remain labeled with their older question sets.
Separate studies stay separate. A 70 in analytics and a 70 in a coding benchmark do not mean the same thing. Our multi-source view combines access to evidence, not incompatible scores into one average.
2. Find the near-best choice
First select a source and task measure. The highest observed score becomes the reference. Keep modes whose score is within your chosen tolerance of that reference; the default is 1 point on a 0–100 scale.
eligible = best observed score − mode score ≤ tolerance
Among eligible modes, we choose the lowest measured dollar cost if every mode has a comparable cost. Otherwise, we use the lowest measured output-token count if every mode has it, using the selected domain’s token count where available. If neither resource is measured, we choose the lowest tested effort. Resource ties favor the higher score, then lower effort.
A near-best pick is a practical starting point, not a universal optimum. Duration is shown as another tradeoff and is not part of this selection rule. “Another mode trades better” means another tested mode scores at least as high with no greater use of the selected resource, with at least one strict improvement.
Example: on DataBench v1.1, GPT-6 Sol scores 61.3 at extra high and 57.0 at max. Extra high is the best observed score and costs less. At a 1-point tolerance it is the near-best pick. Change the tolerance and inspect the actual costs →
3. Show the size of the difference
The A/B score difference is B minus A in percentage points of the displayed 0–100 measure. A 60 to 65 change is +5 points, not a promised 5% productivity gain. Our reading labels are: under 1 point small, 1 to under 5 noticeable, and 5 or more large.
These labels are editorial. They are not confidence intervals or significance tests. Tied or close scores do not establish equal behavior. The source studies generally do not report per-mode uncertainty, so we cannot claim an observed gap will repeat.
An orange chart segment flags an adjacent increase in effort that scored lower in this sweep. Possible reasons include task fit, sampling variation, harness behavior and failures. We do not infer “overthinking” as the cause from a score alone.
Charts default to the full 0–100 score axis. The optional zoom explicitly changes the displayed range to make small differences visible. The near-best shaded region is your chosen practical tolerance.
4. Match the measure to the task
DataBench measures agentic analytics. OckBench separates mathematics, coding and science. The Stet case study separates test passing, equivalence to human patches, AI review and their intersection. Changing the selected measure can change the preferred effort.
For LiveBench, we average task scores inside each published category. The overall score is the equal-weight category mean. Our General agent, Coding assistant, Research & analysis and Writing & support profiles apply the same transparent category weights as Agent Fit. A profile is withheld if its required categories are absent from an older question set; missing categories are never treated as zero.
5. Two calculators, two kinds of evidence
Scale a measured benchmark cost
scenario cost = number of tasks × published average cost per task
This repeats the source’s average workload at the source’s observed cost. It is not a price quote for another agent or a prediction of successful task count. We keep measured duration separate and do not turn it into a throughput promise.
Estimate your own reasoning-token bill
cost = requests × [input tokens × input rate
+ (visible output + reasoning tokens) × output rate] / 1,000,000
Rates are dollars per million tokens. Where a model has a matched market profile, rates start from our September 23 snapshot; otherwise, you must enter them. Token counts start as clearly labeled editable examples. Effort labels never automatically generate a token estimate.
This arithmetic assumes standard text-token billing with thinking charged as output. It excludes caching, tool fees, retries, context-dependent premiums, hosting and taxes. Check your endpoint’s billing and limits. Empty, negative or invalid values withhold the estimate; a valid zero remains zero.
6. Four sources, with their limits attached
Hex DataBench v1.1
Analytics agents working in a synthetic business warehouse. The published blended score reflects task rubrics, with API cost and median duration measured in the Hex harness.
Interpretation limits: The v1 design contains 100 tasks. An LLM judge evaluates the work. Version 1.1 changed judge calibration on September 17; older v1 scores are not mixed in. No confidence intervals are published. These costs and times describe this harness, not every agent.
LiveBench
Checkable benchmark tasks grouped into reasoning, coding, agentic coding, mathematics, data, language and instruction following. Category means and task-weighted Agent Fit are calculated from the published task scores.
Interpretation limits: Each sweep uses one question set and one model identity. The question-set date is not a model launch or evaluation date. Unknown default efforts are excluded. Older sweeps do not establish the behavior of newer model versions. No per-mode cost or confidence interval is inferred.
OckBench
200 selected problems: 100 mathematics, 60 coding and 40 science. Accuracy and average output tokens expose quality versus token use; domain results show where effort helps.
Interpretation limits: Mathematics uses an LLM judge; code runs against tests and science checks answers. Sampling follows model-specific configurations. Opus 5 medium includes nine persistent zero-output failures, counted as incorrect but omitted from average output tokens. Token counts are not dollars or latency.
Stet coding case study
GPT-5.5 in Codex on 26 matched GraphQL-go-tools repository tasks. Test outcomes, equivalence to human changes and AI review outcomes answer different coding-quality questions.
Interpretation limits: Exploratory evidence: one repository, one seed per task, stitched runs and uncalibrated GPT-5.4 judging. The publisher calls this inspect-grade, not decision-grade evidence. Xhigh costs retain pre-regrade records. Do not generalize small gaps.
We report publisher measurements and derived comparisons; we have not independently rerun these evaluations. The downloadable snapshot preserves the source IDs, study context, retrieval date and hashes of the input files so the calculations can be inspected.
7. What the provider’s effort control means
Effort is usually an adaptive instruction to spend more or less computation, rather than a fixed number of tokens. Supported values and defaults are model-specific. Published benchmark labels may also refer to a harness setting or an explicit thinking-token budget.
| Provider | Control | What to check |
|---|---|---|
| OpenAI | reasoning.effort | Supported levels vary by model; none, minimal, low, medium, high, xhigh and max are not available universally. Official reasoning guide ↗ |
| Anthropic | output_config.effort | Effort controls output investment and adaptive thinking. Defaults and support vary by model. Older thinking budgets are a different control. Official effort guide ↗ |
thinking_level / thinking_budget | Use the control supported by the model. A level and a numeric thinking budget are not interchangeable. Official thinking guide ↗ |
8. Use the result as a testable shortlist
Pick a task-relevant sweep, choose your tolerance, then compare the recommended setting with the best-scoring mode on examples from your own agent. Record success, tool use, cost and duration with a consistent harness. Prefer a setting only when the tradeoff fits your workload.
This is a dated snapshot, not a live monitor. New provider versions, benchmark revisions and prices can change the result. We refresh evidence by preserving source versions, rechecking model identities and recomputing the same formulas. Send a source correction →