ANTHROPIC · PROPRIETARY
Claude Fable 5
Benchmarks, pricing and tradeoffs for agent builders.
Tested configuration: source ID max; display xhigh.
Snapshot · LiveBench question set: 25 June 2026 · View every source ↗
Add to a custom comparison ↗Capability profile
Explain these tests ↗| ReasoningCan it work through constraints and connect clues? LiveBench uses spatial, navigation, perspective-taking and logic-puzzle tasks. Useful for planning, but not a direct test of long-running agents. | 89.7 / 100 |
|---|---|
| CodingCan it write or complete code that passes tests? These are contained programming problems, not entire software projects. | 86.0 / 100 |
| Agentic codingCan it edit code in a tool-using workflow? LiveBench tests JavaScript, TypeScript and Python tasks. Results depend on the benchmark harness as well as the model. | 62.2 / 100 |
| MathematicsCan it solve difficult quantitative problems with checkable answers? Strong math is useful evidence of reasoning, but does not guarantee better writing or tool use. | 96.0 / 100 |
| Data analysisCan it join and reformat tables and reason about event sequences? Useful for agents that process structured business information. | 80.5 / 100 |
| LanguageCan it interpret word relationships, reconstruct plots and correct typos? This is a narrow language test, not a full measure of writing quality or multilingual coverage. | 90.7 / 100 |
| Instruction followingCan it follow requested constraints while rewriting, simplifying, summarizing and composing text? Relevant to agents that must return a specific format. | 75.8 / 100 |
What the evidence says
Claude Fable 5 scores 83.0 on the equal-weight LiveBench category average. Its highest measured category is mathematics (96.0); its lowest is agentic coding (62.2). Categories differ in difficulty, so these are test results, not universal strengths or weaknesses.
Our General agent profile produces an Agent Fit score of 79.0 by emphasizing reasoning, instruction following, agentic coding, coding and data analysis. See the exact weights or adjust the profile in the interactive comparison.
Independent preference evidence
A matching Arena configuration is not verified in this snapshot. Evidence Consensus is withheld. A result for another reasoning setting or version is not substituted.
The LiveBench row ID and display label disagree about reasoning effort. Both labels are preserved above; this model is excluded from Consensus until that mismatch is resolved.
API features to verify
The catalog lists tool-calling support and image input. These are declared features, not evaluations of tool accuracy or visual understanding.
This is a proprietary model accessed through a provider. Availability, rate limits, data handling and service tiers depend on the endpoint and account.
Claude Fable 5 API pricing
At the baseline of 1,000 calls, 10,000 input tokens and 2,000 billed output tokens per call, the text API estimate is $200. Reasoning belongs in billed output. Caching, tools, retries, long-context premiums, service tiers, hosting and taxes are excluded.
Trace this model to the source
- LiveBench row ID
claude-fable-5-max-effort- Question-set release
- 2026-06-25 · Source CSV ↗
- Catalog ID
- anthropic/claude-fable-5 ↗
- Arena configuration
- No verified matching observation
- Reviewed snapshot
- 2026-09-23
Compare alternatives
For production decisions, test representative tasks with a fixed harness and success criteria. Neither a large context window nor a strong benchmark score guarantees reliability in your own agent.