Claude Fable 5
Anthropic
- Released
- June 9, 2026
- Data date
- July 30, 2026
Category view
Position within the category
This overview uses only published data from matching cohorts. Missing values never change a rank.
The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.
- Position
- Rank 3 of 26
- Index score
- 83.2 / 100
- Coverage
- 4 / 4
Leaderboard
Measurements
Comparable benchmark results
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
SWE-Marathon v1.0
24% · Rank 2 of 4
Task: Binary resolution, pass@1 · Comparison cohort: swe-marathon-v1-0:official:pass1 · Data date: August 9, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
20 multi-hour tasks. The named coding agent is part of every result row.
Two refusals and three unavailable provider-ban trials were counted as zero.
LiveBench
82.97 points · Rank 1 of 4
Comparison cohort: livebench-2026-06-25-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
23 objectively scored tasks across seven categories.
LiveBench
89.7 points · Rank 3 of 4
Task: Reasoning · Comparison cohort: livebench-2026-06-25-reasoning-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Reasoning category inside the same LiveBench release.
LiveBench
86 points · Rank 1 of 4
Task: Coding · Comparison cohort: livebench-2026-06-25-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Coding category inside the same LiveBench release.
LiveBench
62.2 points · Rank 1 of 4
Task: Agentic coding · Comparison cohort: livebench-2026-06-25-agentic-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Agentic coding category inside the same LiveBench release.
LiveBench
96 points · Rank 2 of 4
Task: Mathematics · Comparison cohort: livebench-2026-06-25-mathematics-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Mathematics category inside the same LiveBench release.
LiveBench
80.5 points · Rank 2 of 4
Task: Data analysis · Comparison cohort: livebench-2026-06-25-data-analysis-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Data analysis category inside the same LiveBench release.
LiveBench
90.7 points · Rank 1 of 4
Task: Language · Comparison cohort: livebench-2026-06-25-language-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Language category inside the same LiveBench release.
LiveBench
75.8 points · Rank 1 of 4
Task: Instruction following · Comparison cohort: livebench-2026-06-25-instruction-following-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Instruction following category inside the same LiveBench release.
SWE-Bench Pro
80% · Rank 2 of 5
Comparison cohort: swe-pro-openai-gpt-5-6-release · Data date: September 2, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
Comparison table published by OpenAI. Not a Gradually test.
After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.
PRBench Finance
53.86% · Rank 1 of 4
Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
600 finance tasks with task-specific rubrics and o4-mini as judge.
PRBench Legal
52.56% · Rank 1 of 4
Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
500 legal tasks with task-specific rubrics and o4-mini as judge.
EnigmaEval
39.28% · Rank 1 of 4
Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Complex multimodal puzzles with exact answer matching and one attempt.
Profile
Specifications and access
Published information about this model. Unknown values are not estimated.
- Model type
- ProprietarySource
- Context window
- 1,000,000 tokensSource
- Knowledge cutoff
- January 2026Source
- Notes
- Mythos-class tier above Opus, most capable widely-released model, 1M context (new tokenizer, ~30% more tokens per text vs. pre-Opus-4.7 models), 128K output, adaptive thinking only, safety classifiers with refusal fallback to Opus 4.8. Briefly suspended 2026-06-12 to 2026-06-30 under a US export control directive after a jailbreak was found; generally available again since 2026-07-01.Source
Pricing
Published prices
Prices remain tied to their documented unit and source.
Head-to-head comparisons
Compare this model
Each matchup compares this model with exactly one other model from the same category.
More models
Models from the same selection
All AI models
Evidence
Primary sources and data date
Every statement links to its underlying documentation or leaderboard.
- Model metadataSource
- Context window dataSource
- Knowledge cutoff dataSource
- API pricingSource
- SWE-Marathon (retrieved August 9, 2026)Source
- LiveBench (retrieved September 2, 2026)Source
- OpenAI (retrieved September 2, 2026)Source
- Scale AI (retrieved August 26, 2026)Source
- Scale AI (retrieved August 26, 2026)Source
- Scale AI (retrieved August 26, 2026)Source