Claude Opus 4.8
Not tested yet
Anthropic
The index averages rank percentiles from 4 documented comparison cohorts.
Entries 1-12 of 15. Page 1 of 2.
Each chart compares models in the same documented test and highlights the current model.
By task type
Overall
Note (Claude Opus 4.8): Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Scale: 0% - 100%
Overall
What this benchmark measures: 1,161 complex, often multimodal puzzles test creative problem solving beyond well-defined academic tasks.
Context: Puzzle solving is a narrow specialty and does not replace professional or everyday task evaluation.
Scale: 0% - 100%
Overall
What this benchmark measures: Computer tasks in realistic desktop environments.
Context: The result belongs to the complete agent system and is not an isolated model score.
Note (Claude Opus 4.8): Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Anthropic's release-page footnote reports 82.3% after an updated run. The associated system-card table still contains 82.8%.
Scale: 0% - 100%
Overall
What this benchmark measures: 108 long computer workflows with up to 500 steps.
Context: It evaluates a complete system. Binary and partial success rates stay separate.
Note (Claude Opus 4.8): Claude used about 244,000 output tokens. This is a system result, not an isolated model score.
Note: Claude used about 244,000 output tokens. This is a system result, not an isolated model score.
Scale: 0% - 100%
Overall
What this benchmark measures: 1,266 hard-to-find research tasks with short, unambiguous answers.
Context: The benchmark rewards persistent fact finding and only partly represents open-ended research reports or typical user questions.
Note (Claude Opus 4.8): Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Note: Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Scale: 0% - 100%
Binary resolution, pass@1
What this benchmark measures: 20 multi-hour repository tasks test very long software work with binary end-to-end resolution.
Context: Scores belong to the v1.0 agent setup. No v1.1 leaderboard was posted at the data cutoff.
Note: Two refusals and three unavailable provider-ban trials were counted as zero.
Scale: 0% - 100%
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Every observation retains its source value and published test conditions.
Our model tests
Compare how the models respond to the same prompt. Each test shows the first attempt, with no subsequent fixes to the generated code. These results do not contribute to an overall score.
A community classic for free-form SVG drawing.
Claude Opus 4.8
Not tested yet
Generate an SVG of a pelican riding a bicycle
Token limit including reasoning: 8,192 tokens
Task origin (Simon Willison)Profile
Published information about this model. Existing estimates are explicitly labeled.
Pricing
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
Cost example
$1.75
100 requests with 1,000 input and 500 output tokens each calculate to $1.75. Input accounts for $0.5; output accounts for $1.25.
The calculation uses documented token prices and no cache discount. It excludes extra tools, tax, reasoning tokens, and further hidden output tokens.
Open sourceHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
Evidence
Every statement links to its underlying documentation or leaderboard.
Base API prices without caching or batch discounts. Higher context tiers and other rates are listed under Costs.