ARC-AGI-3
7.78%
Retrieved August 26, 2026
ARC PrizeOpenAI
Category view
This overview uses only published data from matching cohorts. Missing values never change a rank.
The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
69.2% · Rank 3 of 5
Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 verified prompts without tools. Kaggle independently reproduced the results.
95.14% · Rank 2 of 5
Task: Accepted cases · Comparison cohort: almanbench-v0-1:official:overall · Data date: August 9, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,029 public tasks with one direct run per row. The target is the Alman language specification.
81.05 points · Rank 2 of 4
Comparison cohort: livebench-2026-06-25-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
23 objectively scored tasks across seven categories.
91.7 points · Rank 1 of 4
Task: Reasoning · Comparison cohort: livebench-2026-06-25-reasoning-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Reasoning category inside the same LiveBench release.
83.9 points · Rank 2 of 4
Task: Coding · Comparison cohort: livebench-2026-06-25-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Coding category inside the same LiveBench release.
56.2 points · Rank 2 of 4
Task: Agentic coding · Comparison cohort: livebench-2026-06-25-agentic-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Agentic coding category inside the same LiveBench release.
96.2 points · Rank 1 of 4
Task: Mathematics · Comparison cohort: livebench-2026-06-25-mathematics-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Mathematics category inside the same LiveBench release.
79.8 points · Rank 3 of 4
Task: Data analysis · Comparison cohort: livebench-2026-06-25-data-analysis-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Data analysis category inside the same LiveBench release.
87.7 points · Rank 2 of 4
Task: Language · Comparison cohort: livebench-2026-06-25-language-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Language category inside the same LiveBench release.
71.8 points · Rank 2 of 4
Task: Instruction following · Comparison cohort: livebench-2026-06-25-instruction-following-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Instruction following category inside the same LiveBench release.
64.6% · Rank 3 of 5
Comparison cohort: swe-pro-openai-gpt-5-6-release · Data date: September 2, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
Comparison table published by OpenAI. Not a Gradually test.
After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.
50.45% · Rank 2 of 4
Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
600 finance tasks with task-specific rubrics and o4-mini as judge.
50.5% · Rank 2 of 4
Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
500 legal tasks with task-specific rubrics and o4-mini as judge.
37.12% · Rank 2 of 4
Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Complex multimodal puzzles with exact answer matching and one attempt.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
7.78%
Retrieved August 26, 2026
ARC Prize6.99%
Retrieved August 26, 2026
ARC Prize2.15%
Retrieved August 26, 2026
ARC Prize1.07%
Retrieved August 26, 2026
ARC Prize0.33%
Retrieved August 26, 2026
ARC Prize92.5%
Retrieved September 2, 2026
ARC Prize90%
Retrieved September 2, 2026
ARC Prize85.4%
Retrieved September 2, 2026
ARC Prize67.1%
Retrieved September 2, 2026
ARC Prize42.5%
Retrieved September 2, 2026
ARC Prize92.2%
Retrieved September 2, 2026 · Editorial selection from the vendor table, not a complete extract of the comparison cohort.
OpenAI90.4%
Retrieved September 2, 2026 · Editorial selection from the vendor table, not a complete extract of the comparison cohort.
OpenAIHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.