METR Time Horizon
384.15 minutes
Retrieved August 26, 2026
METRCategory view
This overview uses only published data from matching cohorts. Missing values never change a rank.
The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
82.1% · Rank 1 of 3
Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.86 rating · Rank 2 of 3
Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.81 rating · Rank 2 of 3
Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.83 rating · Rank 2 of 3
Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.94 rating · Rank 3 of 3
Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.97 rating · Rank 2 of 3
Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
76.2% · Rank 2 of 3
Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.82 rating · Rank 2 of 3
Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.76 rating · Rank 1 of 3
Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.78 rating · Rank 3 of 3
Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.94 rating · Rank 3 of 3
Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
1 rating · Rank 1 of 3
Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
77.5% · Rank 1 of 5
Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 verified prompts without tools. Kaggle independently reproduced the results.
46.44% · Rank 1 of 2
Comparison cohort: hle-final-scale-public-temperature-0 · Data date: September 2, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
All 2,500 public questions, temperature 0, and an automatic answer extractor.
85.9% · Rank 3 of 5
Comparison cohort: browsecomp-openai-gpt-5-6-table · Data date: September 2, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
Cross-vendor comparison table with browsing tools. Exact agent systems differ.
Editorial selection from the vendor table, not a complete extract of the comparison cohort.
41.87% · Rank 4 of 4
Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
600 finance tasks with task-specific rubrics and o4-mini as judge.
44.02% · Rank 4 of 4
Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
500 legal tasks with task-specific rubrics and o4-mini as judge.
64.74% · Rank 1 of 2
Task: Overall · Comparison cohort: scale-multinrc-native-languages · Data date: August 26, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
Native, non-translated reasoning tasks with automatic evaluation.
36.78% · Rank 3 of 4
Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Complex multimodal puzzles with exact answer matching and one attempt.
28.97% · Rank 2 of 2
Task: Overall · Comparison cohort: scale-visual-toolbench-apr · Data date: August 26, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
1,204 tasks, six tools, and at most 20 tool calls per task.
Individual values
No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.
Task: Overall · Data date: August 26, 2026
Scale 0 to 100. Not ranked
Multi-turn conversations with binary task-specific rubrics.
Task: Overall · Data date: August 26, 2026
Scale 0 to 100. Not ranked
1,490 tasks with weighted rubrics and Claude 4 Sonnet as judge.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
Head-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.