METR Time Horizon
341.74 minutes
Retrieved August 26, 2026
METROpenAI
Category view
This overview uses only published data from matching cohorts. Missing values never change a rank.
This model does not yet have enough published values for an overall rank. It covers 1 of 4 required comparison cohorts.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
36.24% · Rank 2 of 2
Comparison cohort: hle-final-scale-public-temperature-0 · Data date: September 2, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
All 2,500 public questions, temperature 0, and an automatic answer extractor.
45.63% · Rank 3 of 4
Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
600 finance tasks with task-specific rubrics and o4-mini as judge.
44.35% · Rank 3 of 4
Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
500 legal tasks with task-specific rubrics and o4-mini as judge.
58.29% · Rank 2 of 2
Task: Overall · Comparison cohort: scale-multinrc-native-languages · Data date: August 26, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
Native, non-translated reasoning tasks with automatic evaluation.
29.17% · Rank 1 of 2
Task: Overall · Comparison cohort: scale-visual-toolbench-apr · Data date: August 26, 2026
2 of 2 model versions shown in this chart. A higher value ranks first.
1,204 tasks, six tools, and at most 20 tool calls per task.
Individual values
No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.
Task: Weighted total · Data date: August 9, 2026
Scale 0 to 100. Not ranked
98 project cases. Functional tests run before dynamic exploit checks and security scoring.
Task: Repair without hints · Data date: August 9, 2026
Scale 0 to 100. Not ranked
98 project cases. Functional tests run before dynamic exploit checks and security scoring.
Task: Repair with security hints · Data date: August 9, 2026
Scale 0 to 100. Not ranked
98 project cases. Functional tests run before dynamic exploit checks and security scoring.
Task: Generation without hints · Data date: August 9, 2026
Scale 0 to 100. Not ranked
98 project cases. Functional tests run before dynamic exploit checks and security scoring.
Task: Generation with security hints · Data date: August 9, 2026
Scale 0 to 100. Not ranked
98 project cases. Functional tests run before dynamic exploit checks and security scoring.
Task: Overall · Data date: August 26, 2026
Scale 0 to 100. Not ranked
758 single-turn image-text tasks with structured yes-no rubrics.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
341.74 minutes
Retrieved August 26, 2026
METR53.88 minutes
Retrieved August 26, 2026
METR68.8%
Retrieved August 9, 2026
EnterpriseRAG-Bench55.95%
Retrieved August 9, 2026
EnterpriseRAG-Bench68.41%
Retrieved August 9, 2026
EnterpriseRAG-Bench9.01 points
Retrieved August 9, 2026
EnterpriseRAG-Bench51.4%
Retrieved August 9, 2026
EnterpriseRAG-Bench42.94%
Retrieved August 9, 2026
EnterpriseRAG-Bench46.03%
Retrieved August 9, 2026
EnterpriseRAG-Bench9.32 points
Retrieved August 9, 2026
EnterpriseRAG-Bench60.6%
Retrieved August 9, 2026
EnterpriseRAG-Bench61.12%
Retrieved August 9, 2026
EnterpriseRAG-Bench55.76%
Retrieved August 9, 2026
EnterpriseRAG-Bench2 points
Retrieved August 9, 2026
EnterpriseRAG-Bench25%
Retrieved August 9, 2026
SWE-EVO33.89%
Retrieved August 9, 2026
SWE-EVO25%
Retrieved August 9, 2026
SWE-EVO33.98%
Retrieved August 9, 2026
SWE-EVOHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.