TUA-Bench
64.7%
Retrieved August 9, 2026
TUA-BenchOpenAI
Category view
This overview uses only published data from matching cohorts. Missing values never change a rank.
The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
79.5% · Rank 3 of 3
Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.89 rating · Rank 1 of 3
Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.84 rating · Rank 1 of 3
Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.84 rating · Rank 1 of 3
Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.97 rating · Rank 1 of 3
Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
0.96 rating · Rank 3 of 3
Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
62.8% · Rank 3 of 3
Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.79 rating · Rank 3 of 3
Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.61 rating · Rank 3 of 3
Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.8 rating · Rank 1 of 3
Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
0.95 rating · Rank 2 of 3
Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
1 rating · Rank 1 of 3
Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
95.43% · Rank 1 of 5
Task: Accepted cases · Comparison cohort: almanbench-v0-1:official:overall · Data date: August 9, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,029 public tasks with one direct run per row. The target is the Alman language specification.
12% · Rank 3 of 4
Task: Binary resolution, pass@1 · Comparison cohort: swe-marathon-v1-0:official:pass1 · Data date: August 9, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
20 multi-hour tasks. The named coding agent is part of every result row.
80.19 points · Rank 3 of 4
Comparison cohort: livebench-2026-06-25-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
23 objectively scored tasks across seven categories.
89.7 points · Rank 3 of 4
Task: Reasoning · Comparison cohort: livebench-2026-06-25-reasoning-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Reasoning category inside the same LiveBench release.
82.1 points · Rank 3 of 4
Task: Coding · Comparison cohort: livebench-2026-06-25-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Coding category inside the same LiveBench release.
54 points · Rank 4 of 4
Task: Agentic coding · Comparison cohort: livebench-2026-06-25-agentic-coding-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Agentic coding category inside the same LiveBench release.
95.9 points · Rank 3 of 4
Task: Mathematics · Comparison cohort: livebench-2026-06-25-mathematics-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Mathematics category inside the same LiveBench release.
81.6 points · Rank 1 of 4
Task: Data analysis · Comparison cohort: livebench-2026-06-25-data-analysis-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Data analysis category inside the same LiveBench release.
87.4 points · Rank 3 of 4
Task: Language · Comparison cohort: livebench-2026-06-25-language-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Language category inside the same LiveBench release.
70.7 points · Rank 3 of 4
Task: Instruction following · Comparison cohort: livebench-2026-06-25-instruction-following-max · Data date: September 2, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Instruction following category inside the same LiveBench release.
84.4% · Rank 4 of 5
Comparison cohort: browsecomp-openai-gpt-5-6-table · Data date: September 2, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
Cross-vendor comparison table with browsing tools. Exact agent systems differ.
Editorial selection from the vendor table, not a complete extract of the comparison cohort.
Individual values
No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.
Data date: August 26, 2026
Scale 0 to 100. Not ranked
Complete agent system with up to 500 steps.
Claude used about 244,000 output tokens. This is a system result, not an isolated model score.
Data date: September 2, 2026
Scale 0 to 100. Not ranked
Comparison table published by Anthropic. Not a Gradually test.
After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
64.7%
Retrieved August 9, 2026
TUA-Bench57.7%
Retrieved August 9, 2026
TUA-Bench68.3%
Retrieved August 9, 2026
TUA-Bench42.5%
Retrieved August 9, 2026
TUA-Bench64.2%
Retrieved August 9, 2026
TUA-Bench57.2%
Retrieved August 9, 2026
TUA-Bench66.7%
Retrieved August 9, 2026
TUA-Bench46.7%
Retrieved August 9, 2026
TUA-Bench62.4%
Retrieved August 9, 2026
TUA-Bench54.2%
Retrieved August 9, 2026
TUA-Bench67.5%
Retrieved August 9, 2026
TUA-Bench40%
Retrieved August 9, 2026
TUA-BenchHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.