Gemini 3.5 Flash
- Released
- May 19, 2026
- Data date
- July 30, 2026
Category view
Position within the category
This overview uses only published data from matching cohorts. Missing values never change a rank.
The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.
- Position
- Rank 11 of 26
- Index score
- 58.7 / 100
- Coverage
- 4 / 4
Leaderboard
Measurements
Comparable benchmark results
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
τ-bench Airline Clean
80.8% · Rank 2 of 3
Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
τ-bench Airline Clean
0.86 rating · Rank 2 of 3
Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
τ-bench Airline Clean
0.78 rating · Rank 3 of 3
Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
τ-bench Airline Clean
0.82 rating · Rank 3 of 3
Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
τ-bench Airline Clean
0.97 rating · Rank 1 of 3
Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
τ-bench Airline Clean
0.99 rating · Rank 1 of 3
Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.
GAIA Reliability
79.2% · Rank 1 of 3
Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
GAIA Reliability
0.84 rating · Rank 1 of 3
Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
GAIA Reliability
0.75 rating · Rank 2 of 3
Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
GAIA Reliability
0.8 rating · Rank 1 of 3
Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
GAIA Reliability
0.96 rating · Rank 1 of 3
Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
GAIA Reliability
1 rating · Rank 1 of 3
Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.
SimpleQA Verified v2
70.4% · Rank 2 of 5
Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 verified prompts without tools. Kaggle independently reproduced the results.
SWE-Marathon v1.0
7% · Rank 4 of 4
Task: Binary resolution, pass@1 · Comparison cohort: swe-marathon-v1-0:official:pass1 · Data date: August 9, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
20 multi-hour tasks. The named coding agent is part of every result row.
EnigmaEval
25.41% · Rank 4 of 4
Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
Complex multimodal puzzles with exact answer matching and one attempt.
Profile
Specifications and access
Published information about this model. Unknown values are not estimated.
Pricing
Published prices
Prices remain tied to their documented unit and source.
Head-to-head comparisons
Compare this model
Each matchup compares this model with exactly one other model from the same category.
More models
Models from the same selection
All AI models
Evidence
Primary sources and data date
Every statement links to its underlying documentation or leaderboard.