Skip to main content
Language modelActive

Gemini 3.5 Flash

Google

Released
May 19, 2026
Data date
July 30, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.

Position
Rank 11 of 26
Index score
58.7 / 100
Coverage
4 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
10GPT-5.6 Terra59.1 / 100
11Gemini 3.5 Flash, Current model58.7 / 100
12Claude Sonnet 557.7 / 100
25MiMo-V2.5-Pro12.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

τ-bench Airline Clean

80.8% · Rank 2 of 3

Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026

1.Gemini 3.1 Pro Preview82.1%
2.Gemini 3.5 Flash, Current model80.8%
3.GPT-5.579.5%

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.86 rating · Rank 2 of 3

Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026

1.GPT-5.50.89 rating
2.Gemini 3.1 Pro Preview0.86 rating
2.Gemini 3.5 Flash, Current model0.86 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.78 rating · Rank 3 of 3

Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026

1.GPT-5.50.84 rating
2.Gemini 3.1 Pro Preview0.81 rating
3.Gemini 3.5 Flash, Current model0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.82 rating · Rank 3 of 3

Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026

1.GPT-5.50.84 rating
2.Gemini 3.1 Pro Preview0.83 rating
3.Gemini 3.5 Flash, Current model0.82 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.97 rating · Rank 1 of 3

Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026

1.Gemini 3.5 Flash, Current model0.97 rating
1.GPT-5.50.97 rating
3.Gemini 3.1 Pro Preview0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.99 rating · Rank 1 of 3

Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026

1.Gemini 3.5 Flash, Current model0.99 rating
2.Gemini 3.1 Pro Preview0.97 rating
3.GPT-5.50.96 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

GAIA Reliability

79.2% · Rank 1 of 3

Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026

1.Gemini 3.5 Flash, Current model79.2%
2.Gemini 3.1 Pro Preview76.2%
3.GPT-5.562.8%

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.84 rating · Rank 1 of 3

Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026

1.Gemini 3.5 Flash, Current model0.84 rating
2.Gemini 3.1 Pro Preview0.82 rating
3.GPT-5.50.79 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.75 rating · Rank 2 of 3

Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview0.76 rating
2.Gemini 3.5 Flash, Current model0.75 rating
3.GPT-5.50.61 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.8 rating · Rank 1 of 3

Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026

1.Gemini 3.5 Flash, Current model0.8 rating
1.GPT-5.50.8 rating
3.Gemini 3.1 Pro Preview0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.96 rating · Rank 1 of 3

Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026

1.Gemini 3.5 Flash, Current model0.96 rating
2.GPT-5.50.95 rating
3.Gemini 3.1 Pro Preview0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

1 rating · Rank 1 of 3

Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview1 rating
1.Gemini 3.5 Flash, Current model1 rating
1.GPT-5.51 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

SimpleQA Verified v2

70.4% · Rank 2 of 5

Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026

1.Gemini 3.1 Pro Preview77.5%
2.Gemini 3.5 Flash, Current model70.4%
3.GPT-5.6 Sol69.2%
4.Claude Opus 563.3%
5.Grok 4.553.8%

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 verified prompts without tools. Kaggle independently reproduced the results.

Source: KaggleParticipants: 5

SWE-Marathon v1.0

7% · Rank 4 of 4

Task: Binary resolution, pass@1 · Comparison cohort: swe-marathon-v1-0:official:pass1 · Data date: August 9, 2026

1.Grok 4.529%
2.Claude Fable 524%
3.GPT-5.512%
4.Gemini 3.5 Flash, Current model7%

4 of 4 model versions shown in this chart. A higher value ranks first.

20 multi-hour tasks. The named coding agent is part of every result row.

Source: SWE-MarathonParticipants: 4

EnigmaEval

25.41% · Rank 4 of 4

Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026

1.Claude Fable 539.28%
2.GPT-5.6 Sol37.12%
3.Gemini 3.1 Pro Preview36.78%
4.Gemini 3.5 Flash, Current model25.41%

4 of 4 model versions shown in this chart. A higher value ranks first.

Complex multimodal puzzles with exact answer matching and one attempt.

Source: Scale AIParticipants: 4

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,048,576 tokensSource
Notes
Released May 19, 2026. Google does not publish its parameter count, architecture, or knowledge cutoff.Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$1.5 per 1M tokensSource
API output
$9 per 1M tokensSource
Cache read
$0.15 per 1M tokensSource

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • Model metadataSource
  • API pricingSource
  • Princeton HAL (retrieved August 5, 2026)Source
  • Princeton HAL (retrieved August 6, 2026)Source
  • Kaggle (retrieved August 9, 2026)Source
  • SWE-Marathon (retrieved August 9, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source