Skip to main content
Language modelPreview

Gemini 3.1 Pro Preview

Google

Released
February 19, 2026
Data date
July 30, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.

Position
Rank 16 of 26
Index score
41.3 / 100
Coverage
4 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
15Grok 4.546.2 / 100
16Gemini 3.1 Pro Preview, Current model41.3 / 100
17GLM-5.3-Flash35.6 / 100
25MiMo-V2.5-Pro12.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

τ-bench Airline Clean

82.1% · Rank 1 of 3

Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026

1.Gemini 3.1 Pro Preview, Current model82.1%
2.Gemini 3.5 Flash80.8%
3.GPT-5.579.5%

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.86 rating · Rank 2 of 3

Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026

1.GPT-5.50.89 rating
2.Gemini 3.1 Pro Preview, Current model0.86 rating
2.Gemini 3.5 Flash0.86 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.81 rating · Rank 2 of 3

Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026

1.GPT-5.50.84 rating
2.Gemini 3.1 Pro Preview, Current model0.81 rating
3.Gemini 3.5 Flash0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.83 rating · Rank 2 of 3

Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026

1.GPT-5.50.84 rating
2.Gemini 3.1 Pro Preview, Current model0.83 rating
3.Gemini 3.5 Flash0.82 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.94 rating · Rank 3 of 3

Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026

1.Gemini 3.5 Flash0.97 rating
1.GPT-5.50.97 rating
3.Gemini 3.1 Pro Preview, Current model0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.97 rating · Rank 2 of 3

Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026

1.Gemini 3.5 Flash0.99 rating
2.Gemini 3.1 Pro Preview, Current model0.97 rating
3.GPT-5.50.96 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

GAIA Reliability

76.2% · Rank 2 of 3

Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026

1.Gemini 3.5 Flash79.2%
2.Gemini 3.1 Pro Preview, Current model76.2%
3.GPT-5.562.8%

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.82 rating · Rank 2 of 3

Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026

1.Gemini 3.5 Flash0.84 rating
2.Gemini 3.1 Pro Preview, Current model0.82 rating
3.GPT-5.50.79 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.76 rating · Rank 1 of 3

Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview, Current model0.76 rating
2.Gemini 3.5 Flash0.75 rating
3.GPT-5.50.61 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.78 rating · Rank 3 of 3

Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026

1.Gemini 3.5 Flash0.8 rating
1.GPT-5.50.8 rating
3.Gemini 3.1 Pro Preview, Current model0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.94 rating · Rank 3 of 3

Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026

1.Gemini 3.5 Flash0.96 rating
2.GPT-5.50.95 rating
3.Gemini 3.1 Pro Preview, Current model0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

1 rating · Rank 1 of 3

Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview, Current model1 rating
1.Gemini 3.5 Flash1 rating
1.GPT-5.51 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

SimpleQA Verified v2

77.5% · Rank 1 of 5

Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026

1.Gemini 3.1 Pro Preview, Current model77.5%
2.Gemini 3.5 Flash70.4%
3.GPT-5.6 Sol69.2%
4.Claude Opus 563.3%
5.Grok 4.553.8%

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 verified prompts without tools. Kaggle independently reproduced the results.

Source: KaggleParticipants: 5

Humanity's Last Exam

46.44% · Rank 1 of 2

Comparison cohort: hle-final-scale-public-temperature-0 · Data date: September 2, 2026

1.Gemini 3.1 Pro Preview, Current model46.44%
2.GPT-5.436.24%

2 of 2 model versions shown in this chart. A higher value ranks first.

All 2,500 public questions, temperature 0, and an automatic answer extractor.

Source: Scale AIParticipants: 2

BrowseComp

85.9% · Rank 3 of 5

Comparison cohort: browsecomp-openai-gpt-5-6-table · Data date: September 2, 2026

1.Claude Mythos 588%
2.GPT-5.6 Terra87.5%
3.Gemini 3.1 Pro Preview, Current model85.9%
4.GPT-5.584.4%
5.GPT-5.6 Luna83.3%

5 of 5 model versions shown in this chart. A higher value ranks first.

Cross-vendor comparison table with browsing tools. Exact agent systems differ.

Editorial selection from the vendor table, not a complete extract of the comparison cohort.

Source: OpenAIParticipants: 5

PRBench Finance

41.87% · Rank 4 of 4

Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026

1.Claude Fable 553.86%
2.GPT-5.6 Sol50.45%
3.GPT-5.445.63%
4.Gemini 3.1 Pro Preview, Current model41.87%

4 of 4 model versions shown in this chart. A higher value ranks first.

600 finance tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

44.02% · Rank 4 of 4

Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026

1.Claude Fable 552.56%
2.GPT-5.6 Sol50.5%
3.GPT-5.444.35%
4.Gemini 3.1 Pro Preview, Current model44.02%

4 of 4 model versions shown in this chart. A higher value ranks first.

500 legal tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

MultiNRC

64.74% · Rank 1 of 2

Task: Overall · Comparison cohort: scale-multinrc-native-languages · Data date: August 26, 2026

1.Gemini 3.1 Pro Preview, Current model64.74%
2.GPT-5.458.29%

2 of 2 model versions shown in this chart. A higher value ranks first.

Native, non-translated reasoning tasks with automatic evaluation.

Source: Scale AIParticipants: 2

EnigmaEval

36.78% · Rank 3 of 4

Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026

1.Claude Fable 539.28%
2.GPT-5.6 Sol37.12%
3.Gemini 3.1 Pro Preview, Current model36.78%
4.Gemini 3.5 Flash25.41%

4 of 4 model versions shown in this chart. A higher value ranks first.

Complex multimodal puzzles with exact answer matching and one attempt.

Source: Scale AIParticipants: 4

VisualToolBench

28.97% · Rank 2 of 2

Task: Overall · Comparison cohort: scale-visual-toolbench-apr · Data date: August 26, 2026

1.GPT-5.429.17%
2.Gemini 3.1 Pro Preview, Current model28.97%

2 of 2 model versions shown in this chart. A higher value ranks first.

1,204 tasks, six tools, and at most 20 tool calls per task.

Source: Scale AIParticipants: 2

Individual values

Published individual values

No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.

MultiChallenge

Task: Overall · Data date: August 26, 2026

MultiChallenge71.37%

Scale 0 to 100. Not ranked

Multi-turn conversations with binary task-specific rubrics.

TutorBench

Task: Overall · Data date: August 26, 2026

TutorBench52.99%

Scale 0 to 100. Not ranked

1,490 tasks with weighted rubrics and Claude 4 Sonnet as judge.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,048,576 tokensSource
Notes
2x reasoning improvement over Gemini 3 Pro, GPQA Diamond 94.3%, SWE-bench 80.6%Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$2 per 1M tokensSource
API input
$4 per 1M tokens (above 200,000 context tokens)Source
API output
$12 per 1M tokensSource
API output
$18 per 1M tokens (above 200,000 context tokens)Source
Cache read
$0.2 per 1M tokensSource
Cache read
$0.4 per 1M tokens (above 200,000 context tokens)Source

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

METR Time Horizon

384.15 minutes

Retrieved August 26, 2026

METR

METR Time Horizon

89.8 minutes

Retrieved August 26, 2026

METR

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • Model metadataSource
  • API pricingSource
  • Princeton HAL (retrieved August 5, 2026)Source
  • Princeton HAL (retrieved August 6, 2026)Source
  • METR (retrieved August 26, 2026)Source
  • Kaggle (retrieved August 9, 2026)Source
  • Scale AI (retrieved September 2, 2026)Source
  • OpenAI (retrieved September 2, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source