Skip to main content
Language modelActive

GPT-5.5

OpenAI

Released
April 2026
Data date
July 30, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.

Position
Rank 13 of 26
Index score
57.2 / 100
Coverage
4 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
12Claude Sonnet 557.7 / 100
13GPT-5.5, Current model57.2 / 100
14GLM-5.354.8 / 100
25MiMo-V2.5-Pro12.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

τ-bench Airline Clean

79.5% · Rank 3 of 3

Task: Accuracy · Comparison cohort: hal-tau-airline-clean-hal-reliability-overall · Data date: August 5, 2026

1.Gemini 3.1 Pro Preview82.1%
2.Gemini 3.5 Flash80.8%
3.GPT-5.5, Current model79.5%

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.89 rating · Rank 1 of 3

Task: Reliability · Comparison cohort: hal-tau-airline-clean-hal-reliability-reliability · Data date: August 5, 2026

1.GPT-5.5, Current model0.89 rating
2.Gemini 3.1 Pro Preview0.86 rating
2.Gemini 3.5 Flash0.86 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.84 rating · Rank 1 of 3

Task: Consistency · Comparison cohort: hal-tau-airline-clean-hal-reliability-consistency · Data date: August 5, 2026

1.GPT-5.5, Current model0.84 rating
2.Gemini 3.1 Pro Preview0.81 rating
3.Gemini 3.5 Flash0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.84 rating · Rank 1 of 3

Task: Predictability · Comparison cohort: hal-tau-airline-clean-hal-reliability-predictability · Data date: August 5, 2026

1.GPT-5.5, Current model0.84 rating
2.Gemini 3.1 Pro Preview0.83 rating
3.Gemini 3.5 Flash0.82 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.97 rating · Rank 1 of 3

Task: Robustness · Comparison cohort: hal-tau-airline-clean-hal-reliability-robustness · Data date: August 5, 2026

1.Gemini 3.5 Flash0.97 rating
1.GPT-5.5, Current model0.97 rating
3.Gemini 3.1 Pro Preview0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

τ-bench Airline Clean

0.96 rating · Rank 3 of 3

Task: Safety · Comparison cohort: hal-tau-airline-clean-hal-reliability-safety · Data date: August 5, 2026

1.Gemini 3.5 Flash0.99 rating
2.Gemini 3.1 Pro Preview0.97 rating
3.GPT-5.5, Current model0.96 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

26 cleaned airline tasks. Princeton HAL recomputed every metric from scratch on this subset.

Source: Princeton HALParticipants: 3

GAIA Reliability

62.8% · Rank 3 of 3

Task: Accuracy · Comparison cohort: hal-gaia-reliability-hal-reliability-overall · Data date: August 6, 2026

1.Gemini 3.5 Flash79.2%
2.Gemini 3.1 Pro Preview76.2%
3.GPT-5.5, Current model62.8%

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.79 rating · Rank 3 of 3

Task: Reliability · Comparison cohort: hal-gaia-reliability-hal-reliability-reliability · Data date: August 6, 2026

1.Gemini 3.5 Flash0.84 rating
2.Gemini 3.1 Pro Preview0.82 rating
3.GPT-5.5, Current model0.79 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.61 rating · Rank 3 of 3

Task: Consistency · Comparison cohort: hal-gaia-reliability-hal-reliability-consistency · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview0.76 rating
2.Gemini 3.5 Flash0.75 rating
3.GPT-5.5, Current model0.61 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.8 rating · Rank 1 of 3

Task: Predictability · Comparison cohort: hal-gaia-reliability-hal-reliability-predictability · Data date: August 6, 2026

1.Gemini 3.5 Flash0.8 rating
1.GPT-5.5, Current model0.8 rating
3.Gemini 3.1 Pro Preview0.78 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

0.95 rating · Rank 2 of 3

Task: Robustness · Comparison cohort: hal-gaia-reliability-hal-reliability-robustness · Data date: August 6, 2026

1.Gemini 3.5 Flash0.96 rating
2.GPT-5.5, Current model0.95 rating
3.Gemini 3.1 Pro Preview0.94 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

GAIA Reliability

1 rating · Rank 1 of 3

Task: Safety · Comparison cohort: hal-gaia-reliability-hal-reliability-safety · Data date: August 6, 2026

1.Gemini 3.1 Pro Preview1 rating
1.Gemini 3.5 Flash1 rating
1.GPT-5.5, Current model1 rating

3 of 3 model versions shown in this chart. A higher value ranks first.

165 GAIA tasks with repeated runs and perturbations. HAL does not publish cost for this evaluation.

Source: Princeton HALParticipants: 3

AlmanBench v0.1

95.43% · Rank 1 of 5

Task: Accepted cases · Comparison cohort: almanbench-v0-1:official:overall · Data date: August 9, 2026

1.GPT-5.5, Current model95.43%
2.GPT-5.6 Sol95.14%
3.Claude Opus 594.95%
4.DeepSeek-V4-Flash93%
5.Kimi K390.96%

5 of 5 model versions shown in this chart. A higher value ranks first.

1,029 public tasks with one direct run per row. The target is the Alman language specification.

Source: Alman InstitutParticipants: 5

SWE-Marathon v1.0

12% · Rank 3 of 4

Task: Binary resolution, pass@1 · Comparison cohort: swe-marathon-v1-0:official:pass1 · Data date: August 9, 2026

1.Grok 4.529%
2.Claude Fable 524%
3.GPT-5.5, Current model12%
4.Gemini 3.5 Flash7%

4 of 4 model versions shown in this chart. A higher value ranks first.

20 multi-hour tasks. The named coding agent is part of every result row.

Source: SWE-MarathonParticipants: 4

LiveBench

80.19 points · Rank 3 of 4

Comparison cohort: livebench-2026-06-25-max · Data date: September 2, 2026

1.Claude Fable 582.97 points
2.GPT-5.6 Sol81.05 points
3.GPT-5.5, Current model80.19 points
4.GPT-5.6 Terra77.94 points

4 of 4 model versions shown in this chart. A higher value ranks first.

23 objectively scored tasks across seven categories.

Source: LiveBenchParticipants: 4

LiveBench

89.7 points · Rank 3 of 4

Task: Reasoning · Comparison cohort: livebench-2026-06-25-reasoning-max · Data date: September 2, 2026

1.GPT-5.6 Sol91.7 points
2.GPT-5.6 Terra90.6 points
3.Claude Fable 589.7 points
3.GPT-5.5, Current model89.7 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Reasoning category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

82.1 points · Rank 3 of 4

Task: Coding · Comparison cohort: livebench-2026-06-25-coding-max · Data date: September 2, 2026

1.Claude Fable 586 points
2.GPT-5.6 Sol83.9 points
3.GPT-5.5, Current model82.1 points
4.GPT-5.6 Terra78.2 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Coding category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

54 points · Rank 4 of 4

Task: Agentic coding · Comparison cohort: livebench-2026-06-25-agentic-coding-max · Data date: September 2, 2026

1.Claude Fable 562.2 points
2.GPT-5.6 Sol56.2 points
3.GPT-5.6 Terra54.9 points
4.GPT-5.5, Current model54 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Agentic coding category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

95.9 points · Rank 3 of 4

Task: Mathematics · Comparison cohort: livebench-2026-06-25-mathematics-max · Data date: September 2, 2026

1.GPT-5.6 Sol96.2 points
2.Claude Fable 596 points
3.GPT-5.5, Current model95.9 points
4.GPT-5.6 Terra94.9 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Mathematics category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

81.6 points · Rank 1 of 4

Task: Data analysis · Comparison cohort: livebench-2026-06-25-data-analysis-max · Data date: September 2, 2026

1.GPT-5.5, Current model81.6 points
2.Claude Fable 580.5 points
3.GPT-5.6 Sol79.8 points
4.GPT-5.6 Terra79.3 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Data analysis category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

87.4 points · Rank 3 of 4

Task: Language · Comparison cohort: livebench-2026-06-25-language-max · Data date: September 2, 2026

1.Claude Fable 590.7 points
2.GPT-5.6 Sol87.7 points
3.GPT-5.5, Current model87.4 points
4.GPT-5.6 Terra82.9 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Language category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

70.7 points · Rank 3 of 4

Task: Instruction following · Comparison cohort: livebench-2026-06-25-instruction-following-max · Data date: September 2, 2026

1.Claude Fable 575.8 points
2.GPT-5.6 Sol71.8 points
3.GPT-5.5, Current model70.7 points
4.GPT-5.6 Terra64.6 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Instruction following category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

BrowseComp

84.4% · Rank 4 of 5

Comparison cohort: browsecomp-openai-gpt-5-6-table · Data date: September 2, 2026

1.Claude Mythos 588%
2.GPT-5.6 Terra87.5%
3.Gemini 3.1 Pro Preview85.9%
4.GPT-5.5, Current model84.4%
5.GPT-5.6 Luna83.3%

5 of 5 model versions shown in this chart. A higher value ranks first.

Cross-vendor comparison table with browsing tools. Exact agent systems differ.

Editorial selection from the vendor table, not a complete extract of the comparison cohort.

Source: OpenAIParticipants: 5

Individual values

Published individual values

No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.

OSWorld 2.0

Data date: August 26, 2026

OSWorld 2.013%

Scale 0 to 100. Not ranked

Complete agent system with up to 500 steps.

Claude used about 244,000 output tokens. This is a system result, not an isolated model score.

SWE-Bench Pro

Data date: September 2, 2026

SWE-Bench Pro58.6%

Scale 0 to 100. Not ranked

Comparison table published by Anthropic. Not a Gradually test.

After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,050,000 tokensSource
Knowledge cutoff
December 1, 2025Source
Notes
1M context, agentic workflows, Terminal-Bench 2.0 82.7%, parameter count not disclosedSource

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$5 per 1M tokensSource
API input
$10 per 1M tokens (above 272,000 context tokens)Source
API output
$30 per 1M tokensSource
API output
$45 per 1M tokens (above 272,000 context tokens)Source
Cache read
$0.5 per 1M tokensSource
Cache read
$1 per 1M tokens (above 272,000 context tokens)Source

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

TUA-Bench

64.7%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

57.7%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

68.3%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

42.5%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

64.2%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

57.2%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

66.7%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

46.7%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

62.4%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

54.2%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

67.5%

Retrieved August 9, 2026

TUA-Bench

TUA-Bench

40%

Retrieved August 9, 2026

TUA-Bench

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • Model metadataSource
  • API pricingSource
  • Princeton HAL (retrieved August 5, 2026)Source
  • Princeton HAL (retrieved August 6, 2026)Source
  • Alman Institut (retrieved August 9, 2026)Source
  • TUA-Bench (retrieved August 9, 2026)Source
  • SWE-Marathon (retrieved August 9, 2026)Source
  • LiveBench (retrieved September 2, 2026)Source
  • OpenAI (retrieved September 2, 2026)Source
  • OSWorld 2.0 (retrieved August 26, 2026)Source
  • Anthropic (retrieved September 2, 2026)Source