Skip to main content
Language modelActive

GPT-5.4

OpenAI

Released
March 2026
Data date
July 30, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

This model does not yet have enough published values for an overall rank. It covers 1 of 4 required comparison cohorts.

Coverage
1 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
25MiMo-V2.5-Pro12.5 / 100
GPT-5.4, Current modelNot ranked

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

Humanity's Last Exam

36.24% · Rank 2 of 2

Comparison cohort: hle-final-scale-public-temperature-0 · Data date: September 2, 2026

1.Gemini 3.1 Pro Preview46.44%
2.GPT-5.4, Current model36.24%

2 of 2 model versions shown in this chart. A higher value ranks first.

All 2,500 public questions, temperature 0, and an automatic answer extractor.

Source: Scale AIParticipants: 2

PRBench Finance

45.63% · Rank 3 of 4

Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026

1.Claude Fable 553.86%
2.GPT-5.6 Sol50.45%
3.GPT-5.4, Current model45.63%
4.Gemini 3.1 Pro Preview41.87%

4 of 4 model versions shown in this chart. A higher value ranks first.

600 finance tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

44.35% · Rank 3 of 4

Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026

1.Claude Fable 552.56%
2.GPT-5.6 Sol50.5%
3.GPT-5.4, Current model44.35%
4.Gemini 3.1 Pro Preview44.02%

4 of 4 model versions shown in this chart. A higher value ranks first.

500 legal tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

MultiNRC

58.29% · Rank 2 of 2

Task: Overall · Comparison cohort: scale-multinrc-native-languages · Data date: August 26, 2026

1.Gemini 3.1 Pro Preview64.74%
2.GPT-5.4, Current model58.29%

2 of 2 model versions shown in this chart. A higher value ranks first.

Native, non-translated reasoning tasks with automatic evaluation.

Source: Scale AIParticipants: 2

VisualToolBench

29.17% · Rank 1 of 2

Task: Overall · Comparison cohort: scale-visual-toolbench-apr · Data date: August 26, 2026

1.GPT-5.4, Current model29.17%
2.Gemini 3.1 Pro Preview28.97%

2 of 2 model versions shown in this chart. A higher value ranks first.

1,204 tasks, six tools, and at most 20 tool calls per task.

Source: Scale AIParticipants: 2

Individual values

Published individual values

No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.

SecCodeBench-V2

Task: Weighted total · Data date: August 9, 2026

SecCodeBench-V259.74%

Scale 0 to 100. Not ranked

98 project cases. Functional tests run before dynamic exploit checks and security scoring.

SecCodeBench-V2

Task: Repair without hints · Data date: August 9, 2026

SecCodeBench-V261.41%

Scale 0 to 100. Not ranked

98 project cases. Functional tests run before dynamic exploit checks and security scoring.

SecCodeBench-V2

Task: Repair with security hints · Data date: August 9, 2026

SecCodeBench-V280.04%

Scale 0 to 100. Not ranked

98 project cases. Functional tests run before dynamic exploit checks and security scoring.

SecCodeBench-V2

Task: Generation without hints · Data date: August 9, 2026

SecCodeBench-V250.97%

Scale 0 to 100. Not ranked

98 project cases. Functional tests run before dynamic exploit checks and security scoring.

SecCodeBench-V2

Task: Generation with security hints · Data date: August 9, 2026

SecCodeBench-V267.82%

Scale 0 to 100. Not ranked

98 project cases. Functional tests run before dynamic exploit checks and security scoring.

VISTA

Task: Overall · Data date: August 26, 2026

VISTA50.89%

Scale 0 to 100. Not ranked

758 single-turn image-text tasks with structured yes-no rubrics.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,050,000 tokensSource
Knowledge cutoff
August 31, 2025Source
Notes
1M context, parameter count not disclosedSource

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$2.5 per 1M tokensSource
API input
$5 per 1M tokens (above 272,000 context tokens)Source
API output
$15 per 1M tokensSource
API output
$22.5 per 1M tokens (above 272,000 context tokens)Source
Cache read
$0.25 per 1M tokensSource
Cache read
$0.5 per 1M tokens (above 272,000 context tokens)Source

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

METR Time Horizon

341.74 minutes

Retrieved August 26, 2026

METR

METR Time Horizon

53.88 minutes

Retrieved August 26, 2026

METR

SWE-EVO

25%

Retrieved August 9, 2026

SWE-EVO

SWE-EVO

33.89%

Retrieved August 9, 2026

SWE-EVO

SWE-EVO

25%

Retrieved August 9, 2026

SWE-EVO

SWE-EVO

33.98%

Retrieved August 9, 2026

SWE-EVO

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • Model metadataSource
  • API pricingSource
  • METR (retrieved August 26, 2026)Source
  • SecCodeBench (retrieved August 9, 2026)Source
  • EnterpriseRAG-Bench (retrieved August 9, 2026)Source
  • SWE-EVO (retrieved August 9, 2026)Source
  • Scale AI (retrieved September 2, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source