Skip to main content

Head-to-head comparison

GPT-5.4 vs. Gemini 3.7 Flash

Compare published benchmark results from matching versions and cohorts, API pricing, and technical specifications. Any documented test setup differences remain visible.

2

shared result series

2

distinct benchmarks

2026-08-16

latest retrieval

Change model selection

Specifications and cost

General specifications and standard API pricing

Undisclosed parameter counts and knowledge cutoffs stay marked as unavailable. The table shows direct standard token pricing, including documented context tiers. Batch, fast, priority, regional surcharges, tool calls, and cloud platform pricing are excluded.

Specifications and API pricing for GPT-5.4 and Gemini 3.7 Flash
AttributeGPT-5.4Gemini 3.7 Flash
ProviderOpenAIGoogle
ReleasedMar 2026Aug 13, 2026
AvailabilityActiveActive
Model typeProprietaryProprietary
ParametersNot publishedNot published
ArchitectureNot publishedNot published
Context window1,050,000 tokens1,048,576 tokens
Knowledge cutoffAugust 31, 2025March 2026
Technical sourcesModelModel, Knowledge cutoff
API pricing per 1M tokens
API input$2.5 ≤272K / $5 >272K$0.75
API output$15 ≤272K / $22.5 >272K$3.75
Cache read$0.25 ≤272K / $0.5 >272K$0.07
Cache write (5 min.)Not listedNot listed
Cache write (1 hr.)Not listedNot listed
Price verified07/30/2026Source08/16/2026Source
GPT-5.4Gemini 3.7 Flash
API inputUSD per 1M tokens, base tier
GPT-5.4$2.5
Gemini 3.7 Flash$0.75Lower price
API outputUSD per 1M tokens, base tier
GPT-5.4$15
Gemini 3.7 Flash$3.75Lower price
Cache readUSD per 1M tokens, base tier
GPT-5.4$0.25
Gemini 3.7 Flash$0.075Lower price

Performance

Shared benchmarks

Only published results with the same benchmark version, task, metric, and comparison cohort are paired. Different reasoning levels, agents, harnesses, or output limits appear directly in the table.

2 shared result series
GPT-5.4Gemini 3.7 Flash
Harvey's Legal Agent BenchmarkOverall · Task fully resolved. Higher is better. No documented setup difference
GPT-5.40 %
Gemini 3.7 Flash8.75 %Winner
Vibe Code Bench v1.1Overall. Higher is better. No documented setup difference
GPT-5.467.421 %
Gemini 3.7 Flash70.395 %Winner
Benchmark scores for GPT-5.4 and Gemini 3.7 Flash
Benchmark and taskGPT-5.4Gemini 3.7 FlashComparabilitySource
Harvey's Legal Agent BenchmarkOverall · Task fully resolved
0 %
8.75 %Winner
1task resolution rateHigher is betterCohort: vals-legal-agent-benchmark:1:overallNo documented setup differenceWinner: Gemini 3.7 FlashGPT-5.4: Snapshot: gpt-5.4-2026-03-05
Gemini 3.7 Flash: No further settings disclosed
Vals AIType: Independent evaluationLeaderboard updated: 2026-08-14Import ID: vals-hlab-2026-08-16-bae90566df2fRetrieved 2026-08-16
Vibe Code Bench v1.1Overall
67.421 %
70.395 %Winner
1.1accuracyHigher is betterCohort: vals-vibe-code-bench:1.1:overallNo documented setup differenceWinner: Gemini 3.7 FlashGPT-5.4: Snapshot: gpt-5.4-2026-03-05, Harness: OpenHands
Gemini 3.7 Flash: Harness: OpenHands
Vals AIType: Independent evaluationLeaderboard updated: 2026-08-13Import ID: vals-vibe-code-2026-08-16-a73c9770bcc4Retrieved 2026-08-16

How to read this comparison

Winning one benchmark is not an overall verdict

Your workload, budget, and required context length matter most. A coding benchmark says little about visual understanding or agent performance.

Prices use official API rates per 1M tokens. Web subscriptions and cloud platform surcharges can differ.

Every result links to its measurement source and retrieval date. Labels distinguish vendor reports, official benchmark leaderboards, and independent evaluations. Disputed or archived results are excluded.