Skip to main content
Language modelActive

GPT-5.6 Luna

OpenAI

Released
June 26, 2026
Data date
August 11, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.

Position
Rank 8 of 26
Index score
66.8 / 100
Coverage
4 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
7Grok 4.670.2 / 100
8GPT-5.6 Luna, Current model66.8 / 100
9Gemini 3.6 Flash63 / 100
25MiMo-V2.5-Pro12.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

BrowseComp

83.3% · Rank 5 of 5

Comparison cohort: browsecomp-openai-gpt-5-6-table · Data date: September 2, 2026

1.Claude Mythos 588%
2.GPT-5.6 Terra87.5%
3.Gemini 3.1 Pro Preview85.9%
4.GPT-5.584.4%
5.GPT-5.6 Luna, Current model83.3%

5 of 5 model versions shown in this chart. A higher value ranks first.

Cross-vendor comparison table with browsing tools. Exact agent systems differ.

Editorial selection from the vendor table, not a complete extract of the comparison cohort.

Source: OpenAIParticipants: 5

SWE-Bench Pro

62.7% · Rank 5 of 5

Comparison cohort: swe-pro-openai-gpt-5-6-release · Data date: September 2, 2026

1.Claude Mythos 580.3%
2.Claude Fable 580%
3.GPT-5.6 Sol64.6%
4.GPT-5.6 Terra63.4%
5.GPT-5.6 Luna, Current model62.7%

5 of 5 model versions shown in this chart. A higher value ranks first.

Comparison table published by OpenAI. Not a Gradually test.

After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.

Source: OpenAIParticipants: 5

Individual values

Published individual values

No exactly matching published comparison cohort is available for these values. The bar shows only the documented scale, not a rank.

ARC-AGI-3

Task: RHAE overall score · Data date: August 26, 2026

ARC-AGI-30.18%

Scale 0 to 100. Not ranked

Official ARC Prize harness on unseen interactive environments. Reasoning levels remain separate.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,050,000 tokensSource
Knowledge cutoff
February 16, 2026Source
Notes
Fast and cheapest tier of the GPT-5.6 family. Generally available from 2026-07-09. Effort-selectable on ChatGPT Work and Codex for Plus, Pro, Business, and Enterprise users. $1 input / $6 output per 1M tokens. Terminal-Bench 2.1 84.3%.Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$0.2 per 1M tokensSource
API input
$0.4 per 1M tokens (above 272,000 context tokens)Source
API output
$1.2 per 1M tokensSource
API output
$1.8 per 1M tokens (above 272,000 context tokens)Source
Cache write
$0.25 per 1M tokensSource
Cache write
$0.5 per 1M tokens (above 272,000 context tokens)Source
Cache read
$0.02 per 1M tokensSource
Cache read
$0.04 per 1M tokens (above 272,000 context tokens)Source

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

ARC-AGI-2

59.5%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

47.6%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

29.3%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

7.4%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

5.1%

Retrieved September 2, 2026

ARC Prize

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.