Skip to main content
Language modelActive

GPT-5.6 Sol

OpenAI

Released
June 26, 2026
Data date
August 24, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

The index averages rank percentiles from 4 documented comparison cohorts. Only models with complete coverage receive a position.

Position
Rank 4 of 26
Index score
81.7 / 100
Coverage
4 / 4

Leaderboard

1Claude Opus 588.9 / 100
2Gemini 3.8 Flash88.5 / 100
3Claude Fable 583.2 / 100
4GPT-5.6 Sol, Current model81.7 / 100
5Gemini 3.7 Flash76.4 / 100
25MiMo-V2.5-Pro12.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

SimpleQA Verified v2

69.2% · Rank 3 of 5

Task: Kaggle score · Comparison cohort: simpleqa-verified-v2:official:overall · Data date: August 9, 2026

1.Gemini 3.1 Pro Preview77.5%
2.Gemini 3.5 Flash70.4%
3.GPT-5.6 Sol, Current model69.2%
4.Claude Opus 563.3%
5.Grok 4.553.8%

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 verified prompts without tools. Kaggle independently reproduced the results.

Source: KaggleParticipants: 5

AlmanBench v0.1

95.14% · Rank 2 of 5

Task: Accepted cases · Comparison cohort: almanbench-v0-1:official:overall · Data date: August 9, 2026

1.GPT-5.595.43%
2.GPT-5.6 Sol, Current model95.14%
3.Claude Opus 594.95%
4.DeepSeek-V4-Flash93%
5.Kimi K390.96%

5 of 5 model versions shown in this chart. A higher value ranks first.

1,029 public tasks with one direct run per row. The target is the Alman language specification.

Source: Alman InstitutParticipants: 5

LiveBench

81.05 points · Rank 2 of 4

Comparison cohort: livebench-2026-06-25-max · Data date: September 2, 2026

1.Claude Fable 582.97 points
2.GPT-5.6 Sol, Current model81.05 points
3.GPT-5.580.19 points
4.GPT-5.6 Terra77.94 points

4 of 4 model versions shown in this chart. A higher value ranks first.

23 objectively scored tasks across seven categories.

Source: LiveBenchParticipants: 4

LiveBench

91.7 points · Rank 1 of 4

Task: Reasoning · Comparison cohort: livebench-2026-06-25-reasoning-max · Data date: September 2, 2026

1.GPT-5.6 Sol, Current model91.7 points
2.GPT-5.6 Terra90.6 points
3.Claude Fable 589.7 points
3.GPT-5.589.7 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Reasoning category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

83.9 points · Rank 2 of 4

Task: Coding · Comparison cohort: livebench-2026-06-25-coding-max · Data date: September 2, 2026

1.Claude Fable 586 points
2.GPT-5.6 Sol, Current model83.9 points
3.GPT-5.582.1 points
4.GPT-5.6 Terra78.2 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Coding category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

56.2 points · Rank 2 of 4

Task: Agentic coding · Comparison cohort: livebench-2026-06-25-agentic-coding-max · Data date: September 2, 2026

1.Claude Fable 562.2 points
2.GPT-5.6 Sol, Current model56.2 points
3.GPT-5.6 Terra54.9 points
4.GPT-5.554 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Agentic coding category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

96.2 points · Rank 1 of 4

Task: Mathematics · Comparison cohort: livebench-2026-06-25-mathematics-max · Data date: September 2, 2026

1.GPT-5.6 Sol, Current model96.2 points
2.Claude Fable 596 points
3.GPT-5.595.9 points
4.GPT-5.6 Terra94.9 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Mathematics category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

79.8 points · Rank 3 of 4

Task: Data analysis · Comparison cohort: livebench-2026-06-25-data-analysis-max · Data date: September 2, 2026

1.GPT-5.581.6 points
2.Claude Fable 580.5 points
3.GPT-5.6 Sol, Current model79.8 points
4.GPT-5.6 Terra79.3 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Data analysis category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

87.7 points · Rank 2 of 4

Task: Language · Comparison cohort: livebench-2026-06-25-language-max · Data date: September 2, 2026

1.Claude Fable 590.7 points
2.GPT-5.6 Sol, Current model87.7 points
3.GPT-5.587.4 points
4.GPT-5.6 Terra82.9 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Language category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

LiveBench

71.8 points · Rank 2 of 4

Task: Instruction following · Comparison cohort: livebench-2026-06-25-instruction-following-max · Data date: September 2, 2026

1.Claude Fable 575.8 points
2.GPT-5.6 Sol, Current model71.8 points
3.GPT-5.570.7 points
4.GPT-5.6 Terra64.6 points

4 of 4 model versions shown in this chart. A higher value ranks first.

Instruction following category inside the same LiveBench release.

Source: LiveBenchParticipants: 4

SWE-Bench Pro

64.6% · Rank 3 of 5

Comparison cohort: swe-pro-openai-gpt-5-6-release · Data date: September 2, 2026

1.Claude Mythos 580.3%
2.Claude Fable 580%
3.GPT-5.6 Sol, Current model64.6%
4.GPT-5.6 Terra63.4%
5.GPT-5.6 Luna62.7%

5 of 5 model versions shown in this chart. A higher value ranks first.

Comparison table published by OpenAI. Not a Gradually test.

After an audit, OpenAI estimates that about 30% of the public tasks are broken. The result therefore remains a disputed secondary signal. The rows are an editorial selection from the respective comparison table.

Source: OpenAIParticipants: 5

PRBench Finance

50.45% · Rank 2 of 4

Task: Overall · Comparison cohort: scale-prbench-finance-full · Data date: August 26, 2026

1.Claude Fable 553.86%
2.GPT-5.6 Sol, Current model50.45%
3.GPT-5.445.63%
4.Gemini 3.1 Pro Preview41.87%

4 of 4 model versions shown in this chart. A higher value ranks first.

600 finance tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

50.5% · Rank 2 of 4

Task: Overall · Comparison cohort: scale-prbench-legal-full · Data date: August 26, 2026

1.Claude Fable 552.56%
2.GPT-5.6 Sol, Current model50.5%
3.GPT-5.444.35%
4.Gemini 3.1 Pro Preview44.02%

4 of 4 model versions shown in this chart. A higher value ranks first.

500 legal tasks with task-specific rubrics and o4-mini as judge.

Source: Scale AIParticipants: 4

EnigmaEval

37.12% · Rank 2 of 4

Task: Overall · Comparison cohort: scale-enigma-eval-2026-07-23 · Data date: August 26, 2026

1.Claude Fable 539.28%
2.GPT-5.6 Sol, Current model37.12%
3.Gemini 3.1 Pro Preview36.78%
4.Gemini 3.5 Flash25.41%

4 of 4 model versions shown in this chart. A higher value ranks first.

Complex multimodal puzzles with exact answer matching and one attempt.

Source: Scale AIParticipants: 4

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Model type
ProprietarySource
Context window
1,050,000 tokensSource
Knowledge cutoff
February 16, 2026Source
Notes
Flagship of the GPT-5.6 family (Sol, Terra, Luna). Previewed 2026-06-26, generally available from 2026-07-09 after the US government lifted the restricted-access framework. Available in ChatGPT (Plus, Pro, Business, Enterprise from medium effort setting; Pro and Enterprise also get Sol Pro), Codex, and the API. Terminal-Bench 2.1 88.8% (Sol Ultra mode with four parallel sub-agents: 91.9%), Artificial Analysis Coding Agent Index 80, SWE-Bench Pro 64.6% (about 15 points behind Claude Mythos 5 and Fable 5). New at GA: programmatic tool calling with model-written JavaScript in a V8 sandbox.Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

API input
$4 per 1M tokensSource
API input
$8 per 1M tokens (above 272,000 context tokens)Source
API output
$20 per 1M tokensSource
API output
$30 per 1M tokens (above 272,000 context tokens)Source
Cache write
$5 per 1M tokensSource
Cache write
$10 per 1M tokens (above 272,000 context tokens)Source
Cache read
$0.4 per 1M tokensSource
Cache read
$0.8 per 1M tokens (above 272,000 context tokens)Source

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

ARC-AGI-3

7.78%

Retrieved August 26, 2026

ARC Prize

ARC-AGI-3

6.99%

Retrieved August 26, 2026

ARC Prize

ARC-AGI-3

2.15%

Retrieved August 26, 2026

ARC Prize

ARC-AGI-3

1.07%

Retrieved August 26, 2026

ARC Prize

ARC-AGI-3

0.33%

Retrieved August 26, 2026

ARC Prize

ARC-AGI-2

92.5%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

90%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

85.4%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

67.1%

Retrieved September 2, 2026

ARC Prize

ARC-AGI-2

42.5%

Retrieved September 2, 2026

ARC Prize

BrowseComp

92.2%

Retrieved September 2, 2026 · Editorial selection from the vendor table, not a complete extract of the comparison cohort.

OpenAI

BrowseComp

90.4%

Retrieved September 2, 2026 · Editorial selection from the vendor table, not a complete extract of the comparison cohort.

OpenAI

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • Model metadataSource
  • Context window dataSource
  • Knowledge cutoff dataSource
  • API pricingSource
  • ARC Prize (retrieved August 26, 2026)Source
  • Kaggle (retrieved August 9, 2026)Source
  • Alman Institut (retrieved August 9, 2026)Source
  • LiveBench (retrieved September 2, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source
  • Scale AI (retrieved August 26, 2026)Source