Skip to main content
Image modelNo status listed

GPT Image 2

OpenAI

Released
April 2026
Data date
August 15, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.

Position
Rank 2 of 28
Index score
71 / 100
Coverage
4 / 4

Leaderboard

1MAI-Image-2.5 Flash76.6 / 100
2GPT Image 2, Current model71 / 100
3MAI-Image-2.567.6 / 100
28Stable Diffusion 3.5 Large2.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

Text to Image Arena

1,371 Elo · Rank 1 of 25

Sample: 14,328 · Data date: August 26, 2026

1.GPT Image 2, Current model1,371 Elo
2.Reve 2.11,322 Elo
3.Nano Banana 21,321 Elo
25.FLUX.2 Klein 9B1,144 Elo

4 of 25 model versions shown in this chart. A higher value ranks first.

Image Editing Arena

1,257 Elo · Rank 2 of 19

Sample: 18,356 · Data date: August 26, 2026

1.Reve 2.11,263 Elo
2.GPT Image 2, Current model1,257 Elo
2.MAI-Image-2.51,257 Elo
4.GPT Image 1.51,251 Elo
19.FLUX.2 Flex1,162 Elo

5 of 19 model versions shown in this chart. A higher value ranks first.

GenExam Mathematics, strict

50.3% · Rank 3 of 6

Sample: 151 · Data date: August 26, 2026

1.Nano Banana 256.3%
2.Nano Banana Pro55.6%
3.GPT Image 2, Current model50.3%
4.GPT Image 1.526.5%
5.FLUX.2 Max6.6%
6.Seedream 4.02.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Physics, strict

79.6% · Rank 1 of 6

Sample: 113 · Data date: August 26, 2026

1.GPT Image 2, Current model79.6%
2.Nano Banana Pro75.2%
3.Nano Banana 274.3%
4.GPT Image 1.546%
5.FLUX.2 Max8.8%
6.Seedream 4.03.5%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Chemistry, strict

69.5% · Rank 1 of 6

Sample: 118 · Data date: August 26, 2026

1.GPT Image 2, Current model69.5%
2.Nano Banana Pro60.2%
3.Nano Banana 252.5%
4.GPT Image 1.539%
5.FLUX.2 Max6.8%
6.Seedream 4.05.9%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Biology, strict

89.1% · Rank 1 of 6

Sample: 156 · Data date: August 26, 2026

1.GPT Image 2, Current model89.1%
2.Nano Banana Pro75.6%
3.Nano Banana 266%
4.GPT Image 1.556.4%
5.Seedream 4.018.6%
6.FLUX.2 Max11.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Geography, strict

84.8% · Rank 1 of 6

Sample: 66 · Data date: August 26, 2026

1.GPT Image 2, Current model84.8%
2.Nano Banana Pro75.8%
3.Nano Banana 269.7%
4.GPT Image 1.560.6%
5.FLUX.2 Max15.2%
6.Seedream 4.010.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Computer science, strict

73.5% · Rank 1 of 6

Sample: 102 · Data date: August 26, 2026

1.GPT Image 2, Current model73.5%
2.Nano Banana Pro65.7%
3.Nano Banana 256.9%
4.GPT Image 1.536.3%
5.FLUX.2 Max8.8%
6.Seedream 4.06.9%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Engineering, strict

79.3% · Rank 1 of 6

Sample: 111 · Data date: August 26, 2026

1.GPT Image 2, Current model79.3%
2.Nano Banana Pro71.2%
3.Nano Banana 267.6%
4.GPT Image 1.544.1%
5.Seedream 4.011.7%
6.FLUX.2 Max10.8%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Economics, strict

83.1% · Rank 2 of 6

Sample: 77 · Data date: August 26, 2026

1.Nano Banana Pro88.3%
2.GPT Image 2, Current model83.1%
3.Nano Banana 263.6%
4.GPT Image 1.542.9%
5.Seedream 4.05.2%
6.FLUX.2 Max2.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Music, strict

64.6% · Rank 1 of 6

Sample: 65 · Data date: August 26, 2026

1.GPT Image 2, Current model64.6%
2.Nano Banana Pro61.5%
3.Nano Banana 250.8%
4.GPT Image 1.529.2%
5.FLUX.2 Max6.2%
6.Seedream 4.00%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam History, strict

82.9% · Rank 2 of 6

Sample: 41 · Data date: August 26, 2026

1.Nano Banana Pro97.6%
2.GPT Image 2, Current model82.9%
2.Nano Banana 282.9%
4.GPT Image 1.551.2%
5.FLUX.2 Max7.3%
5.Seedream 4.07.3%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam overall, strict

74.6% · Rank 1 of 6

Sample: 1,000 · Data date: August 26, 2026

1.GPT Image 2, Current model74.6%
2.Nano Banana Pro72.7%
3.Nano Banana 264.1%
4.GPT Image 1.543.2%
5.FLUX.2 Max8.5%
6.Seedream 4.07.2%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Mathematics, relaxed

85.2% · Rank 3 of 6

Sample: 151 · Data date: August 26, 2026

1.Nano Banana 287.8%
2.Nano Banana Pro86.3%
3.GPT Image 2, Current model85.2%
4.GPT Image 1.565.8%
5.FLUX.2 Max49.1%
6.Seedream 4.039.8%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Physics, relaxed

95.6% · Rank 2 of 6

Sample: 113 · Data date: August 26, 2026

1.Nano Banana 295.7%
2.GPT Image 2, Current model95.6%
3.Nano Banana Pro95.1%
4.GPT Image 1.585.4%
5.FLUX.2 Max63.2%
6.Seedream 4.049%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Chemistry, relaxed

92% · Rank 1 of 6

Sample: 118 · Data date: August 26, 2026

1.GPT Image 2, Current model92%
2.Nano Banana 290%
3.Nano Banana Pro88.7%
4.GPT Image 1.578.1%
5.FLUX.2 Max54%
6.Seedream 4.046.1%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Biology, relaxed

97.5% · Rank 1 of 6

Sample: 156 · Data date: August 26, 2026

1.GPT Image 2, Current model97.5%
2.Nano Banana Pro95.9%
3.Nano Banana 295.2%
4.GPT Image 1.591.9%
5.FLUX.2 Max74.5%
6.Seedream 4.071%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Geography, relaxed

97.6% · Rank 1 of 6

Sample: 66 · Data date: August 26, 2026

1.GPT Image 2, Current model97.6%
2.Nano Banana Pro96.5%
3.Nano Banana 294.8%
4.GPT Image 1.592.5%
5.FLUX.2 Max76.3%
6.Seedream 4.065.1%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Computer science, relaxed

93.3% · Rank 1 of 6

Sample: 102 · Data date: August 26, 2026

1.GPT Image 2, Current model93.3%
2.Nano Banana Pro91.7%
3.Nano Banana 288.8%
4.GPT Image 1.575.8%
5.FLUX.2 Max56.5%
6.Seedream 4.052.2%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Engineering, relaxed

96.5% · Rank 1 of 6

Sample: 111 · Data date: August 26, 2026

1.GPT Image 2, Current model96.5%
2.Nano Banana 295.8%
3.Nano Banana Pro95.1%
4.GPT Image 1.586.4%
5.FLUX.2 Max68.9%
6.Seedream 4.060%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Economics, relaxed

97.7% · Rank 1 of 6

Sample: 77 · Data date: August 26, 2026

1.GPT Image 2, Current model97.7%
2.Nano Banana Pro97.2%
3.Nano Banana 294.2%
4.GPT Image 1.585.5%
5.FLUX.2 Max61.5%
6.Seedream 4.056%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Music, relaxed

89.1% · Rank 2 of 6

Sample: 65 · Data date: August 26, 2026

1.Nano Banana Pro91%
2.GPT Image 2, Current model89.1%
3.Nano Banana 286.9%
4.GPT Image 1.570.8%
5.FLUX.2 Max47%
6.Seedream 4.034.5%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam History, relaxed

97.1% · Rank 3 of 6

Sample: 41 · Data date: August 26, 2026

1.Nano Banana Pro99.9%
2.Nano Banana 297.3%
3.GPT Image 2, Current model97.1%
4.GPT Image 1.590.9%
5.FLUX.2 Max68%
6.Seedream 4.056.7%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam overall, relaxed

93.8% · Rank 1 of 6

Sample: 1,000 · Data date: August 26, 2026

1.GPT Image 2, Current model93.8%
2.Nano Banana Pro93.7%
3.Nano Banana 292.6%
4.GPT Image 1.582.3%
5.FLUX.2 Max61.9%
6.Seedream 4.053%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

Qwen Image Bench quality

58.65 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model58.65 points
2.Nano Banana Pro55.67 points
3.GPT Image 1.555.14 points
4.Nano Banana 254.77 points
5.Qwen Image 2.0 Pro54.39 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench aesthetics

67.53 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model67.53 points
2.Nano Banana 261.08 points
3.GPT Image 1.560.88 points
4.Nano Banana Pro60.26 points
5.Qwen Image 2.0 Pro58.67 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench alignment

65.85 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model65.85 points
2.Nano Banana 262.4 points
3.GPT Image 1.561.72 points
4.Nano Banana Pro61.25 points
5.Qwen Image 2.0 Pro59.28 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench real-world fidelity

57.38 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model57.38 points
2.Nano Banana 254.28 points
3.Nano Banana Pro54.07 points
4.GPT Image 1.553.95 points
5.Qwen Image 2.0 Pro51.83 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench creative generation

75.23 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model75.23 points
2.Nano Banana 267.05 points
3.GPT Image 1.566.35 points
4.Nano Banana Pro66.23 points
5.Qwen Image 2.0 Pro64.94 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench overall

64.69 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 2, Current model64.69 points
2.Nano Banana 259.82 points
3.GPT Image 1.559.65 points
4.Nano Banana Pro59.45 points
5.Qwen Image 2.0 Pro57.84 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

GRADE reasoning

82.2 points · Rank 1 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 2, Current model82.2 points
2.Nano Banana Pro77.5 points
3.Nano Banana 272.6 points
4.GPT Image 1.554.5 points
5.FLUX.2 Max47.8 points
6.FLUX.2 Pro38.9 points
7.Seedream 4.032.4 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE consistency

94.4 points · Rank 1 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 2, Current model94.4 points
2.Nano Banana Pro89.5 points
3.Nano Banana 286.4 points
4.GPT Image 1.582.3 points
5.FLUX.2 Max67.2 points
6.FLUX.2 Pro55.5 points
7.Seedream 4.053.2 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE readability

98.8 points · Rank 1 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 2, Current model98.8 points
2.Nano Banana 295.9 points
3.Nano Banana Pro95.8 points
4.GPT Image 1.590.7 points
5.Seedream 4.077 points
6.FLUX.2 Pro70.3 points
7.FLUX.2 Max68.6 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE accuracy

56 points · Rank 1 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 2, Current model56 points
2.Nano Banana Pro46.2 points
3.Nano Banana 239.6 points
4.GPT Image 1.516 points
5.FLUX.2 Max11.9 points
6.FLUX.2 Pro4.4 points
7.Seedream 4.03.1 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

Gradually sample test: Typography and layout

35.8 rank points · Rank 20 of 28

Sample: 1 · Data date: August 18, 2026

1.MAI-Image-2.593.8 rank points
18.Seedream 5.0 Pro40.7 rank points
19.Krea 2 Medium Turbo39.5 rank points
20.GPT Image 2, Current model35.8 rank points
20.Reve 2.135.8 rank points
22.GPT Image 1.533.3 rank points
28.Stable Diffusion 3.5 Large0 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Product photography

95.1 rank points · Rank 1 of 28

Sample: 1 · Data date: August 18, 2026

1.GPT Image 2, Current model95.1 rank points
2.MAI-Image-2.5 Flash93.8 rank points
3.MAI-Image-2.586.4 rank points
28.Stable Diffusion 3.5 Large3.7 rank points

4 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Character and detail

76.6 rank points · Rank 5 of 28

Sample: 1 · Data date: August 18, 2026

1.MAI-Image-2.5 Flash86.4 rank points
3.FLUX.2 Max80.3 rank points
4.HiDream-O1-Image-1.579 rank points
5.GPT Image 2, Current model76.6 rank points
5.Qwen Image 2.0 Pro76.6 rank points
7.Recraft V4.1 Utility72.9 rank points
28.Stable Diffusion 3.5 Large3.7 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Infographic

76.6 rank points · Rank 7 of 28

Sample: 1 · Data date: August 18, 2026

1.Reve 2.197.5 rank points
5.Krea 2 Medium Turbo82.7 rank points
6.MAI-Image-2.5 Flash80.3 rank points
7.GPT Image 2, Current model76.6 rank points
8.HiDream-O1-Image-1.574.1 rank points
9.Seedream 4.067.9 rank points
28.Firefly Image Model 51.2 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Weights
ProprietarySource
Access
API, WebSource
Maximum output
4KSource
Benchmark configuration
HighSource
Generator
gpt-image-2Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

Representative price
$0.21 per imageSource
Together AI (openai/gpt-image-2)
$0.05 per imageSource

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

LMArena Text to Image: 3D modeling

1,360.37 points

Rank 1 · 6,635 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Art

1,370.44 points

Rank 1 · 8,762 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Cartoon

1,396.03 points

Rank 1 · 28,210 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Commercial design

1,389.56 points

Rank 1 · 27,138 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Overall

1,380.47 points

Rank 1 · 69,194 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Photorealism

1,378.29 points

Rank 1 · 26,719 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Portraits

1,428.87 points

Rank 1 · 13,521 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Text rendering

1,424.05 points

Rank 1 · 26,226 samples · Retrieved August 26, 2026

LMArena

LMArena Image Editing: Multi-image editing

1,454.12 points

Rank 1 · 72,134 samples · Retrieved August 26, 2026

LMArena

LMArena Image Editing: Overall

1,462.8 points

Rank 1 · 199,275 samples · Retrieved August 26, 2026

LMArena

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • OpenAISource
  • Image pricingSource
  • Together AISource
  • Artificial Analysis (retrieved August 26, 2026)Source
  • LMArena (retrieved August 26, 2026)Source
  • LMArena (retrieved August 26, 2026)Source
  • GenExam (retrieved August 26, 2026)Source
  • Qwen Image Bench (retrieved August 26, 2026)Source
  • GRADE (retrieved August 26, 2026)Source
  • Gradually-Bildtest (retrieved August 18, 2026)Source