Skip to main content
Image modelDeprecated

GPT Image 1.5

OpenAI

Released
December 2025
Data date
August 15, 2026

Category view

Position within the category

This overview uses only published data from matching cohorts. Missing values never change a rank.

This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.

Position
Rank 17 of 28
Index score
51.2 / 100
Coverage
4 / 4

Leaderboard

1MAI-Image-2.5 Flash76.6 / 100
2GPT Image 271 / 100
3MAI-Image-2.567.6 / 100
16Ideogram 4.051.9 / 100
17GPT Image 1.5, Current model51.2 / 100
18FLUX.2 Flex47.2 / 100
28Stable Diffusion 3.5 Large2.5 / 100

Measurements

Comparable benchmark results

Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.

Text to Image Arena

1,310 Elo · Rank 4 of 25

Sample: 13,674 · Data date: August 26, 2026

1.GPT Image 21,371 Elo
2.Reve 2.11,322 Elo
3.Nano Banana 21,321 Elo
4.GPT Image 1.5, Current model1,310 Elo
5.MAI-Image-2.51,303 Elo
6.Nano Banana Pro1,298 Elo
25.FLUX.2 Klein 9B1,144 Elo

7 of 25 model versions shown in this chart. A higher value ranks first.

Image Editing Arena

1,251 Elo · Rank 4 of 19

Sample: 11,694 · Data date: August 26, 2026

1.Reve 2.11,263 Elo
2.GPT Image 21,257 Elo
2.MAI-Image-2.51,257 Elo
4.GPT Image 1.5, Current model1,251 Elo
5.Nano Banana 21,250 Elo
6.Seedream 5.0 Pro1,248 Elo
19.FLUX.2 Flex1,162 Elo

7 of 19 model versions shown in this chart. A higher value ranks first.

GenExam Mathematics, strict

26.5% · Rank 4 of 6

Sample: 151 · Data date: August 26, 2026

1.Nano Banana 256.3%
2.Nano Banana Pro55.6%
3.GPT Image 250.3%
4.GPT Image 1.5, Current model26.5%
5.FLUX.2 Max6.6%
6.Seedream 4.02.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Physics, strict

46% · Rank 4 of 6

Sample: 113 · Data date: August 26, 2026

1.GPT Image 279.6%
2.Nano Banana Pro75.2%
3.Nano Banana 274.3%
4.GPT Image 1.5, Current model46%
5.FLUX.2 Max8.8%
6.Seedream 4.03.5%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Chemistry, strict

39% · Rank 4 of 6

Sample: 118 · Data date: August 26, 2026

1.GPT Image 269.5%
2.Nano Banana Pro60.2%
3.Nano Banana 252.5%
4.GPT Image 1.5, Current model39%
5.FLUX.2 Max6.8%
6.Seedream 4.05.9%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Biology, strict

56.4% · Rank 4 of 6

Sample: 156 · Data date: August 26, 2026

1.GPT Image 289.1%
2.Nano Banana Pro75.6%
3.Nano Banana 266%
4.GPT Image 1.5, Current model56.4%
5.Seedream 4.018.6%
6.FLUX.2 Max11.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Geography, strict

60.6% · Rank 4 of 6

Sample: 66 · Data date: August 26, 2026

1.GPT Image 284.8%
2.Nano Banana Pro75.8%
3.Nano Banana 269.7%
4.GPT Image 1.5, Current model60.6%
5.FLUX.2 Max15.2%
6.Seedream 4.010.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Computer science, strict

36.3% · Rank 4 of 6

Sample: 102 · Data date: August 26, 2026

1.GPT Image 273.5%
2.Nano Banana Pro65.7%
3.Nano Banana 256.9%
4.GPT Image 1.5, Current model36.3%
5.FLUX.2 Max8.8%
6.Seedream 4.06.9%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Engineering, strict

44.1% · Rank 4 of 6

Sample: 111 · Data date: August 26, 2026

1.GPT Image 279.3%
2.Nano Banana Pro71.2%
3.Nano Banana 267.6%
4.GPT Image 1.5, Current model44.1%
5.Seedream 4.011.7%
6.FLUX.2 Max10.8%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Economics, strict

42.9% · Rank 4 of 6

Sample: 77 · Data date: August 26, 2026

1.Nano Banana Pro88.3%
2.GPT Image 283.1%
3.Nano Banana 263.6%
4.GPT Image 1.5, Current model42.9%
5.Seedream 4.05.2%
6.FLUX.2 Max2.6%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Music, strict

29.2% · Rank 4 of 6

Sample: 65 · Data date: August 26, 2026

1.GPT Image 264.6%
2.Nano Banana Pro61.5%
3.Nano Banana 250.8%
4.GPT Image 1.5, Current model29.2%
5.FLUX.2 Max6.2%
6.Seedream 4.00%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam History, strict

51.2% · Rank 4 of 6

Sample: 41 · Data date: August 26, 2026

1.Nano Banana Pro97.6%
2.GPT Image 282.9%
2.Nano Banana 282.9%
4.GPT Image 1.5, Current model51.2%
5.FLUX.2 Max7.3%
5.Seedream 4.07.3%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam overall, strict

43.2% · Rank 4 of 6

Sample: 1,000 · Data date: August 26, 2026

1.GPT Image 274.6%
2.Nano Banana Pro72.7%
3.Nano Banana 264.1%
4.GPT Image 1.5, Current model43.2%
5.FLUX.2 Max8.5%
6.Seedream 4.07.2%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Mathematics, relaxed

65.8% · Rank 4 of 6

Sample: 151 · Data date: August 26, 2026

1.Nano Banana 287.8%
2.Nano Banana Pro86.3%
3.GPT Image 285.2%
4.GPT Image 1.5, Current model65.8%
5.FLUX.2 Max49.1%
6.Seedream 4.039.8%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Physics, relaxed

85.4% · Rank 4 of 6

Sample: 113 · Data date: August 26, 2026

1.Nano Banana 295.7%
2.GPT Image 295.6%
3.Nano Banana Pro95.1%
4.GPT Image 1.5, Current model85.4%
5.FLUX.2 Max63.2%
6.Seedream 4.049%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Chemistry, relaxed

78.1% · Rank 4 of 6

Sample: 118 · Data date: August 26, 2026

1.GPT Image 292%
2.Nano Banana 290%
3.Nano Banana Pro88.7%
4.GPT Image 1.5, Current model78.1%
5.FLUX.2 Max54%
6.Seedream 4.046.1%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Biology, relaxed

91.9% · Rank 4 of 6

Sample: 156 · Data date: August 26, 2026

1.GPT Image 297.5%
2.Nano Banana Pro95.9%
3.Nano Banana 295.2%
4.GPT Image 1.5, Current model91.9%
5.FLUX.2 Max74.5%
6.Seedream 4.071%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Geography, relaxed

92.5% · Rank 4 of 6

Sample: 66 · Data date: August 26, 2026

1.GPT Image 297.6%
2.Nano Banana Pro96.5%
3.Nano Banana 294.8%
4.GPT Image 1.5, Current model92.5%
5.FLUX.2 Max76.3%
6.Seedream 4.065.1%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Computer science, relaxed

75.8% · Rank 4 of 6

Sample: 102 · Data date: August 26, 2026

1.GPT Image 293.3%
2.Nano Banana Pro91.7%
3.Nano Banana 288.8%
4.GPT Image 1.5, Current model75.8%
5.FLUX.2 Max56.5%
6.Seedream 4.052.2%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Engineering, relaxed

86.4% · Rank 4 of 6

Sample: 111 · Data date: August 26, 2026

1.GPT Image 296.5%
2.Nano Banana 295.8%
3.Nano Banana Pro95.1%
4.GPT Image 1.5, Current model86.4%
5.FLUX.2 Max68.9%
6.Seedream 4.060%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Economics, relaxed

85.5% · Rank 4 of 6

Sample: 77 · Data date: August 26, 2026

1.GPT Image 297.7%
2.Nano Banana Pro97.2%
3.Nano Banana 294.2%
4.GPT Image 1.5, Current model85.5%
5.FLUX.2 Max61.5%
6.Seedream 4.056%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam Music, relaxed

70.8% · Rank 4 of 6

Sample: 65 · Data date: August 26, 2026

1.Nano Banana Pro91%
2.GPT Image 289.1%
3.Nano Banana 286.9%
4.GPT Image 1.5, Current model70.8%
5.FLUX.2 Max47%
6.Seedream 4.034.5%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam History, relaxed

90.9% · Rank 4 of 6

Sample: 41 · Data date: August 26, 2026

1.Nano Banana Pro99.9%
2.Nano Banana 297.3%
3.GPT Image 297.1%
4.GPT Image 1.5, Current model90.9%
5.FLUX.2 Max68%
6.Seedream 4.056.7%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

GenExam overall, relaxed

82.3% · Rank 4 of 6

Sample: 1,000 · Data date: August 26, 2026

1.GPT Image 293.8%
2.Nano Banana Pro93.7%
3.Nano Banana 292.6%
4.GPT Image 1.5, Current model82.3%
5.FLUX.2 Max61.9%
6.Seedream 4.053%

6 of 6 model versions shown in this chart. A higher value ranks first.

1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.

Source: GenExamParticipants: 6

Qwen Image Bench quality

55.14 points · Rank 3 of 5

Data date: August 26, 2026

1.GPT Image 258.65 points
2.Nano Banana Pro55.67 points
3.GPT Image 1.5, Current model55.14 points
4.Nano Banana 254.77 points
5.Qwen Image 2.0 Pro54.39 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench aesthetics

60.88 points · Rank 3 of 5

Data date: August 26, 2026

1.GPT Image 267.53 points
2.Nano Banana 261.08 points
3.GPT Image 1.5, Current model60.88 points
4.Nano Banana Pro60.26 points
5.Qwen Image 2.0 Pro58.67 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench alignment

61.72 points · Rank 3 of 5

Data date: August 26, 2026

1.GPT Image 265.85 points
2.Nano Banana 262.4 points
3.GPT Image 1.5, Current model61.72 points
4.Nano Banana Pro61.25 points
5.Qwen Image 2.0 Pro59.28 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench real-world fidelity

53.95 points · Rank 4 of 5

Data date: August 26, 2026

1.GPT Image 257.38 points
2.Nano Banana 254.28 points
3.Nano Banana Pro54.07 points
4.GPT Image 1.5, Current model53.95 points
5.Qwen Image 2.0 Pro51.83 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench creative generation

66.35 points · Rank 3 of 5

Data date: August 26, 2026

1.GPT Image 275.23 points
2.Nano Banana 267.05 points
3.GPT Image 1.5, Current model66.35 points
4.Nano Banana Pro66.23 points
5.Qwen Image 2.0 Pro64.94 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

Qwen Image Bench overall

59.65 points · Rank 3 of 5

Data date: August 26, 2026

1.GPT Image 264.69 points
2.Nano Banana 259.82 points
3.GPT Image 1.5, Current model59.65 points
4.Nano Banana Pro59.45 points
5.Qwen Image 2.0 Pro57.84 points

5 of 5 model versions shown in this chart. A higher value ranks first.

1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.

WISE Verified overall

82.5% · Rank 2 of 4

Sample: 1,000 · Data date: August 26, 2026

1.Nano Banana Pro87.6%
2.GPT Image 1.5, Current model82.5%
3.FLUX.2 Klein 9B44%
4.Stable Diffusion 3.5 Large40.4%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE culture

89% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro89.75%
2.GPT Image 1.5, Current model89%
3.FLUX.2 Klein 9B49%
3.Stable Diffusion 3.5 Large49%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE time

69.17% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro81.67%
2.GPT Image 1.5, Current model69.17%
3.Stable Diffusion 3.5 Large40.83%
4.FLUX.2 Klein 9B39.17%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE space

88.33% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro93.33%
2.GPT Image 1.5, Current model88.33%
3.FLUX.2 Klein 9B55%
4.Stable Diffusion 3.5 Large44.17%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE biology

80% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro81.67%
2.GPT Image 1.5, Current model80%
3.FLUX.2 Klein 9B38.33%
4.Stable Diffusion 3.5 Large30%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE physics

75.83% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro86.67%
2.GPT Image 1.5, Current model75.83%
3.FLUX.2 Klein 9B48.33%
4.Stable Diffusion 3.5 Large37.5%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

WISE chemistry

77.5% · Rank 2 of 4

Data date: August 26, 2026

1.Nano Banana Pro87.5%
2.GPT Image 1.5, Current model77.5%
3.FLUX.2 Klein 9B22.5%
4.Stable Diffusion 3.5 Large20.83%

4 of 4 model versions shown in this chart. A higher value ranks first.

1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.

Source: WISE VerifiedParticipants: 4

GRADE reasoning

54.5 points · Rank 4 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 282.2 points
2.Nano Banana Pro77.5 points
3.Nano Banana 272.6 points
4.GPT Image 1.5, Current model54.5 points
5.FLUX.2 Max47.8 points
6.FLUX.2 Pro38.9 points
7.Seedream 4.032.4 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE consistency

82.3 points · Rank 4 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 294.4 points
2.Nano Banana Pro89.5 points
3.Nano Banana 286.4 points
4.GPT Image 1.5, Current model82.3 points
5.FLUX.2 Max67.2 points
6.FLUX.2 Pro55.5 points
7.Seedream 4.053.2 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE readability

90.7 points · Rank 4 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 298.8 points
2.Nano Banana 295.9 points
3.Nano Banana Pro95.8 points
4.GPT Image 1.5, Current model90.7 points
5.Seedream 4.077 points
6.FLUX.2 Pro70.3 points
7.FLUX.2 Max68.6 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GRADE accuracy

16 points · Rank 4 of 7

Sample: 520 · Data date: August 26, 2026

1.GPT Image 256 points
2.Nano Banana Pro46.2 points
3.Nano Banana 239.6 points
4.GPT Image 1.5, Current model16 points
5.FLUX.2 Max11.9 points
6.FLUX.2 Pro4.4 points
7.Seedream 4.03.1 points

7 of 7 model versions shown in this chart. A higher value ranks first.

520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.

Source: GRADEParticipants: 7

GEBench Chinese, single-step

83.79 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro84.5 points
2.GPT Image 1.5, Current model83.79 points
3.FLUX.2 Pro68.83 points
4.Wan 2.6 Text to Image64.2 points
5.Seedream 4.062.04 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench Chinese, multi-step

56.97 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro68.65 points
2.GPT Image 1.5, Current model56.97 points
3.FLUX.2 Pro55.07 points
4.Wan 2.6 Text to Image50.11 points
5.Seedream 4.048.64 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench Chinese, fictional app

60.11 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro65.75 points
2.GPT Image 1.5, Current model60.11 points
3.FLUX.2 Pro58.13 points
4.Wan 2.6 Text to Image52.72 points
5.Seedream 4.049.28 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench Chinese, real app

55.65 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro64.35 points
2.GPT Image 1.5, Current model55.65 points
3.FLUX.2 Pro55.41 points
4.Seedream 4.050.93 points
5.Wan 2.6 Text to Image50.4 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench Chinese, grounding

53.33 points · Rank 4 of 5

Data date: August 26, 2026

1.Nano Banana Pro64.83 points
2.Wan 2.6 Text to Image59.58 points
3.Seedream 4.053.53 points
4.GPT Image 1.5, Current model53.33 points
5.FLUX.2 Pro50.24 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench Chinese, overall

63.22 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro69.62 points
2.GPT Image 1.5, Current model63.22 points
3.FLUX.2 Pro57.54 points
4.Wan 2.6 Text to Image55.4 points
5.Seedream 4.052.88 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, single-step

80.8 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro84.32 points
2.GPT Image 1.5, Current model80.8 points
3.FLUX.2 Pro61 points
4.Wan 2.6 Text to Image60.17 points
5.Seedream 4.053.28 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, multi-step

58.87 points · Rank 2 of 5

Data date: August 26, 2026

1.Nano Banana Pro69.51 points
2.GPT Image 1.5, Current model58.87 points
3.FLUX.2 Pro52.17 points
4.Wan 2.6 Text to Image44.36 points
5.Seedream 4.037.57 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, fictional app

63.68 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 1.5, Current model63.68 points
2.FLUX.2 Pro49.92 points
3.Wan 2.6 Text to Image49.55 points
4.Seedream 4.047.92 points
5.Nano Banana Pro46.33 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, real app

58.93 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 1.5, Current model58.93 points
2.Seedream 4.049.36 points
3.Nano Banana Pro47.2 points
4.FLUX.2 Pro47.16 points
5.Wan 2.6 Text to Image44.8 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, grounding

49.23 points · Rank 3 of 5

Data date: August 26, 2026

1.Nano Banana Pro58.64 points
2.Wan 2.6 Text to Image53.36 points
3.GPT Image 1.5, Current model49.23 points
4.FLUX.2 Pro45.67 points
5.Seedream 4.044.17 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GEBench English, overall

63.16 points · Rank 1 of 5

Data date: August 26, 2026

1.GPT Image 1.5, Current model63.16 points
2.Nano Banana Pro61.2 points
3.FLUX.2 Pro51.18 points
4.Wan 2.6 Text to Image50.45 points
5.Seedream 4.046.46 points

5 of 5 model versions shown in this chart. A higher value ranks first.

A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.

Source: GEBenchParticipants: 5

GenAI-Bench overall preference

1,047 Elo · Rank 2 of 3

Data date: August 26, 2026

1.Nano Banana 21,073 Elo
2.GPT Image 1.5, Current model1,047 Elo
3.Nano Banana Pro1,021 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

GenAI-Bench visual quality

975 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana 21,129 Elo
2.Nano Banana Pro1,043 Elo
3.GPT Image 1.5, Current model975 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Infographics factuality

985 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,102 Elo
2.Nano Banana 21,074 Elo
3.GPT Image 1.5, Current model985 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

General image editing

995 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,051 Elo
2.Nano Banana 21,047 Elo
3.GPT Image 1.5, Current model995 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Character editing

1,025 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,050 Elo
2.Nano Banana 21,049 Elo
3.GPT Image 1.5, Current model1,025 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Creative editing

1,017 Elo · Rank 2 of 3

Data date: August 26, 2026

1.Nano Banana 21,031 Elo
2.GPT Image 1.5, Current model1,017 Elo
3.Nano Banana Pro1,004 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Object and environment editing

976 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,042 Elo
2.Nano Banana 21,018 Elo
3.GPT Image 1.5, Current model976 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Editing with 1-3 input images

1,014 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,056 Elo
2.Nano Banana 21,016 Elo
3.GPT Image 1.5, Current model1,014 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Stylization

996 Elo · Rank 3 of 3

Data date: August 26, 2026

1.Nano Banana Pro1,045 Elo
2.Nano Banana 21,031 Elo
3.GPT Image 1.5, Current model996 Elo

3 of 3 model versions shown in this chart. A higher value ranks first.

Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.

Gradually sample test: Typography and layout

33.3 rank points · Rank 22 of 28

Sample: 1 · Data date: August 18, 2026

1.MAI-Image-2.593.8 rank points
20.GPT Image 235.8 rank points
20.Reve 2.135.8 rank points
22.GPT Image 1.5, Current model33.3 rank points
23.Seedream 4.018.5 rank points
24.Wan 2.6 Text to Image14.8 rank points
28.Stable Diffusion 3.5 Large0 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Product photography

37 rank points · Rank 19 of 28

Sample: 1 · Data date: August 18, 2026

1.GPT Image 295.1 rank points
17.Qwen Image 2.0 Pro42 rank points
18.Midjourney V8.140.7 rank points
19.GPT Image 1.5, Current model37 rank points
20.Nano Banana 229.6 rank points
21.Wan 2.6 Text to Image25.9 rank points
28.Stable Diffusion 3.5 Large3.7 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Character and detail

85.2 rank points · Rank 2 of 28

Sample: 1 · Data date: August 18, 2026

1.MAI-Image-2.5 Flash86.4 rank points
2.GPT Image 1.5, Current model85.2 rank points
3.FLUX.2 Max80.3 rank points
4.HiDream-O1-Image-1.579 rank points
28.Stable Diffusion 3.5 Large3.7 rank points

5 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Gradually sample test: Infographic

49.4 rank points · Rank 14 of 28

Sample: 1 · Data date: August 18, 2026

1.Reve 2.197.5 rank points
12.Nano Banana 2 Lite58 rank points
13.Qwen Image 2.0 Pro53.1 rank points
14.GPT Image 1.5, Current model49.4 rank points
14.Nano Banana 249.4 rank points
14.Wan 2.6 Text to Image49.4 rank points
28.Firefly Image Model 51.2 rank points

7 of 28 model versions shown in this chart. A higher value ranks first.

One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.

Profile

Specifications and access

Published information about this model. Unknown values are not estimated.

Weights
ProprietarySource
Access
API, WebSource
Maximum output
1,536 × 1,024 pixelsSource
Benchmark configuration
HighSource
Generator
gpt-image-1.5Source

Pricing

Published prices

Prices remain tied to their documented unit and source.

Representative price
$0.13 per imageSource
Together AI (openai/gpt-image-1.5)
$0.03 per imageSource

Measurements

Other published benchmarks

The stored dataset does not contain an exactly matching comparison cohort for these values.

LMArena Text to Image: 3D modeling

1,215.44 points

Rank 11 · 14,763 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Art

1,223.91 points

Rank 11 · 19,937 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Cartoon

1,243.08 points

Rank 10 · 55,661 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Commercial design

1,240.91 points

Rank 10 · 55,926 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Overall

1,238.66 points

Rank 11 · 142,427 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Photorealism

1,248.78 points

Rank 11 · 57,631 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Portraits

1,261.42 points

Rank 9 · 28,153 samples · Retrieved August 26, 2026

LMArena

LMArena Text to Image: Text rendering

1,254.59 points

Rank 10 · 51,503 samples · Retrieved August 26, 2026

LMArena

LMArena Image Editing: Multi-image editing

1,341.51 points

Rank 9 · 164,316 samples · Retrieved August 26, 2026

LMArena

LMArena Image Editing: Overall

1,370.21 points

Rank 13 · 531,138 samples · Retrieved August 26, 2026

LMArena

Head-to-head comparisons

Compare this model

Each matchup compares this model with exactly one other model from the same category.

More models

Models from the same selection

All AI models

Evidence

Primary sources and data date

Every statement links to its underlying documentation or leaderboard.

  • OpenAISource
  • Image pricingSource
  • Together AISource
  • Artificial Analysis (retrieved August 26, 2026)Source
  • LMArena (retrieved August 26, 2026)Source
  • LMArena (retrieved August 26, 2026)Source
  • GenExam (retrieved August 26, 2026)Source
  • Qwen Image Bench (retrieved August 26, 2026)Source
  • WISE Verified (retrieved August 26, 2026)Source
  • GRADE (retrieved August 26, 2026)Source
  • GEBench (retrieved August 26, 2026)Source
  • Gemini 3.1 Flash Image model card (retrieved August 26, 2026)Source
  • Gradually-Bildtest (retrieved August 18, 2026)Source