Skip to main content

Compare AI image models, model by model

36 image generation models and 630 head-to-head matchups. Each page shows only exact shared evidence. The fixed Gradually sample leaderboard currently covers 30 models, with separate research benchmarks and preference arenas where available.

image models
36
head-to-head matchups
630
benchmark results
656
separate benchmark tasks
78

Current leaderboard

AI image model leaderboard

This equal-weight sample ranking averages four fixed tasks covering typography, product photography, character detail, and infographic design. 30 of 36 current comparison models contribute one archived first output per task. New checkpoints remain outside the ranking until they complete the same protocol. It remains a documented n = 1 sample per model and task.

Not yet tested in this sample: Midjourney V8.2, Qwen-Image-2.1, MAI-Image-2.6, MAI-Image-2.6 Flash, Muse Image, and Grok Imagine Image 2.0. Their current standing appears in the LMArena ranking below where available.

1.Krea 2 Large75.6 / 100, Rank 1
2.GPT Image 274.2 / 100
3.MAI-Image-2.5 Flash73.0 / 100
4.Krea 2 Medium Turbo65.2 / 100
5.Luma Uni-1.1 Max63.8 / 100
6.MAI-Image-2.561.8 / 100
7.GPT Image 2.5 Sunburst60.9 / 100
8.Reve 2.160.1 / 100
9.Seedream 5.0 Pro57.2 / 100
10.HiDream-O1-Image-1.556.3 / 100
11.GPT Image 2.5 Flare56.3 / 100
12.HunyuanImage 3.0 Instruct55.8 / 100
13.FLUX.2 Max55.8 / 100
14.Nano Banana Pro53.8 / 100
15.Nano Banana 2 Lite53.2 / 100
16.Recraft V4.1 Utility51.4 / 100
17.Nano Banana 251.1 / 100
18.FLUX.2 Flex49.4 / 100
19.GPT Image 1.548.3 / 100
20.Qwen Image 2.0 Pro47.7 / 100
21.Firefly Image Model 546.0 / 100
21.FLUX.2 Pro46.0 / 100
23.Ideogram 4.045.4 / 100
24.Wan 2.6 Text to Image41.4 / 100
25.Seedream 4.037.7 / 100
26.Grok Imagine Image Quality32.5 / 100
27.Midjourney V8.128.5 / 100
28.FLUX.2 Klein 9B24.4 / 100
29.HiDream-O1-Image19.3 / 100
30.Stable Diffusion 3.5 Large8.3 / 100

Blind preference votes

LMArena text-to-image ranking

LMArena lets visitors choose between two anonymous images for the same prompt. The rating aggregates many thousands of these votes across all kinds of motifs. Our sample above scores four fixed tasks with one image each. Both lists measure something different, so their leaders can differ.

1.GPT Image 2.5 Sunburst1,421, Rank 1
2.GPT Image 21,381
3.Reve 2.11,301
4.Nano Banana 21,261
5.Seedream 5.0 Pro1,257
6.MAI-Image-2.51,254
7.Nano Banana 2 Lite1,250
8.Nano Banana Pro1,246
9.GPT Image 1.51,239
10.Qwen Image 2.0 Pro1,191
11.FLUX.2 Max1,162
12.FLUX.2 Flex1,157
13.FLUX.2 Pro1,154
14.Wan 2.6 Text to Image1,137
15.HiDream-O1-Image1,117
16.Krea 2 Large1,108
17.FLUX.2 Klein 9B1,070
18.Stable Diffusion 3.5 Large938

18 of 36 comparison models have an overall rating. Positions refer to the models in this catalog. Data retrieved September 19, 2026.LMArena

Model selection

36 important models and configurations

Quality tiers such as Max, Pro, and High remain separate when an arena reports distinct scores and prices. This avoids misleading comparisons between different runtime tiers.

For the full technical profile of every model, browse the AI model directory.

Scope

Image model or AI image generator?

An image model is the underlying generation system, such as GPT Image 2.5 Sunburst, Midjourney V8.2, or Nano Banana 2. A generator is the app around one or more models, including its interface, subscription, editing workflow, storage, and usage rights.

Use this page to compare model performance. Use the generator guide to choose a complete product.

Frequently asked questions about the image model comparison

How arena coverage, sample images, variants, and data freshness work.

It compares separate arena results, task-specific prices, technical specifications, availability, and documented sample images. The hub leaderboard averages only the four fixed Gradually sample tasks. Research benchmarks and preference arenas remain separate.

No. All 630 matchup pages follow the same evidence rules, but a page shows only measurements shared by the exact two checkpoints. The Gradually leaderboard currently covers 30 of 36 models with four archived first outputs. New checkpoints remain unranked until they complete the same protocol.

483 of 630 matchups have at least one shared measurement series in the stored snapshot. 134 of the 147 gaps involve Midjourney V8.2, Firefly Image Model 5, Midjourney V8.1, or Qwen-Image-2.1. These models have no stored arena result. The remaining 13 pairs have results, but not from the same arena and category.

No. Prompts and selection rules are standardized, but provider resolution, quality tiers, prompt processing, and randomness can still differ.

Max, Pro, High, and other tiers can have different scores, prices, and runtime settings. Combining them would hide meaningful differences.

The stored arena snapshot was retrieved on September 28, 2026. Model facts and sample images carry separate source dates.

Changelog

The latest updates and improvements to our image model comparison
v1.5September 22, 2026

35-model roster and re-verified pricing

  • Expanded the catalog to 35 models and 595 head-to-head matchups
  • Re-verified provider prices and lifecycle notices against current provider pages
v1.4September 19, 2026

GPT Image 2.5 Sunburst and Midjourney V8.2

  • Expanded the catalog to 30 models and 435 head-to-head matchups
  • Added exact LMArena results for GPT Image 2.5 Sunburst where available
  • Kept GPT Image 2 and Midjourney V8.1 samples under their original checkpoints
v1.3August 28, 2026

Complete image model leaderboard

  • Added an equal-weight leaderboard for all 28 comparison models
  • Combined only the four shared Gradually sample tasks
  • Kept the n = 1 sample ranking separate from research benchmarks and preference arenas
v1.2August 18, 2026

Complete matchup coverage and research benchmarks

  • Added four shared Gradually rankings from three model-blind orderings of all 28 first outputs
  • Integrated exact model results from GenExam, Qwen Image Bench, and WISE Verified
  • Kept n = 1 sample rankings, academic benchmarks, and preference arenas visibly separate
v1.1August 18, 2026

FAQ, changelog, and internal links

  • Added an FAQ that quantifies shared arena coverage and explains data gaps
  • Added this changelog so material updates remain visible
  • Added the hub and all matchup pages to navigation, the sitemap, and related comparisons
v1.0August 17, 2026

Initial release

  • Published 28 current image models and 378 head-to-head matchups
  • Kept text-to-image generation and image editing in separate benchmark series
  • Combined pricing, specifications, sources, and documented sample images