LMArena Text to Image: 3D modeling
1,215.44 points
Rank 11 · 14,763 samples · Retrieved August 26, 2026
LMArenaOpenAI
Category view
This overview uses only published data from matching cohorts. Missing values never change a rank.
This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
1,310 Elo · Rank 4 of 25
Sample: 13,674 · Data date: August 26, 2026
7 of 25 model versions shown in this chart. A higher value ranks first.
1,251 Elo · Rank 4 of 19
Sample: 11,694 · Data date: August 26, 2026
7 of 19 model versions shown in this chart. A higher value ranks first.
26.5% · Rank 4 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
46% · Rank 4 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
39% · Rank 4 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
56.4% · Rank 4 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
60.6% · Rank 4 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
36.3% · Rank 4 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
44.1% · Rank 4 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
42.9% · Rank 4 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
29.2% · Rank 4 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
51.2% · Rank 4 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
43.2% · Rank 4 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
65.8% · Rank 4 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
85.4% · Rank 4 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
78.1% · Rank 4 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
91.9% · Rank 4 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
92.5% · Rank 4 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
75.8% · Rank 4 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
86.4% · Rank 4 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
85.5% · Rank 4 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
70.8% · Rank 4 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
90.9% · Rank 4 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
82.3% · Rank 4 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
55.14 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
60.88 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
61.72 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
53.95 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
66.35 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
59.65 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
82.5% · Rank 2 of 4
Sample: 1,000 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
89% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
69.17% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
88.33% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
80% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
75.83% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
77.5% · Rank 2 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
54.5 points · Rank 4 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
82.3 points · Rank 4 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
90.7 points · Rank 4 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
16 points · Rank 4 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
83.79 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
56.97 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
60.11 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
55.65 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
53.33 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
63.22 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
80.8 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
58.87 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
63.68 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
58.93 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
49.23 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
63.16 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
1,047 Elo · Rank 2 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
975 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
985 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
995 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,025 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,017 Elo · Rank 2 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
976 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,014 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
996 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
33.3 rank points · Rank 22 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
37 rank points · Rank 19 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
85.2 rank points · Rank 2 of 28
Sample: 1 · Data date: August 18, 2026
5 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
49.4 rank points · Rank 14 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
1,215.44 points
Rank 11 · 14,763 samples · Retrieved August 26, 2026
LMArena1,223.91 points
Rank 11 · 19,937 samples · Retrieved August 26, 2026
LMArena1,243.08 points
Rank 10 · 55,661 samples · Retrieved August 26, 2026
LMArena1,240.91 points
Rank 10 · 55,926 samples · Retrieved August 26, 2026
LMArena1,238.66 points
Rank 11 · 142,427 samples · Retrieved August 26, 2026
LMArena1,248.78 points
Rank 11 · 57,631 samples · Retrieved August 26, 2026
LMArena1,261.42 points
Rank 9 · 28,153 samples · Retrieved August 26, 2026
LMArena1,254.59 points
Rank 10 · 51,503 samples · Retrieved August 26, 2026
LMArena1,341.51 points
Rank 9 · 164,316 samples · Retrieved August 26, 2026
LMArena1,370.21 points
Rank 13 · 531,138 samples · Retrieved August 26, 2026
LMArenaHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.