LMArena Text to Image: 3D modeling
1,230.04 points
Rank 7 · 13,757 samples · Retrieved August 26, 2026
LMArenaCategory view
This overview uses only published data from matching cohorts. Missing values never change a rank.
This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
1,298 Elo · Rank 6 of 25
Sample: 13,840 · Data date: August 26, 2026
7 of 25 model versions shown in this chart. A higher value ranks first.
1,246 Elo · Rank 7 of 19
Sample: 11,029 · Data date: August 26, 2026
7 of 19 model versions shown in this chart. A higher value ranks first.
55.6% · Rank 2 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
75.2% · Rank 2 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
60.2% · Rank 2 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
75.6% · Rank 2 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
75.8% · Rank 2 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
65.7% · Rank 2 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
71.2% · Rank 2 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
88.3% · Rank 1 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
61.5% · Rank 2 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
97.6% · Rank 1 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
72.7% · Rank 2 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
86.3% · Rank 2 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
95.1% · Rank 3 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
88.7% · Rank 3 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
95.9% · Rank 2 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
96.5% · Rank 2 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
91.7% · Rank 2 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
95.1% · Rank 3 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
97.2% · Rank 2 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
91% · Rank 1 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
99.9% · Rank 1 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
93.7% · Rank 2 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
55.67 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
60.26 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
61.25 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
54.07 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
66.23 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
59.45 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
1,000 prompts, five top-level dimensions, and 56 detailed facets. Published top-five scores use the same deterministic judge configuration.
87.6% · Rank 1 of 4
Sample: 1,000 · Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
89.75% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
81.67% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
93.33% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
81.67% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
86.67% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
87.5% · Rank 1 of 4
Data date: August 26, 2026
4 of 4 model versions shown in this chart. A higher value ranks first.
1,000 knowledge and realism tasks. Binary item judgments are averaged by category, and the overall score follows the published weighting.
77.5 points · Rank 2 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
89.5 points · Rank 2 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
95.8 points · Rank 3 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
46.2 points · Rank 2 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
84.5 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
68.65 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
65.75 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
64.35 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
64.83 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
69.62 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
84.32 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
69.51 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
46.33 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
47.2 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
58.64 points · Rank 1 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
61.2 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
1,021 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,043 Elo · Rank 2 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,102 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,051 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,050 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,004 Elo · Rank 3 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,042 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,056 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
1,045 Elo · Rank 1 of 3
Data date: August 26, 2026
3 of 3 model versions shown in this chart. A higher value ranks first.
Google's model card reports Elo scores with uncertainty intervals for text-to-image generation and image editing. These rows use standard Gemini 3.1 Flash Image without tools.
93.8 rank points · Rank 1 of 28
Sample: 1 · Data date: August 18, 2026
5 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
82.7 rank points · Rank 4 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
21 rank points · Rank 24 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
42 rank points · Rank 17 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
1,230.04 points
Rank 7 · 13,757 samples · Retrieved August 26, 2026
LMArena1,224.75 points
Rank 10 · 19,671 samples · Retrieved August 26, 2026
LMArena1,240.57 points
Rank 11 · 56,045 samples · Retrieved August 26, 2026
LMArena1,245.1 points
Rank 9 · 50,186 samples · Retrieved August 26, 2026
LMArena1,245.59 points
Rank 10 · 138,756 samples · Retrieved August 26, 2026
LMArena1,260.12 points
Rank 7 · 59,290 samples · Retrieved August 26, 2026
LMArena1,258.04 points
Rank 11 · 31,463 samples · Retrieved August 26, 2026
LMArena1,269.2 points
Rank 9 · 46,529 samples · Retrieved August 26, 2026
LMArena1,363.81 points
Rank 6 · 189,338 samples · Retrieved August 26, 2026
LMArena1,389.6 points
Rank 8 · 517,418 samples · Retrieved August 26, 2026
LMArenaHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.