LMArena Text to Image: 3D modeling
1,149.39 points
Rank 28 · 17,241 samples · Retrieved August 26, 2026
LMArenaBlack Forest Labs
Category view
This overview uses only published data from matching cohorts. Missing values never change a rank.
This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.
Measurements
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
1,208 Elo · Rank 22 of 25
Sample: 8,915 · Data date: August 26, 2026
7 of 25 model versions shown in this chart. A higher value ranks first.
1,170 Elo · Rank 16 of 19
Sample: 11,529 · Data date: August 26, 2026
7 of 19 model versions shown in this chart. A higher value ranks first.
38.9 points · Rank 6 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
55.5 points · Rank 6 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
70.3 points · Rank 6 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
4.4 points · Rank 6 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
68.83 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
55.07 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
58.13 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
55.41 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
50.24 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
57.54 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
61 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
52.17 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
49.92 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
47.16 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
45.67 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
51.18 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
50.6 rank points · Rank 11 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
43.2 rank points · Rank 16 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
70.4 rank points · Rank 8 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
24.7 rank points · Rank 22 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Profile
Published information about this model. Unknown values are not estimated.
Pricing
Prices remain tied to their documented unit and source.
Measurements
The stored dataset does not contain an exactly matching comparison cohort for these values.
1,149.39 points
Rank 28 · 17,241 samples · Retrieved August 26, 2026
LMArena1,163.05 points
Rank 26 · 25,069 samples · Retrieved August 26, 2026
LMArena1,161.02 points
Rank 27 · 66,010 samples · Retrieved August 26, 2026
LMArena1,156.46 points
Rank 26 · 64,392 samples · Retrieved August 26, 2026
LMArena1,154.79 points
Rank 26 · 175,196 samples · Retrieved August 26, 2026
LMArena1,151.5 points
Rank 30 · 74,515 samples · Retrieved August 26, 2026
LMArena1,146.51 points
Rank 31 · 39,262 samples · Retrieved August 26, 2026
LMArena1,156.45 points
Rank 27 · 59,326 samples · Retrieved August 26, 2026
LMArena1,238.24 points
Rank 20 · 171,464 samples · Retrieved August 26, 2026
LMArena1,244.29 points
Rank 30 · 497,149 samples · Retrieved August 26, 2026
LMArenaHead-to-head comparisons
Each matchup compares this model with exactly one other model from the same category.
More models
All AI models
Evidence
Every statement links to its underlying documentation or leaderboard.