Seedream 4.0
ByteDance
- Released
- September 2025
- Data date
- August 15, 2026
Category view
Position within the category
This overview uses only published data from matching cohorts. Missing values never change a rank.
This position gives equal weight to 4 fixed Gradually image tests. One archived first output contributes for each model and task.
- Position
- Rank 25 of 28
- Index score
- 29.9 / 100
- Coverage
- 4 / 4
Leaderboard
Measurements
Comparable benchmark results
Each chart contains exactly one source, one measurement series, and one stored comparison cohort. Bars show the position. The measured value appears on the right.
Text to Image Arena
1,227 Elo · Rank 13 of 25
Sample: 4,603 · Data date: August 26, 2026
7 of 25 model versions shown in this chart. A higher value ranks first.
Image Editing Arena
1,185 Elo · Rank 15 of 19
Sample: 10,984 · Data date: August 26, 2026
7 of 19 model versions shown in this chart. A higher value ranks first.
GenExam Mathematics, strict
2.6% · Rank 6 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Physics, strict
3.5% · Rank 6 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Chemistry, strict
5.9% · Rank 6 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Biology, strict
18.6% · Rank 5 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Geography, strict
10.6% · Rank 6 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Computer science, strict
6.9% · Rank 6 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Engineering, strict
11.7% · Rank 5 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Economics, strict
5.2% · Rank 5 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Music, strict
0% · Rank 6 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam History, strict
7.3% · Rank 5 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam overall, strict
7.2% · Rank 6 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Mathematics, relaxed
39.8% · Rank 6 of 6
Sample: 151 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Physics, relaxed
49% · Rank 6 of 6
Sample: 113 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Chemistry, relaxed
46.1% · Rank 6 of 6
Sample: 118 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Biology, relaxed
71% · Rank 6 of 6
Sample: 156 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Geography, relaxed
65.1% · Rank 6 of 6
Sample: 66 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Computer science, relaxed
52.2% · Rank 6 of 6
Sample: 102 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Engineering, relaxed
60% · Rank 6 of 6
Sample: 111 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Economics, relaxed
56% · Rank 6 of 6
Sample: 77 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam Music, relaxed
34.5% · Rank 6 of 6
Sample: 65 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam History, relaxed
56.7% · Rank 6 of 6
Sample: 41 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GenExam overall, relaxed
53% · Rank 6 of 6
Sample: 1,000 · Data date: August 26, 2026
6 of 6 model versions shown in this chart. A higher value ranks first.
1,000 multidisciplinary drawing tasks with reference images and fine-grained scoring points. Strict and relaxed scoring remain separate.
GRADE reasoning
32.4 points · Rank 7 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
GRADE consistency
53.2 points · Rank 7 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
GRADE readability
77 points · Rank 5 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
GRADE accuracy
3.1 points · Rank 7 of 7
Sample: 520 · Data date: August 26, 2026
7 of 7 model versions shown in this chart. A higher value ranks first.
520 scientific image-editing tasks across ten academic domains. GRADE scores reasoning, consistency, readability, and domain accuracy separately.
GEBench Chinese, single-step
62.04 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench Chinese, multi-step
48.64 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench Chinese, fictional app
49.28 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench Chinese, real app
50.93 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench Chinese, grounding
53.53 points · Rank 3 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench Chinese, overall
52.88 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, single-step
53.28 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, multi-step
37.57 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, fictional app
47.92 points · Rank 4 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, real app
49.36 points · Rank 2 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, grounding
44.17 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
GEBench English, overall
46.46 points · Rank 5 of 5
Data date: August 26, 2026
5 of 5 model versions shown in this chart. A higher value ranks first.
A bilingual graphical-user-interface benchmark. Chinese and English subsets and the single-step, multi-step, fictional-app, real-app, and grounding categories remain separate.
Gradually sample test: Typography and layout
18.5 rank points · Rank 23 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Gradually sample test: Product photography
19.7 rank points · Rank 23 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Gradually sample test: Character and detail
13.6 rank points · Rank 27 of 28
Sample: 1 · Data date: August 18, 2026
5 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Gradually sample test: Infographic
67.9 rank points · Rank 9 of 28
Sample: 1 · Data date: August 18, 2026
7 of 28 model versions shown in this chart. A higher value ranks first.
One archived, unfiltered first output per model. All 28 images are ranked in three model-blind orderings. The 0-100 value normalizes the mean rank. This remains an n = 1 sample, and an automated judge may have visual preferences of its own.
Profile
Specifications and access
Published information about this model. Unknown values are not estimated.
Pricing
Published prices
Prices remain tied to their documented unit and source.
Head-to-head comparisons
Compare this model
Each matchup compares this model with exactly one other model from the same category.
More models
Models from the same selection
All AI models
Evidence
Primary sources and data date
Every statement links to its underlying documentation or leaderboard.