CyberBench v1.1PatchTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 72.414 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 57,564 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,871 anonymous pairwise votes. | 95% confidence interval: ±3.905The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.082 95% confidence interval: ±4.048The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
CyberBench v1.1PoCTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | 57.627 % | | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 54,884 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,747 anonymous pairwise votes. | 95% confidence interval: ±3.973The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.511 95% confidence interval: ±4.115The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 11,411 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,436 anonymous pairwise votes. | 95% confidence interval: ±6.73The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.526 95% confidence interval: ±6.971The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 11,076 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,641 anonymous pairwise votes. | 95% confidence interval: ±6.825The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.189 95% confidence interval: ±6.982The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,408 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,996 anonymous pairwise votes. | 95% confidence interval: ±10.761The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,471.267 95% confidence interval: ±11.424The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,360 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,992 anonymous pairwise votes. | 95% confidence interval: ±10.912The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.924 95% confidence interval: ±11.42The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 17,165 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,262 anonymous pairwise votes. | 95% confidence interval: ±5.93The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,484.094 95% confidence interval: ±6.23The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 16,030 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,239 anonymous pairwise votes. | 95% confidence interval: ±6.099The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.555 95% confidence interval: ±6.265The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,697 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,539 anonymous pairwise votes. | 95% confidence interval: ±7.296The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,407.791 95% confidence interval: ±7.543The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,329 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,638 anonymous pairwise votes. | 95% confidence interval: ±7.37The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,410.703 95% confidence interval: ±7.575The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 25,176 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,008 anonymous pairwise votes. | 95% confidence interval: ±5.13The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.92 95% confidence interval: ±5.29The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 23,891 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 22,908 anonymous pairwise votes. | 95% confidence interval: ±5.188The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.94 95% confidence interval: ±5.323The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 12,738 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,179 anonymous pairwise votes. | 95% confidence interval: ±6.63The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,404.567 95% confidence interval: ±6.937The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 12,346 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,428 anonymous pairwise votes. | 95% confidence interval: ±6.685The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.063 95% confidence interval: ±6.896The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 43,577 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,226 anonymous pairwise votes. | 95% confidence interval: ±5.137The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.809 95% confidence interval: ±5.308The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 41,744 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,391 anonymous pairwise votes. | 95% confidence interval: ±5.153The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.008 95% confidence interval: ±5.353The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,925 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,392 anonymous pairwise votes. | 95% confidence interval: ±8.43The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.264 95% confidence interval: ±8.778The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,532 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,311 anonymous pairwise votes. | 95% confidence interval: ±8.612The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,466.306 95% confidence interval: ±8.867The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,215 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,006 anonymous pairwise votes. | 95% confidence interval: ±14.661The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.483 95% confidence interval: ±15.039The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,075 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,966 anonymous pairwise votes. | 95% confidence interval: ±14.782The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.263 95% confidence interval: ±15.286The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,001 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 843 anonymous pairwise votes. | 95% confidence interval: ±19.454The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,424.769 95% confidence interval: ±21.221The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 928 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 906 anonymous pairwise votes. | 95% confidence interval: ±20.112The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.595 95% confidence interval: ±20.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 38,586 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,753 anonymous pairwise votes. | 95% confidence interval: ±4.586The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.919 95% confidence interval: ±4.767The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 36,977 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,700 anonymous pairwise votes. | 95% confidence interval: ±4.651The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.083 95% confidence interval: ±4.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 17,828 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,176 anonymous pairwise votes. | 95% confidence interval: ±5.864The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,463.433 95% confidence interval: ±6.065The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 16,866 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,039 anonymous pairwise votes. | 95% confidence interval: ±5.961The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,464.61 95% confidence interval: ±6.133The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 20,066 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 17,961 anonymous pairwise votes. | 95% confidence interval: ±5.588The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.187 95% confidence interval: ±5.811The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 19,090 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 18,051 anonymous pairwise votes. | 95% confidence interval: ±5.678The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.413 95% confidence interval: ±5.881The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 714 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 647 anonymous pairwise votes. | 95% confidence interval: ±23.983The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,398.363 95% confidence interval: ±24.906The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 677 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 635 anonymous pairwise votes. | 95% confidence interval: ±24.225The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,424.242 95% confidence interval: ±25.166The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,105 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 898 anonymous pairwise votes. | 95% confidence interval: ±19.032The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,392.925 95% confidence interval: ±20.977The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 981 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 951 anonymous pairwise votes. | 95% confidence interval: ±19.922The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,378.153 95% confidence interval: ±20.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 4,506 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,139 anonymous pairwise votes. | 95% confidence interval: ±9.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,447.478 95% confidence interval: ±9.874The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 4,328 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,107 anonymous pairwise votes. | 95% confidence interval: ±9.718The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.491 95% confidence interval: ±9.927The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,173 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,423 anonymous pairwise votes. | 95% confidence interval: ±7.148The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,456.345 95% confidence interval: ±7.388The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 8,814 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,509 anonymous pairwise votes. | 95% confidence interval: ±7.218The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.724 95% confidence interval: ±7.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 25,983 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,359 anonymous pairwise votes. | 95% confidence interval: ±5.337The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,447.83 95% confidence interval: ±5.578The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 24,788 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,501 anonymous pairwise votes. | 95% confidence interval: ±5.421The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.636 95% confidence interval: ±5.577The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,058 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,688 anonymous pairwise votes. | 95% confidence interval: ±11.133The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,425.467 95% confidence interval: ±11.801The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,708 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,619 anonymous pairwise votes. | 95% confidence interval: ±11.797The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,441.412 95% confidence interval: ±11.893The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,296 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,975 anonymous pairwise votes. | 95% confidence interval: ±10.89The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.489 95% confidence interval: ±11.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,993 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,828 anonymous pairwise votes. | 95% confidence interval: ±11.459The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.003 95% confidence interval: ±11.761The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 4,069 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,699 anonymous pairwise votes. | 95% confidence interval: ±10.122The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.257 95% confidence interval: ±10.503The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,861 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,774 anonymous pairwise votes. | 95% confidence interval: ±10.291The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,461.307 95% confidence interval: ±10.453The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 10,367 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,363 anonymous pairwise votes. | 95% confidence interval: ±7.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,453.414 95% confidence interval: ±7.284The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,694 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,233 anonymous pairwise votes. | 95% confidence interval: ±7.159The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,446.332 95% confidence interval: ±7.34The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 32,388 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,862 anonymous pairwise votes. | 95% confidence interval: ±4.725The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.239 95% confidence interval: ±4.947The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 30,989 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,837 anonymous pairwise votes. | 95% confidence interval: ±4.791The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.125 95% confidence interval: ±5.023The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,124 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,030 anonymous pairwise votes. | 95% confidence interval: ±18.108The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.601 95% confidence interval: ±18.712The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,093 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,037 anonymous pairwise votes. | 95% confidence interval: ±18.224The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.977 95% confidence interval: ±18.457The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 6,319 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,750 anonymous pairwise votes. | 95% confidence interval: ±8.237The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.519 95% confidence interval: ±8.502The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 6,162 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,790 anonymous pairwise votes. | 95% confidence interval: ±8.305The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.327 95% confidence interval: ±8.593The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 23,687 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,360 anonymous pairwise votes. | 95% confidence interval: ±5.271The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,470.864 95% confidence interval: ±5.508The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 22,470 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,329 anonymous pairwise votes. | 95% confidence interval: ±5.371The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,472.722 95% confidence interval: ±5.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,755 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,596 anonymous pairwise votes. | 95% confidence interval: ±15.298The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.099 95% confidence interval: ±16.059The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,719 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,497 anonymous pairwise votes. | 95% confidence interval: ±15.428The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,429.802 95% confidence interval: ±16.45The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 14,117 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,557 anonymous pairwise votes. | 95% confidence interval: ±6.338The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,415.672 95% confidence interval: ±6.557The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 13,716 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,763 anonymous pairwise votes. | 95% confidence interval: ±6.399The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.277 95% confidence interval: ±6.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |