CyberBenchOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 66.716 % | Vals AISource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 54,243 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 48,963 anonymous pairwise votes. | 95% confidence interval: ±4.009The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.816 95% confidence interval: ±4.15The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
CyberBenchPatchTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 72.414 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 51,599 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 48,705 anonymous pairwise votes. | 95% confidence interval: ±4.08The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.462 95% confidence interval: ±4.21The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
CyberBenchPoCTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | 57.627 % | | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 10,742 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,911 anonymous pairwise votes. | 95% confidence interval: ±6.908The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.854 95% confidence interval: ±7.128The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 10,411 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,060 anonymous pairwise votes. | 95% confidence interval: ±6.986The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.968 95% confidence interval: ±7.142The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,871 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,530 anonymous pairwise votes. | 95% confidence interval: ±11.839The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,472.677 95% confidence interval: ±12.483The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,757 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,509 anonymous pairwise votes. | 95% confidence interval: ±12.143The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.62 95% confidence interval: ±12.459The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 16,165 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 14,346 anonymous pairwise votes. | 95% confidence interval: ±6.076The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,483.101 95% confidence interval: ±6.376The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 15,113 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 14,310 anonymous pairwise votes. | 95% confidence interval: ±6.226The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.449 95% confidence interval: ±6.409The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,051 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,972 anonymous pairwise votes. | 95% confidence interval: ±7.495The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.18 95% confidence interval: ±7.747The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 8,660 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,029 anonymous pairwise votes. | 95% confidence interval: ±7.584The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.903 95% confidence interval: ±7.781The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 23,924 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,931 anonymous pairwise votes. | 95% confidence interval: ±5.227The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.287 95% confidence interval: ±5.383The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 22,700 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,747 anonymous pairwise votes. | 95% confidence interval: ±5.28The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.894 95% confidence interval: ±5.421The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 11,880 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,423 anonymous pairwise votes. | 95% confidence interval: ±6.804The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,405.349 95% confidence interval: ±7.122The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 11,440 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,655 anonymous pairwise votes. | 95% confidence interval: ±6.877The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,406.152 95% confidence interval: ±7.06The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 41,089 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 36,983 anonymous pairwise votes. | 95% confidence interval: ±5.248The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.271 95% confidence interval: ±5.416The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 39,187 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 37,039 anonymous pairwise votes. | 95% confidence interval: ±5.263The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.7 95% confidence interval: ±5.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,439 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,904 anonymous pairwise votes. | 95% confidence interval: ±8.815The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.425 95% confidence interval: ±9.147The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,063 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,907 anonymous pairwise votes. | 95% confidence interval: ±8.974The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,467.744 95% confidence interval: ±9.211The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,116 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,909 anonymous pairwise votes. | 95% confidence interval: ±15.107The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.246 95% confidence interval: ±15.542The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,986 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,872 anonymous pairwise votes. | 95% confidence interval: ±15.212The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.512 95% confidence interval: ±15.76The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 968 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 834 anonymous pairwise votes. | 95% confidence interval: ±19.953The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.985 95% confidence interval: ±21.432The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 892 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 881 anonymous pairwise votes. | 95% confidence interval: ±20.588The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.001 95% confidence interval: ±20.992The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 36,098 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 32,576 anonymous pairwise votes. | 95% confidence interval: ±4.693The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.19 95% confidence interval: ±4.879The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 34,551 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 32,479 anonymous pairwise votes. | 95% confidence interval: ±4.754The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.492 95% confidence interval: ±4.934The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 16,819 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,327 anonymous pairwise votes. | 95% confidence interval: ±6.001The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,464.306 95% confidence interval: ±6.191The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 15,908 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,135 anonymous pairwise votes. | 95% confidence interval: ±6.095The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,465.194 95% confidence interval: ±6.261The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 18,672 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,772 anonymous pairwise votes. | 95% confidence interval: ±5.73The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.139 95% confidence interval: ±5.955The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 17,742 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,813 anonymous pairwise votes. | 95% confidence interval: ±5.821The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.908 95% confidence interval: ±6.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 696 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 633 anonymous pairwise votes. | 95% confidence interval: ±24.329The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,397.389 95% confidence interval: ±25.351The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 649 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 600 anonymous pairwise votes. | 95% confidence interval: ±24.768The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.217 95% confidence interval: ±25.99The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,069 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 867 anonymous pairwise votes. | 95% confidence interval: ±19.478The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,394.225 95% confidence interval: ±21.434The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 960 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 909 anonymous pairwise votes. | 95% confidence interval: ±20.291The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,374.943 95% confidence interval: ±21.551The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 4,235 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,892 anonymous pairwise votes. | 95% confidence interval: ±9.866The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.174 95% confidence interval: ±10.185The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 4,061 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,872 anonymous pairwise votes. | 95% confidence interval: ±10.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.091 95% confidence interval: ±10.242The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 8,607 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,776 anonymous pairwise votes. | 95% confidence interval: ±7.337The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.608 95% confidence interval: ±7.61The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 8,161 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,968 anonymous pairwise votes. | 95% confidence interval: ±7.429The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,460.882 95% confidence interval: ±7.588The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 23,973 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,561 anonymous pairwise votes. | 95% confidence interval: ±5.5The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.456 95% confidence interval: ±5.733The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 22,832 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,664 anonymous pairwise votes. | 95% confidence interval: ±5.577The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.036 95% confidence interval: ±5.731The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,868 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,516 anonymous pairwise votes. | 95% confidence interval: ±11.525The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,425.091 95% confidence interval: ±12.209The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,505 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,459 anonymous pairwise votes. | 95% confidence interval: ±12.249The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,441.301 95% confidence interval: ±12.311The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,010 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,746 anonymous pairwise votes. | 95% confidence interval: ±11.504The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.608 95% confidence interval: ±11.935The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 2,693 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,607 anonymous pairwise votes. | 95% confidence interval: ±12.095The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,443.348 95% confidence interval: ±12.313The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,762 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,429 anonymous pairwise votes. | 95% confidence interval: ±10.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.604 95% confidence interval: ±10.945The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 3,584 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,495 anonymous pairwise votes. | 95% confidence interval: ±10.705The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,462.945 95% confidence interval: ±10.866The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,810 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,820 anonymous pairwise votes. | 95% confidence interval: ±7.2The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,452.356 95% confidence interval: ±7.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 9,139 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,674 anonymous pairwise votes. | 95% confidence interval: ±7.327The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.683 95% confidence interval: ±7.508The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 30,319 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 27,030 anonymous pairwise votes. | 95% confidence interval: ±4.83The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.067 95% confidence interval: ±5.056The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 28,895 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 26,956 anonymous pairwise votes. | 95% confidence interval: ±4.904The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.651 95% confidence interval: ±5.123The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,085 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,005 anonymous pairwise votes. | 95% confidence interval: ±18.414The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,439.462 95% confidence interval: ±18.965The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,058 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,014 anonymous pairwise votes. | 95% confidence interval: ±18.524The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,430.358 95% confidence interval: ±18.866The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,710 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,250 anonymous pairwise votes. | 95% confidence interval: ±8.611The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.186 95% confidence interval: ±8.845The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 5,539 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,207 anonymous pairwise votes. | 95% confidence interval: ±8.697The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.38 95% confidence interval: ±8.983The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 22,292 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 20,105 anonymous pairwise votes. | 95% confidence interval: ±5.393The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,471.075 95% confidence interval: ±5.634The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 21,173 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 20,031 anonymous pairwise votes. | 95% confidence interval: ±5.481The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,473.269 95% confidence interval: ±5.664The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,687 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,548 anonymous pairwise votes. | 95% confidence interval: ±15.717The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.026 95% confidence interval: ±16.368The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 1,656 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,443 anonymous pairwise votes. | 95% confidence interval: ±15.837The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.772 95% confidence interval: ±16.888The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 13,175 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,742 anonymous pairwise votes. | 95% confidence interval: ±6.494The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,417.982 95% confidence interval: ±6.712The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterDeepSeek-V4-Pro: Snapshot: deepseek-v4-pro-high-preview, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlDeepSeek-V4-Pro: Style-controlled Arena rating from 12,751 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,861 anonymous pairwise votes. | 95% confidence interval: ±6.569The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,421.395 95% confidence interval: ±6.717The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |