CyberBenchOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-2026-03-05 | | 66.716 % | Vals AISource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 63,615 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 48,963 anonymous pairwise votes. | 95% confidence interval: ±3.807The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.816 95% confidence interval: ±4.15The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
CyberBenchPatchTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-2026-03-05 | | 72.414 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 60,657 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 48,705 anonymous pairwise votes. | 95% confidence interval: ±3.881The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.462 95% confidence interval: ±4.21The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
CyberBenchPoCTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-2026-03-05 | | 61.017 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 13,060 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,911 anonymous pairwise votes. | 95% confidence interval: ±6.534The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.854 95% confidence interval: ±7.128The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 12,346 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,060 anonymous pairwise votes. | 95% confidence interval: ±6.654The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.968 95% confidence interval: ±7.142The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,684 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,530 anonymous pairwise votes. | 95% confidence interval: ±10.702The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,472.677 95% confidence interval: ±12.483The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,482 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,509 anonymous pairwise votes. | 95% confidence interval: ±10.923The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.62 95% confidence interval: ±12.459The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 17,704 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 14,346 anonymous pairwise votes. | 95% confidence interval: ±5.956The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,483.101 95% confidence interval: ±6.376The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 16,441 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 14,310 anonymous pairwise votes. | 95% confidence interval: ±6.001The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.449 95% confidence interval: ±6.409The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 10,326 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,972 anonymous pairwise votes. | 95% confidence interval: ±7.048The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.18 95% confidence interval: ±7.747The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 10,108 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,029 anonymous pairwise votes. | 95% confidence interval: ±7.189The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.903 95% confidence interval: ±7.781The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 29,058 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,931 anonymous pairwise votes. | 95% confidence interval: ±4.99The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.287 95% confidence interval: ±5.383The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 27,549 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,747 anonymous pairwise votes. | 95% confidence interval: ±5.043The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.894 95% confidence interval: ±5.421The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 13,621 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,423 anonymous pairwise votes. | 95% confidence interval: ±6.441The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,405.349 95% confidence interval: ±7.122The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 12,946 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,655 anonymous pairwise votes. | 95% confidence interval: ±6.606The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,406.152 95% confidence interval: ±7.06The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 48,871 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 36,983 anonymous pairwise votes. | 95% confidence interval: ±4.961The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.271 95% confidence interval: ±5.416The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 46,411 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 37,039 anonymous pairwise votes. | 95% confidence interval: ±5.055The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.7 95% confidence interval: ±5.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 6,107 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,904 anonymous pairwise votes. | 95% confidence interval: ±8.598The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.425 95% confidence interval: ±9.147The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 5,765 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,907 anonymous pairwise votes. | 95% confidence interval: ±8.748The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,467.744 95% confidence interval: ±9.211The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,327 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,909 anonymous pairwise votes. | 95% confidence interval: ±14.655The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.246 95% confidence interval: ±15.542The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,288 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,872 anonymous pairwise votes. | 95% confidence interval: ±14.8The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.512 95% confidence interval: ±15.76The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 964 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 834 anonymous pairwise votes. | 95% confidence interval: ±20.361The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.985 95% confidence interval: ±21.432The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,056 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 881 anonymous pairwise votes. | 95% confidence interval: ±19.432The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.001 95% confidence interval: ±20.992The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 41,790 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 32,576 anonymous pairwise votes. | 95% confidence interval: ±4.591The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.19 95% confidence interval: ±4.879The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 39,303 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 32,479 anonymous pairwise votes. | 95% confidence interval: ±4.628The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.492 95% confidence interval: ±4.934The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 19,850 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,327 anonymous pairwise votes. | 95% confidence interval: ±5.732The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,464.306 95% confidence interval: ±6.191The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 18,552 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,135 anonymous pairwise votes. | 95% confidence interval: ±5.797The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,465.194 95% confidence interval: ±6.261The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 21,742 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,772 anonymous pairwise votes. | 95% confidence interval: ±5.529The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.139 95% confidence interval: ±5.955The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 20,284 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,813 anonymous pairwise votes. | 95% confidence interval: ±5.598The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.908 95% confidence interval: ±6.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 652 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 633 anonymous pairwise votes. | 95% confidence interval: ±24.878The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,397.389 95% confidence interval: ±25.351The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 644 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 600 anonymous pairwise votes. | 95% confidence interval: ±25.629The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.217 95% confidence interval: ±25.99The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,132 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 867 anonymous pairwise votes. | 95% confidence interval: ±19.624The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,394.225 95% confidence interval: ±21.434The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,077 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 909 anonymous pairwise votes. | 95% confidence interval: ±20.212The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,374.943 95% confidence interval: ±21.551The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 5,056 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,892 anonymous pairwise votes. | 95% confidence interval: ±9.309The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.174 95% confidence interval: ±10.185The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,896 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,872 anonymous pairwise votes. | 95% confidence interval: ±9.43The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.091 95% confidence interval: ±10.242The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 10,310 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,776 anonymous pairwise votes. | 95% confidence interval: ±6.907The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.608 95% confidence interval: ±7.61The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 9,887 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 7,968 anonymous pairwise votes. | 95% confidence interval: ±7.029The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,460.882 95% confidence interval: ±7.588The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 27,399 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,561 anonymous pairwise votes. | 95% confidence interval: ±5.365The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.456 95% confidence interval: ±5.733The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 26,046 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,664 anonymous pairwise votes. | 95% confidence interval: ±5.438The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.036 95% confidence interval: ±5.731The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,352 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,516 anonymous pairwise votes. | 95% confidence interval: ±10.677The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,425.091 95% confidence interval: ±12.209The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,183 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,459 anonymous pairwise votes. | 95% confidence interval: ±10.961The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,441.301 95% confidence interval: ±12.311The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,439 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,746 anonymous pairwise votes. | 95% confidence interval: ±10.952The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.608 95% confidence interval: ±11.935The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,248 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,607 anonymous pairwise votes. | 95% confidence interval: ±11.167The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,443.348 95% confidence interval: ±12.313The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,674 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,429 anonymous pairwise votes. | 95% confidence interval: ±9.706The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.604 95% confidence interval: ±10.945The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,468 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,495 anonymous pairwise votes. | 95% confidence interval: ±9.928The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,462.945 95% confidence interval: ±10.866The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 11,973 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,820 anonymous pairwise votes. | 95% confidence interval: ±6.718The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,452.356 95% confidence interval: ±7.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 11,276 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,674 anonymous pairwise votes. | 95% confidence interval: ±6.791The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.683 95% confidence interval: ±7.508The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 34,555 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 27,030 anonymous pairwise votes. | 95% confidence interval: ±4.704The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.067 95% confidence interval: ±5.056The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 33,108 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 26,956 anonymous pairwise votes. | 95% confidence interval: ±4.748The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.651 95% confidence interval: ±5.123The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,256 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,005 anonymous pairwise votes. | 95% confidence interval: ±17.194The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,439.462 95% confidence interval: ±18.965The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,297 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,014 anonymous pairwise votes. | 95% confidence interval: ±17.11The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,430.358 95% confidence interval: ±18.866The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 6,931 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,250 anonymous pairwise votes. | 95% confidence interval: ±8.018The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.186 95% confidence interval: ±8.845The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 6,518 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,207 anonymous pairwise votes. | 95% confidence interval: ±8.237The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.38 95% confidence interval: ±8.983The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 25,295 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 20,105 anonymous pairwise votes. | 95% confidence interval: ±5.265The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,471.075 95% confidence interval: ±5.634The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 23,828 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 20,031 anonymous pairwise votes. | 95% confidence interval: ±5.304The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,473.269 95% confidence interval: ±5.664The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,989 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,548 anonymous pairwise votes. | 95% confidence interval: ±14.754The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.026 95% confidence interval: ±16.368The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,023 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,443 anonymous pairwise votes. | 95% confidence interval: ±14.616The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.772 95% confidence interval: ±16.888The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 15,387 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,742 anonymous pairwise votes. | 95% confidence interval: ±6.094The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,417.982 95% confidence interval: ±6.712The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 15,037 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,861 anonymous pairwise votes. | 95% confidence interval: ±6.138The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,421.395 95% confidence interval: ±6.717The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |