CyberBench v1.1PatchTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-2026-03-05 | | 72.414 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 67,632 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,871 anonymous pairwise votes. | 95% confidence interval: ±3.702The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.082 95% confidence interval: ±4.048The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
CyberBench v1.1PoCTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-2026-03-05 | | 61.017 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 64,479 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,747 anonymous pairwise votes. | 95% confidence interval: ±3.778The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.511 95% confidence interval: ±4.115The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 13,882 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,436 anonymous pairwise votes. | 95% confidence interval: ±6.376The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.526 95% confidence interval: ±6.971The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 13,092 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,641 anonymous pairwise votes. | 95% confidence interval: ±6.485The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.189 95% confidence interval: ±6.982The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,364 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,996 anonymous pairwise votes. | 95% confidence interval: ±9.833The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,471.267 95% confidence interval: ±11.424The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,159 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,992 anonymous pairwise votes. | 95% confidence interval: ±10.001The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.924 95% confidence interval: ±11.42The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 18,919 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,262 anonymous pairwise votes. | 95% confidence interval: ±5.805The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,484.094 95% confidence interval: ±6.23The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 17,577 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,239 anonymous pairwise votes. | 95% confidence interval: ±5.854The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.555 95% confidence interval: ±6.265The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 11,135 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,539 anonymous pairwise votes. | 95% confidence interval: ±6.844The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,407.791 95% confidence interval: ±7.543The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 10,894 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,638 anonymous pairwise votes. | 95% confidence interval: ±6.988The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,410.703 95% confidence interval: ±7.575The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 30,587 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,008 anonymous pairwise votes. | 95% confidence interval: ±4.916The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.92 95% confidence interval: ±5.29The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 28,991 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 22,908 anonymous pairwise votes. | 95% confidence interval: ±4.966The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.94 95% confidence interval: ±5.323The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 14,616 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,179 anonymous pairwise votes. | 95% confidence interval: ±6.282The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,404.567 95% confidence interval: ±6.937The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 13,917 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,428 anonymous pairwise votes. | 95% confidence interval: ±6.441The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.063 95% confidence interval: ±6.896The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 51,980 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,226 anonymous pairwise votes. | 95% confidence interval: ±4.836The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.809 95% confidence interval: ±5.308The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 49,393 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,391 anonymous pairwise votes. | 95% confidence interval: ±4.933The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.008 95% confidence interval: ±5.353The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 6,707 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,392 anonymous pairwise votes. | 95% confidence interval: ±8.255The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.264 95% confidence interval: ±8.778The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 6,334 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,311 anonymous pairwise votes. | 95% confidence interval: ±8.393The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,466.306 95% confidence interval: ±8.867The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,420 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,006 anonymous pairwise votes. | 95% confidence interval: ±14.25The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.483 95% confidence interval: ±15.039The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,403 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,966 anonymous pairwise votes. | 95% confidence interval: ±14.273The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.263 95% confidence interval: ±15.286The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,001 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 843 anonymous pairwise votes. | 95% confidence interval: ±19.823The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,424.769 95% confidence interval: ±21.221The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,092 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 906 anonymous pairwise votes. | 95% confidence interval: ±19.045The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.595 95% confidence interval: ±20.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 44,698 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,753 anonymous pairwise votes. | 95% confidence interval: ±4.475The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.919 95% confidence interval: ±4.767The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 42,103 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,700 anonymous pairwise votes. | 95% confidence interval: ±4.521The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.083 95% confidence interval: ±4.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 21,002 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,176 anonymous pairwise votes. | 95% confidence interval: ±5.628The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,463.433 95% confidence interval: ±6.065The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 19,649 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,039 anonymous pairwise votes. | 95% confidence interval: ±5.688The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,464.61 95% confidence interval: ±6.133The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 23,384 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 17,961 anonymous pairwise votes. | 95% confidence interval: ±5.39The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.187 95% confidence interval: ±5.811The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 21,896 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 18,051 anonymous pairwise votes. | 95% confidence interval: ±5.474The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.413 95% confidence interval: ±5.881The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 681 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 647 anonymous pairwise votes. | 95% confidence interval: ±24.201The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,398.363 95% confidence interval: ±24.906The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 671 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 635 anonymous pairwise votes. | 95% confidence interval: ±24.937The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,424.242 95% confidence interval: ±25.166The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,161 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 898 anonymous pairwise votes. | 95% confidence interval: ±19.269The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,392.925 95% confidence interval: ±20.977The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,107 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 951 anonymous pairwise votes. | 95% confidence interval: ±19.893The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,378.153 95% confidence interval: ±20.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 5,387 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,139 anonymous pairwise votes. | 95% confidence interval: ±9.028The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,447.478 95% confidence interval: ±9.874The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 5,240 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,107 anonymous pairwise votes. | 95% confidence interval: ±9.116The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.491 95% confidence interval: ±9.927The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 11,096 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,423 anonymous pairwise votes. | 95% confidence interval: ±6.725The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,456.345 95% confidence interval: ±7.388The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 10,659 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,509 anonymous pairwise votes. | 95% confidence interval: ±6.842The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.724 95% confidence interval: ±7.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 29,890 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,359 anonymous pairwise votes. | 95% confidence interval: ±5.219The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,447.83 95% confidence interval: ±5.578The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 28,372 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,501 anonymous pairwise votes. | 95% confidence interval: ±5.295The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.636 95% confidence interval: ±5.577The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,551 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,688 anonymous pairwise votes. | 95% confidence interval: ±10.346The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,425.467 95% confidence interval: ±11.801The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,387 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,619 anonymous pairwise votes. | 95% confidence interval: ±10.6The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,441.412 95% confidence interval: ±11.893The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,730 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,975 anonymous pairwise votes. | 95% confidence interval: ±10.457The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.489 95% confidence interval: ±11.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 3,546 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,828 anonymous pairwise votes. | 95% confidence interval: ±10.639The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.003 95% confidence interval: ±11.761The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 5,046 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,699 anonymous pairwise votes. | 95% confidence interval: ±9.337The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.257 95% confidence interval: ±10.503The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 4,826 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,774 anonymous pairwise votes. | 95% confidence interval: ±9.555The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,461.307 95% confidence interval: ±10.453The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 12,664 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,363 anonymous pairwise votes. | 95% confidence interval: ±6.57The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,453.414 95% confidence interval: ±7.284The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 11,982 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,233 anonymous pairwise votes. | 95% confidence interval: ±6.641The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,446.332 95% confidence interval: ±7.34The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 37,043 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,862 anonymous pairwise votes. | 95% confidence interval: ±4.588The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.239 95% confidence interval: ±4.947The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 35,486 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,837 anonymous pairwise votes. | 95% confidence interval: ±4.642The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.125 95% confidence interval: ±5.023The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,304 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,030 anonymous pairwise votes. | 95% confidence interval: ±16.827The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.601 95% confidence interval: ±18.712The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 1,322 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,037 anonymous pairwise votes. | 95% confidence interval: ±16.86The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.977 95% confidence interval: ±18.457The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 7,658 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,750 anonymous pairwise votes. | 95% confidence interval: ±7.678The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.519 95% confidence interval: ±8.502The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 7,218 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,790 anonymous pairwise votes. | 95% confidence interval: ±7.885The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.327 95% confidence interval: ±8.593The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 26,958 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,360 anonymous pairwise votes. | 95% confidence interval: ±5.151The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,470.864 95% confidence interval: ±5.508The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 25,454 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,329 anonymous pairwise votes. | 95% confidence interval: ±5.185The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,472.722 95% confidence interval: ±5.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,067 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,596 anonymous pairwise votes. | 95% confidence interval: ±14.33The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.099 95% confidence interval: ±16.059The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 2,086 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,497 anonymous pairwise votes. | 95% confidence interval: ±14.273The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,429.802 95% confidence interval: ±16.45The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 16,553 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,557 anonymous pairwise votes. | 95% confidence interval: ±5.94The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,415.672 95% confidence interval: ±6.557The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGPT-5.4: Snapshot: gpt-5.4-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGPT-5.4: Style-controlled Arena rating from 16,142 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,763 anonymous pairwise votes. | 95% confidence interval: ±5.995The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.277 95% confidence interval: ±6.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |