CyberBench v1.1PatchTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 72.414 % | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 46,190 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 51,747 anonymous pairwise votes. | 95% confidence interval: ±3.954The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.511 95% confidence interval: ±4.115The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
CyberBench v1.1PoCTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | 57.627 % | | Vals AISource detailsData as of: 08/03/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 9,088 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 10,641 anonymous pairwise votes. | 95% confidence interval: ±7.031The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.189 95% confidence interval: ±6.982The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 3,491 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,992 anonymous pairwise votes. | 95% confidence interval: ±10.738The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.924 95% confidence interval: ±11.42The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 13,289 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 15,239 anonymous pairwise votes. | 95% confidence interval: ±6.21The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,479.555 95% confidence interval: ±6.265The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 8,808 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,638 anonymous pairwise votes. | 95% confidence interval: ±7.404The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,410.703 95% confidence interval: ±7.575The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 19,425 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 22,908 anonymous pairwise votes. | 95% confidence interval: ±5.278The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,449.94 95% confidence interval: ±5.323The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 11,195 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 11,428 anonymous pairwise votes. | 95% confidence interval: ±6.734The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.063 95% confidence interval: ±6.896The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 34,820 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 39,391 anonymous pairwise votes. | 95% confidence interval: ±5.234The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.008 95% confidence interval: ±5.353The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 5,340 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,311 anonymous pairwise votes. | 95% confidence interval: ±8.777The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,466.306 95% confidence interval: ±8.867The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 1,472 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,966 anonymous pairwise votes. | 95% confidence interval: ±17.019The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.263 95% confidence interval: ±15.286The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 726 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 906 anonymous pairwise votes. | 95% confidence interval: ±22.822The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.595 95% confidence interval: ±20.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 31,069 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 34,700 anonymous pairwise votes. | 95% confidence interval: ±4.663The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.083 95% confidence interval: ±4.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 13,168 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 16,039 anonymous pairwise votes. | 95% confidence interval: ±6.176The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,464.61 95% confidence interval: ±6.133The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 16,713 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 18,051 anonymous pairwise votes. | 95% confidence interval: ±5.696The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.413 95% confidence interval: ±5.881The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 639 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 635 anonymous pairwise votes. | 95% confidence interval: ±25.099The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,424.242 95% confidence interval: ±25.166The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 929 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 951 anonymous pairwise votes. | 95% confidence interval: ±20.203The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,378.153 95% confidence interval: ±20.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 3,917 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 4,107 anonymous pairwise votes. | 95% confidence interval: ±10.17The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,457.491 95% confidence interval: ±9.927The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 7,683 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 8,509 anonymous pairwise votes. | 95% confidence interval: ±7.391The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,459.724 95% confidence interval: ±7.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 21,881 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 23,501 anonymous pairwise votes. | 95% confidence interval: ±5.441The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.636 95% confidence interval: ±5.577The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 2,261 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,619 anonymous pairwise votes. | 95% confidence interval: ±12.768The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,441.412 95% confidence interval: ±11.893The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 2,730 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 2,828 anonymous pairwise votes. | 95% confidence interval: ±11.978The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.003 95% confidence interval: ±11.761The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 3,429 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 3,774 anonymous pairwise votes. | 95% confidence interval: ±10.913The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,461.307 95% confidence interval: ±10.453The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 7,746 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 9,233 anonymous pairwise votes. | 95% confidence interval: ±7.534The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,446.332 95% confidence interval: ±7.34The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 26,754 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 28,837 anonymous pairwise votes. | 95% confidence interval: ±4.816The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.125 95% confidence interval: ±5.023The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 765 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,037 anonymous pairwise votes. | 95% confidence interval: ±22.117The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.977 95% confidence interval: ±18.457The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 5,188 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 5,790 anonymous pairwise votes. | 95% confidence interval: ±8.821The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.327 95% confidence interval: ±8.593The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 18,620 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 21,329 anonymous pairwise votes. | 95% confidence interval: ±5.465The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,472.722 95% confidence interval: ±5.539The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 1,388 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 1,497 anonymous pairwise votes. | 95% confidence interval: ±16.935The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,429.802 95% confidence interval: ±16.45The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash: Snapshot: gemini-3.5-flash-high, Reasoning: high, Harness: Arena text, style control DeepSeek-V4-Flash: Snapshot: deepseek-v4-flash-high-preview, Reasoning: high, Harness: Arena text, style controlGemini 3.5 Flash: Style-controlled Arena rating from 11,905 anonymous pairwise votes.DeepSeek-V4-Flash: Style-controlled Arena rating from 12,763 anonymous pairwise votes. | 95% confidence interval: ±6.479The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.277 95% confidence interval: ±6.549The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |