LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 21,843 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 11,017 anonymous pairwise votes. | 95% confidence interval: ±5.216The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,426.98 95% confidence interval: ±6.556The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 4,096 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,148 anonymous pairwise votes. | 95% confidence interval: ±9.744The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.113 95% confidence interval: ±13.17The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 1,535 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 500 anonymous pairwise votes. | 95% confidence interval: ±15.959The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,448.299 95% confidence interval: ±26.61The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 6,146 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,079 anonymous pairwise votes. | 95% confidence interval: ±8.334The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,478.845 95% confidence interval: ±11.157The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 4,491 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,862 anonymous pairwise votes. | 95% confidence interval: ±9.77The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,396.971 95% confidence interval: ±14.389The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 9,090 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,299 anonymous pairwise votes. | 95% confidence interval: ±7.047The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.656 95% confidence interval: ±8.682The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 5,663 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,453 anonymous pairwise votes. | 95% confidence interval: ±8.825The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,396.875 95% confidence interval: ±12.491The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 16,468 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 8,299 anonymous pairwise votes. | 95% confidence interval: ±6.788The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,421.17 95% confidence interval: ±8.686The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 2,471 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,110 anonymous pairwise votes. | 95% confidence interval: ±12.303The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.278 95% confidence interval: ±17.831The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 729 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 438 anonymous pairwise votes. | 95% confidence interval: ±22.982The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,436.08 95% confidence interval: ±30.551The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 406 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 170 anonymous pairwise votes. | 95% confidence interval: ±29.398The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,442.607 95% confidence interval: ±44.317The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 14,613 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 7,154 anonymous pairwise votes. | 95% confidence interval: ±6.151The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,445.592 95% confidence interval: ±7.857The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 6,015 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,629 anonymous pairwise votes. | 95% confidence interval: ±8.355The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,458.246 95% confidence interval: ±10.39The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 7,798 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,561 anonymous pairwise votes. | 95% confidence interval: ±7.591The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,422.491 95% confidence interval: ±10.335The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 427 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 178 anonymous pairwise votes. | 95% confidence interval: ±28.759The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,394.42 95% confidence interval: ±44.572The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 1,846 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 839 anonymous pairwise votes. | 95% confidence interval: ±14.333The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.785 95% confidence interval: ±21.302The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 3,594 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,823 anonymous pairwise votes. | 95% confidence interval: ±10.263The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.494 95% confidence interval: ±14.359The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 10,464 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 4,661 anonymous pairwise votes. | 95% confidence interval: ±7.088The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,430.723 95% confidence interval: ±9.428The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 943 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 539 anonymous pairwise votes. | 95% confidence interval: ±19.104The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,429.582 95% confidence interval: ±24.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 1,160 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 597 anonymous pairwise votes. | 95% confidence interval: ±17.619The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,442.192 95% confidence interval: ±23.932The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 1,599 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 802 anonymous pairwise votes. | 95% confidence interval: ±15.438The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.47 95% confidence interval: ±21.791The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 3,662 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,803 anonymous pairwise votes. | 95% confidence interval: ±10.462The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.302 95% confidence interval: ±14.433The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 12,751 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,718 anonymous pairwise votes. | 95% confidence interval: ±6.331The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,406.687 95% confidence interval: ±8.36The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 315 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 226 anonymous pairwise votes. | 95% confidence interval: ±32.378The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,407.955 95% confidence interval: ±38.858The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 2,179 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,143 anonymous pairwise votes. | 95% confidence interval: ±13.064The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,415.692 95% confidence interval: ±17.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 8,836 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 4,320 anonymous pairwise votes. | 95% confidence interval: ±7.204The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,470.76 95% confidence interval: ±9.541The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 554 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 377 anonymous pairwise votes. | 95% confidence interval: ±25.992The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,386.845 95% confidence interval: ±32.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.5 Flash-Lite: Snapshot: gemini-3.5-flash-lite, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.5 Flash-Lite: Style-controlled Arena rating from 5,853 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,734 anonymous pairwise votes. | 95% confidence interval: ±8.642The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,405.944 95% confidence interval: ±11.74The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |