LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 65,336 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 11,017 anonymous pairwise votes. | 1,413.556 95% confidence interval: ±3.08The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±6.556The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 12,392 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,148 anonymous pairwise votes. | 1,412.721 95% confidence interval: ±5.917The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±13.17The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 4,483 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 500 anonymous pairwise votes. | 1,431.759 95% confidence interval: ±9.282The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±26.61The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 16,301 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,079 anonymous pairwise votes. | 1,468.587 95% confidence interval: ±5.241The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±11.157The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 10,954 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,862 anonymous pairwise votes. | 1,375.017 95% confidence interval: ±6.318The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±14.389The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 27,473 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,299 anonymous pairwise votes. | 1,426.818 95% confidence interval: ±4.282The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±8.682The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 13,723 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,453 anonymous pairwise votes. | 1,373.228 95% confidence interval: ±5.71The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±12.491The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 47,430 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 8,299 anonymous pairwise votes. | 1,400.876 95% confidence interval: ±4.294The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±8.686The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 5,705 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,110 anonymous pairwise votes. | 1,423.349 95% confidence interval: ±8.346The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±17.831The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 1,858 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 438 anonymous pairwise votes. | 1,434.923 95% confidence interval: ±15.523The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±30.551The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 1,080 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 170 anonymous pairwise votes. | 1,404.512 95% confidence interval: ±19.03The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±44.317The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 39,237 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 7,154 anonymous pairwise votes. | 1,431.026 95% confidence interval: ±3.848The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±7.857The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 16,526 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,629 anonymous pairwise votes. | 1,438.338 95% confidence interval: ±5.247The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±10.39The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 20,253 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,561 anonymous pairwise votes. | 1,399.916 95% confidence interval: ±4.811The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±10.335The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 1,118 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 178 anonymous pairwise votes. | 1,355.204 95% confidence interval: ±19.178The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±44.572The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 4,880 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 839 anonymous pairwise votes. | 1,420.347 95% confidence interval: ±8.946The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±21.302The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 10,595 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,823 anonymous pairwise votes. | 1,427.433 95% confidence interval: ±6.173The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±14.359The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 23,193 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 4,661 anonymous pairwise votes. | 1,416.784 95% confidence interval: ±4.746The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±9.428The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 3,755 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 539 anonymous pairwise votes. | 1,402.864 95% confidence interval: ±9.902The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±24.872The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 3,175 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 597 anonymous pairwise votes. | 1,417.813 95% confidence interval: ±10.98The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±23.932The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 4,505 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 802 anonymous pairwise votes. | 95% confidence interval: ±9.388The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,423.47 95% confidence interval: ±21.791The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 11,201 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,803 anonymous pairwise votes. | 1,420.712 95% confidence interval: ±6.145The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±14.433The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 37,859 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,718 anonymous pairwise votes. | 1,397.941 95% confidence interval: ±3.805The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±8.36The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 1,582 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 226 anonymous pairwise votes. | 95% confidence interval: ±15.021The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,407.955 95% confidence interval: ±38.858The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 7,429 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,143 anonymous pairwise votes. | 1,409.803 95% confidence interval: ±7.118The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±17.832The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 24,350 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 4,320 anonymous pairwise votes. | 1,456.163 95% confidence interval: ±4.498The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±9.541The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 1,998 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 377 anonymous pairwise votes. | 95% confidence interval: ±14.343The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,386.845 95% confidence interval: ±32.418The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterMistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlMistral Large 3: Style-controlled Arena rating from 15,693 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,734 anonymous pairwise votes. | 1,390.534 95% confidence interval: ±5.355The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 95% confidence interval: ±11.74The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/01/2026 |