Code MigrationOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 5.127 % | Vals AISource detailsData as of: 09/27/2026 |
CorpFin v2OverallTest detailsNo documented difference in the test setupVersion: 2Metric: accuracyScoring: Higher is better | | 58.78 % | Vals AISource detailsData as of: 08/12/2026 |
EMBOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 11.185 % | Vals AISource detailsData as of: 09/27/2026 |
Finance Agent (v2)OverallTest detailsNo documented difference in the test setupVersion: 2Metric: accuracyScoring: Higher is better | | 32.103 % | Vals AISource detailsData as of: 09/27/2026 |
GPQA DiamondOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 34.848 % | Vals AISource detailsData as of: 09/01/2026 |
Harvey's Legal Agent BenchmarkOverall · Task fully resolvedTest detailsNo documented difference in the test setupVersion: 1Metric: task resolution rateScoring: Higher is better | 0 % | | Vals AISource detailsData as of: 09/27/2026 |
Legal Research BenchOverall · All-passTest detailsNo documented difference in the test setupVersion: 1Metric: all-pass rateScoring: Higher is better | | 9.135 % | Vals AISource detailsData as of: 09/27/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 119,196 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 11,693 anonymous pairwise votes. | 95% confidence interval: ±3.034The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,426.128 95% confidence interval: ±6.42The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
MedCodeOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 33.752 % | Vals AISource detailsData as of: 09/26/2026 |
MedScribeOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 67.728 % | Vals AISource detailsData as of: 09/26/2026 |
MMLU-ProOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 75.335 % | Vals AISource detailsData as of: 09/01/2026 |
MortgageTaxOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 28.895 % | Vals AISource detailsData as of: 09/01/2026 |
ProofBenchOverallTest detailsNo documented difference in the test setupVersion: 1.1Metric: accuracyScoring: Higher is better | | 9 % | Vals AISource detailsData as of: 09/28/2026 |
SAGEOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 37.613 % | Vals AISource detailsData as of: 09/26/2026 |
SWE-bench VerifiedOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 66.4 % | Vals AISource detailsData as of: 09/01/2026 |
Tax Agent BenchOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 27.291 % | Vals AISource detailsData as of: 09/27/2026 |
TaxEval v2OverallTest detailsNo documented difference in the test setupVersion: 2Metric: accuracyScoring: Higher is better | | 67.988 % | Vals AISource detailsData as of: 09/01/2026 |
Terminal-Bench 2.0OverallTest detailsNo documented difference in the test setupVersion: 2.0Metric: accuracyScoring: Higher is better | | 30.337 % | Vals AISource detailsData as of: 06/04/2026 |
Terminal-Bench 2.1OverallTest detailsNo documented difference in the test setupVersion: 2.1Metric: accuracyScoring: Higher is better | | 38.951 % | Vals AISource detailsData as of: 09/27/2026 |
Vals IndexOverallTest detailsNo documented difference in the test setupVersion: 2Metric: weighted index scoreScoring: Higher is better | | 17.947 % | Vals AISource detailsData as of: 09/27/2026 |
Vals Multimodal IndexOverallTest detailsNo documented difference in the test setupVersion: 1.2Metric: weighted index scoreScoring: Higher is better | | 34.768 % | Vals AISource detailsData as of: 08/11/2026 |
Vibe Code Bench v1.1OverallTest detailsNo documented difference in the test setupVersion: 1.1Metric: accuracyScoring: Higher is betterSettings: Harness: OpenHands | | 2.888 % | Vals AISource detailsData as of: 09/27/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 23,398 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,256 anonymous pairwise votes. | 95% confidence interval: ±5.145The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,430.173 95% confidence interval: ±12.827The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 8,183 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 621 anonymous pairwise votes. | 95% confidence interval: ±7.531The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,446.524 95% confidence interval: ±23.931The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 33,049 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,282 anonymous pairwise votes. | 95% confidence interval: ±4.686The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,477.898 95% confidence interval: ±10.854The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 21,305 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,999 anonymous pairwise votes. | 95% confidence interval: ±5.535The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,397.828 95% confidence interval: ±13.94The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 52,351 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,562 anonymous pairwise votes. | 95% confidence interval: ±4.036The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,442.84 95% confidence interval: ±8.53The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 27,013 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,618 anonymous pairwise votes. | 95% confidence interval: ±5.112The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,394.24 95% confidence interval: ±12.151The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 89,549 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 8,856 anonymous pairwise votes. | 95% confidence interval: ±4.1The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,419.562 95% confidence interval: ±8.482The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 12,336 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,205 anonymous pairwise votes. | 95% confidence interval: ±6.452The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,433.7 95% confidence interval: ±17.235The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 3,917 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 453 anonymous pairwise votes. | 95% confidence interval: ±11.678The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.855 95% confidence interval: ±29.943The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupNo clear leadVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 1,979 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 173 anonymous pairwise votes. | 95% confidence interval: ±14.89The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,447.96 95% confidence interval: ±43.885The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 77,865 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 7,627 anonymous pairwise votes. | 95% confidence interval: ±3.69The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,444.677 95% confidence interval: ±7.686The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 35,115 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,821 anonymous pairwise votes. | 95% confidence interval: ±4.646The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,455.806 95% confidence interval: ±10.148The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 40,860 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 3,827 anonymous pairwise votes. | 95% confidence interval: ±4.393The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,420.763 95% confidence interval: ±10.011The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 2,108 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 188 anonymous pairwise votes. | 95% confidence interval: ±14.845The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,392.677 95% confidence interval: ±43.291The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 9,836 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 893 anonymous pairwise votes. | 95% confidence interval: ±7.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,437.755 95% confidence interval: ±20.721The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 19,920 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,939 anonymous pairwise votes. | 95% confidence interval: ±5.358The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,435.825 95% confidence interval: ±13.971The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 52,750 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 5,065 anonymous pairwise votes. | 95% confidence interval: ±4.296The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.117 95% confidence interval: ±9.119The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 6,192 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 565 anonymous pairwise votes. | 95% confidence interval: ±8.229The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.606 95% confidence interval: ±24.143The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 6,656 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 653 anonymous pairwise votes. | 95% confidence interval: ±8.227The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,438.481 95% confidence interval: ±22.896The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 8,930 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 844 anonymous pairwise votes. | 95% confidence interval: ±7.351The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,421.161 95% confidence interval: ±21.257The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 20,555 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,925 anonymous pairwise votes. | 95% confidence interval: ±5.404The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,432.256 95% confidence interval: ±14.007The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 66,835 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 6,131 anonymous pairwise votes. | 95% confidence interval: ±3.741The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,406.139 95% confidence interval: ±8.159The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 2,318 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 235 anonymous pairwise votes. | 95% confidence interval: ±12.764The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,415.979 95% confidence interval: ±38.135The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 13,461 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 1,242 anonymous pairwise votes. | 95% confidence interval: ±6.15The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,410.34 95% confidence interval: ±17.157The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 47,251 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 4,587 anonymous pairwise votes. | 95% confidence interval: ±4.177The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,468.806 95% confidence interval: ±9.304The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 3,879 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 393 anonymous pairwise votes. | 95% confidence interval: ±10.938The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,392.394 95% confidence interval: ±31.573The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Medium 3.5: Snapshot: mistral-medium-3.5, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 29,992 anonymous pairwise votes.Mistral Medium 3.5: Style-controlled Arena rating from 2,919 anonymous pairwise votes. | 95% confidence interval: ±4.855The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,403.766 95% confidence interval: ±11.409The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |