CorpFin v2OverallTest detailsNo documented difference in the test setupVersion: 2Metric: accuracyScoring: Higher is better | | 61.033 % | Vals AISource detailsData as of: 08/12/2026 |
GPQA DiamondOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 68.434 % | Vals AISource detailsData as of: 09/01/2026 |
LegalBenchOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 79.138 % | Vals AISource detailsData as of: 09/27/2026 |
LiveCodeBench v6OverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 55.337 % | Vals AISource detailsData as of: 09/01/2026 |
LMArena Text, Style ControlOverallTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 119,196 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 77,456 anonymous pairwise votes. | 95% confidence interval: ±3.034The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,413.254 95% confidence interval: ±2.912The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
MedQAOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 82.233 % | Vals AISource detailsData as of: 04/16/2026 |
MMLU-ProOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 79.823 % | Vals AISource detailsData as of: 09/01/2026 |
MMMU-ProOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 66.185 % | Vals AISource detailsData as of: 09/01/2026 |
MortgageTaxOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 52.106 % | Vals AISource detailsData as of: 09/01/2026 |
SAGEOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 24.595 % | Vals AISource detailsData as of: 09/26/2026 |
SWE-bench VerifiedOverallTest detailsNo documented difference in the test setupVersion: 1Metric: accuracyScoring: Higher is better | | 41.4 % | Vals AISource detailsData as of: 09/01/2026 |
TaxEval v2OverallTest detailsNo documented difference in the test setupVersion: 2Metric: accuracyScoring: Higher is better | 72.882 % | | Vals AISource detailsData as of: 09/01/2026 |
Terminal-Bench 2.0OverallTest detailsNo documented difference in the test setupVersion: 2.0Metric: accuracyScoring: Higher is better | | 8.989 % | Vals AISource detailsData as of: 06/04/2026 |
LMArena Text, Style ControlBusiness, management and financeTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 23,398 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 14,597 anonymous pairwise votes. | 95% confidence interval: ±5.145The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,409.446 95% confidence interval: ±5.509The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlChineseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 8,183 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 5,845 anonymous pairwise votes. | 95% confidence interval: ±7.531The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,434.058 95% confidence interval: ±8.22The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCodingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 33,049 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 19,574 anonymous pairwise votes. | 95% confidence interval: ±4.686The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,469.182 95% confidence interval: ±4.902The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlCreative writingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 21,305 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 13,448 anonymous pairwise votes. | 95% confidence interval: ±5.535The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,372.648 95% confidence interval: ±5.824The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 52,351 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 32,060 anonymous pairwise votes. | 95% confidence interval: ±4.036The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,426.615 95% confidence interval: ±4.05The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlEntertainment, sports and mediaTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 27,013 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 16,716 anonymous pairwise votes. | 95% confidence interval: ±5.112The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,371.316 95% confidence interval: ±5.299The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExcluding tiesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 89,549 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 56,501 anonymous pairwise votes. | 95% confidence interval: ±4.1The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,400.18 95% confidence interval: ±4.058The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlExpert promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 12,336 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 7,295 anonymous pairwise votes. | 95% confidence interval: ±6.452The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,427.543 95% confidence interval: ±7.482The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlFrenchTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 3,917 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 2,181 anonymous pairwise votes. | 95% confidence interval: ±11.678The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,432.389 95% confidence interval: ±14.389The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlGermanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 1,979 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 1,257 anonymous pairwise votes. | 95% confidence interval: ±14.89The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,409.467 95% confidence interval: ±17.833The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard promptsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 77,865 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 47,316 anonymous pairwise votes. | 95% confidence interval: ±3.69The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,431.274 95% confidence interval: ±3.63The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlHard prompts, EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 35,115 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 19,638 anonymous pairwise votes. | 95% confidence interval: ±4.646The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,439.274 95% confidence interval: ±4.91The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlInstruction followingTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 40,860 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 24,861 anonymous pairwise votes. | 95% confidence interval: ±4.393The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,400.115 95% confidence interval: ±4.491The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlJapaneseTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 1,365 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 830 anonymous pairwise votes. | 95% confidence interval: ±18.293The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,382.659 95% confidence interval: ±22.106The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlKoreanTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 2,108 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 1,318 anonymous pairwise votes. | 95% confidence interval: ±14.845The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,356.821 95% confidence interval: ±17.794The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLegal and governmentTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 9,836 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 5,879 anonymous pairwise votes. | 95% confidence interval: ±7.024The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,421.357 95% confidence interval: ±8.211The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLife, physical and social scienceTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 19,920 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 12,690 anonymous pairwise votes. | 95% confidence interval: ±5.358The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,428.685 95% confidence interval: ±5.734The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlLonger queriesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 52,750 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 29,146 anonymous pairwise votes. | 95% confidence interval: ±4.296The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,417.29 95% confidence interval: ±4.417The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 6,192 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 4,411 anonymous pairwise votes. | 95% confidence interval: ±8.229The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,401.493 95% confidence interval: ±9.177The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMathematical professionsTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 6,656 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 4,047 anonymous pairwise votes. | 95% confidence interval: ±8.227The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,413.935 95% confidence interval: ±9.781The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMedicine and healthcareTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 8,930 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 5,443 anonymous pairwise votes. | 95% confidence interval: ±7.351The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,430.076 95% confidence interval: ±8.59The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlMulti-turnTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 20,555 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 13,164 anonymous pairwise votes. | 95% confidence interval: ±5.404The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,419.382 95% confidence interval: ±5.74The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlNon-EnglishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 66,835 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 45,384 anonymous pairwise votes. | 95% confidence interval: ±3.741The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,397.089 95% confidence interval: ±3.587The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlPolishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 2,318 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 1,754 anonymous pairwise votes. | 95% confidence interval: ±12.764The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,414.116 95% confidence interval: ±14.298The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlRussianTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 13,461 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 9,292 anonymous pairwise votes. | 95% confidence interval: ±6.15The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,408.043 95% confidence interval: ±6.505The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSoftware and IT servicesTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 47,251 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 29,068 anonymous pairwise votes. | 95% confidence interval: ±4.177The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,455.584 95% confidence interval: ±4.237The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlSpanishTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 3,879 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 2,328 anonymous pairwise votes. | 95% confidence interval: ±10.938The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,405.285 95% confidence interval: ±13.338The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |
LMArena Text, Style ControlWriting, literature and languageTest detailsNo documented difference in the test setupVersion: style-control-v1Metric: style-controlled Bradley-Terry ratingScoring: Higher is betterGemini 3.1 Pro Preview: Snapshot: gemini-3.1-pro-preview, Harness: Arena text, style control Mistral Large 3: Snapshot: mistral-large-3, Harness: Arena text, style controlGemini 3.1 Pro Preview: Style-controlled Arena rating from 29,992 anonymous pairwise votes.Mistral Large 3: Style-controlled Arena rating from 19,002 anonymous pairwise votes. | 95% confidence interval: ±4.855The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | 1,388.294 95% confidence interval: ±4.988The Arena score measures human preference rather than a fixed capability test. The confidence interval and model configuration are part of the result. | LMArenaSource detailsData as of: 09/25/2026 |