Skip to main content

Compare LLMs without mixing incompatible test setups

Compare 36 leading language models across 630 head-to-head matchups. Every result includes shared benchmark cohorts, current API pricing, key specifications, and source links.

models
36
head-to-head matchups
630
active benchmarks
83
active results
14,315

Cross-benchmark view

Overall LLM leaderboard

The index uses 4 benchmark cohorts with comparable documented test setups. It converts the comparison-model scores into rank percentiles and averages them with equal weight. Only the 26 models with results in every cohort receive an overall rank. Models without complete coverage are omitted, and missing data is not treated as a penalty.

1.Claude Opus 588.9 / 100, Rank 1
2.Gemini 3.8 Flash88.5 / 100
3.Claude Fable 583.2 / 100
4.GPT-5.6 Sol81.7 / 100
5.Gemini 3.7 Flash76.4 / 100
6.Kimi K374.5 / 100
7.Grok 4.670.2 / 100
8.GPT-5.6 Luna66.8 / 100
9.Gemini 3.6 Flash63.0 / 100
10.GPT-5.6 Terra59.1 / 100
11.Gemini 3.5 Flash58.7 / 100
12.Claude Sonnet 557.7 / 100
13.GPT-5.557.2 / 100
14.GLM-5.354.8 / 100
15.Grok 4.546.2 / 100
16.Gemini 3.1 Pro Preview41.3 / 100
17.GLM-5.3-Flash35.6 / 100
18.GLM-5.234.1 / 100
19.DeepSeek-V4-Pro24.5 / 100
20.Kimi K2.623.1 / 100
21.GPT-5.4 mini20.2 / 100
22.Gemini 3.5 Flash-Lite16.8 / 100
23.GLM-5.115.4 / 100
24.MiMo-V2.513.5 / 100
25.Grok 4.312.5 / 100
25.MiMo-V2.5-Pro12.5 / 100

This is a transparent orientation index, not an additional benchmark. It does not include price, speed, context length, or results outside the listed 4 benchmark cohorts.

Source entries with incompatible documented configurations are excluded from the composite instead of being reduced to their strongest result.

Current leaders

Which LLMs lead the most important benchmarks?

These five charts keep the benchmarks separate. Each ranking uses one published leaderboard cohort, one task, and one metric. Models without an unambiguous published result in that cohort are omitted instead of receiving an estimated score.

Overall preference

LMArena Text, Style Control

Style-controlled rating from anonymous pairwise comparisons.

1.Claude Fable 51,507.4 points, Rank 1
2.Kimi K31,489 points
3.Gemini 3.1 Pro Preview1,486.7 points
4.GPT-5.6 Sol1,482.9 points
5.Gemini 3.6 Flash1,480.2 points
6.GLM-5.21,471.9 points
7.Grok 4.51,470 points
8.MiMo-V2.5-Pro1,468.3 points
9.GPT-5.6 Terra1,466.2 points
10.GLM-5.11,466 points
11.Claude Sonnet 51,462.2 points
12.Grok 4.61,461.1 points
13.Kimi K2.61,460.6 points
14.Gemini 3.5 Flash-Lite1,456.8 points
15.GPT-5.6 Luna1,452.4 points
16.GPT-5.4 mini1,448.1 points
17.Grok 4.31,442.6 points
18.MiMo-V2.51,433.8 points
19.Mistral Medium 3.51,427 points
20.Mistral Large 31,413.6 points
LMArenaUpdated Sep 1, 2026 · 20 of 36 comparison models with an unambiguous result

The published leaderboard documents model-specific test settings. 6 comparison models have more than one incompatible setup. They are omitted from this ranking instead of selecting the best result.

Reasoning

GPQA Diamond

Difficult scientific questions from the GPQA Diamond dataset.

1.Gemini 3.1 Pro Preview95.5 %, Rank 1
2.GPT-5.6 Sol95.2 %
3.Grok 4.694.7 %
4.Gemini 3.8 Flash94.4 %
5.Gemini 3.7 Flash93.9 %
6.Claude Opus 593.4 %
6.Gemini 3.6 Flash93.4 %
8.Claude Fable 593.2 %
8.GPT-5.593.2 %
10.Grok 4.592.9 %
10.Kimi K392.9 %
12.Gemini 3.5 Flash92.7 %
13.GPT-5.491.7 %
13.GPT-5.6 Luna91.7 %
15.Grok 4.391.4 %
16.GPT-5.6 Terra90.9 %
17.DeepSeek-V4-Pro89.4 %
18.Kimi K2.689.1 %
19.Claude Sonnet 588.9 %
20.GLM-5.388.1 %
21.GLM-5.3-Flash86.4 %
22.GLM-5.285.6 %
23.GLM-5.184.5 %
24.Gemini 3.5 Flash-Lite83.8 %
25.GPT-5.4 mini83.1 %
26.MiMo-V2.5-Pro82.6 %
27.MiMo-V2.581.6 %
Vals AIUpdated Sep 1, 2026 · 27 of 36 comparison models with an unambiguous result

Coding

Vibe Code Bench v1.1

Practical programming tasks from Vibe Code Bench v1.1.

1.Claude Fable 590.4 %, Rank 1
2.GPT-6 Astra89.6 %
3.Claude Opus 588.4 %
4.Kimi K385 %
5.Claude Sonnet 581.3 %
6.GPT-5.6 Sol80.5 %
7.Gemini 3.8 Flash78.7 %
8.GLM-5.378.1 %
9.GPT-5.6 Luna77.1 %
10.Grok 4.676.2 %
11.GPT-5.6 Terra74.6 %
12.Gemini 3.7 Flash70.4 %
13.GPT-5.569.8 %
14.Grok 4.569 %
15.Gemini 3.6 Flash64 %
16.GLM-5.264 %
17.DeepSeek-V4-Pro49.9 %
18.Gemini 3.5 Flash48.7 %
19.GPT-5.4 mini48 %
20.MiMo-V2.542.2 %
21.Kimi K2.637.9 %
22.Gemini 3.5 Flash-Lite37.2 %
23.MiMo-V2.5-Pro34.1 %
24.Gemini 3.1 Pro Preview32 %
25.GLM-5.131.5 %
26.GLM-5.3-Flash30.8 %
27.Grok 4.319.4 %
Vals AIUpdated Sep 3, 2026 · 27 of 36 comparison models with an unambiguous result

The published leaderboard documents model-specific test settings. 1 comparison model has more than one incompatible setup. It is omitted from this ranking instead of selecting the best result.

Agents

Terminal-Bench 2.1

Multi-step tasks completed in a real terminal environment.

1.GPT-6 Astra87.3 %, Rank 1
2.GPT-5.6 Sol85.8 %
3.Claude Opus 584.6 %
4.Gemini 3.8 Flash81.3 %
5.Kimi K380.9 %
6.Claude Fable 580.5 %
7.GPT-5.6 Luna79 %
8.Grok 4.678.3 %
9.Gemini 3.7 Flash77.5 %
9.GPT-5.6 Terra77.5 %
11.GPT-5.576.4 %
12.Claude Sonnet 574.5 %
13.Gemini 3.5 Flash74.2 %
14.Gemini 3.6 Flash73.8 %
15.GLM-5.371.5 %
16.Gemini 3.1 Pro Preview70.8 %
17.GLM-5.267.8 %
17.Grok 4.567.8 %
19.GLM-5.3-Flash62.9 %
20.MiMo-V2.560.7 %
21.MiMo-V2.5-Pro57.3 %
22.GLM-5.156.9 %
23.GPT-5.4 mini54.7 %
24.Kimi K2.653.6 %
25.DeepSeek-V4-Pro50.2 %
25.Gemini 3.5 Flash-Lite50.2 %
27.Grok 4.341.9 %
Vals AIUpdated Sep 3, 2026 · 27 of 36 comparison models with an unambiguous result

Finance

Finance Agent (v2)

Research and analysis tasks from Finance Agent v2.

1.Gemini 3.8 Flash61.4 %, Rank 1
2.Gemini 3.7 Flash59 %
3.Claude Opus 558.6 %
4.Gemini 3.5 Flash57.9 %
5.GLM-5.3-Flash57.9 %
6.Claude Fable 556.3 %
7.Gemini 3.6 Flash56.3 %
8.GLM-5.355.8 %
9.GPT-5.6 Luna55 %
10.GPT-5.6 Terra54.4 %
11.Kimi K354.4 %
12.Claude Sonnet 553.9 %
13.GPT-5.6 Sol53.8 %
14.Grok 4.653.7 %
15.GPT-6 Astra53.5 %
16.GPT-5.551.8 %
17.GLM-5.249.7 %
18.Grok 4.548.3 %
19.Gemini 3.5 Flash-Lite47.4 %
20.GPT-5.4 mini45.4 %
21.Kimi K2.644.9 %
22.GLM-5.144.8 %
23.DeepSeek-V4-Pro44.1 %
24.Gemini 3.1 Pro Preview43 %
25.MiMo-V2.5-Pro41.5 %
26.Grok 4.337.7 %
27.MiMo-V2.536.7 %
Vals AIUpdated Sep 3, 2026 · 27 of 36 comparison models with an unambiguous result

A first-place result applies only to the named benchmark. It is not an overall quality score and says nothing by itself about price, speed, context length, or your own workload.

Decision criteria

What this LLM comparison lets you check

Benchmarks
Shared tasks and versions for reasoning, coding, agents, research, and other capabilities.
API pricing
Input, output, cache-read, and cache-write pricing where the provider publishes it.
Context
Published context windows and maximum output limits without imprecise rounding.
Architecture
Known parameters, active parameters, dense or MoE architecture, and open-weight status.
Availability
API access, lifecycle status, release date, and documented web-search support.
Sources
A retrieval or verification date and a direct source for each volatile result.

Need a broader directory? Browse the open-source LLM directory. For workload-specific cost estimates, use the API cost calculator.

Model selection

The 36 most relevant models

The database focuses on current frontier models and relevant lower-cost options. Superseded generations are excluded.

For the full technical profile of every model, browse the AI model directory.

Historical reference

A comparison from ChatGPT's early model generations

The current comparison matrix intentionally focuses on supported models. This historical reference keeps the documented differences between GPT-3.5 and GPT-4 available for existing applications and migration context.

GPT-3.5 vs. GPT-4: What's the Difference?

Frequently asked questions about the LLM comparison

How the benchmark matching, winners, and data freshness work.

Changelog

The latest updates and improvements to our LLM comparison
v1.9September 4, 2026

GPT-6 Astra added

  • Added GPT-6 Astra with official specifications, pricing, context window, availability, and source links
  • Expanded the comparison to 36 models and 630 head-to-head matchups
v1.8September 3, 2026

Muse Spark 1.3 and Qwen 3.8 Max 0902 added

  • Added Muse Spark 1.3 with current pricing, context, availability, and source links
  • Updated Qwen 3.8 Max to the 0902 snapshot and added it to the comparison roster
  • Expanded the comparison to 35 models and 595 head-to-head matchups
v1.7September 2, 2026

Gemini 3.8 Flash added

  • Added Gemini 3.8 Flash with official specifications, pricing, knowledge cutoff, and source links
  • Expanded the comparison to 33 models and 528 head-to-head matchups
v1.6September 2, 2026

Claude 5.1 and GLM-5.3 added

  • Added GLM-5.3 with official specifications, pricing, availability, and source links
  • Expanded the comparison to 32 models and 496 head-to-head matchups, including Claude Fable 5.1 and Mythos 5.1
v1.5August 29, 2026

GLM-5.3-Flash and cleaner leaderboards

  • Added GLM-5.3-Flash with official specifications, pricing, and source links
  • Expanded the comparison to 29 models and 406 head-to-head matchups
  • Removed models without a published score from leaderboard rows
v1.4August 28, 2026

Complete and cross-benchmark leaderboards

  • Added all 28 comparison models to every individual benchmark leaderboard
  • Added an orientation index from four setup-comparable benchmark cohorts
  • Excluded ambiguous model variants and kept missing results visible without estimation
v1.3August 24, 2026

Related matchups, provider icons, and expanded FAQ

  • Expanded each LLM matchup page to 14 related matchups and placed the hub button below the list
  • Added local provider icons to all 28 model badges without placeholders
  • Expanded the FAQ with benchmark, pricing, parameter, and radar methodology details
v1.2August 21, 2026

Benchmark leaderboards and capability profiles

  • Added five separate leaderboards for overall preference, reasoning, coding, agents, and finance
  • Added evidence-based capability radar charts to the head-to-head comparison pages
  • Kept every ranking within one published cohort, task, metric, and source snapshot
v1.1August 18, 2026

FAQ, changelog, and internal links

  • Added an FAQ about benchmark coverage, methodology, and database freshness
  • Added this changelog so material updates remain visible
  • Added the hub and all matchup pages to navigation, the sitemap, and related comparisons
v1.0August 17, 2026

Initial release

  • Published 28 current models and 378 head-to-head matchups
  • Combined shared benchmarks, API pricing, specifications, and primary sources
  • Added the comparison picker and ten prioritized matchups