Compare LLMs without mixing incompatible test setups
Compare 36 leading language models across 630 head-to-head matchups. Every result includes shared benchmark cohorts, current API pricing, key specifications, and source links.
- models
- 36
- head-to-head matchups
- 630
- active benchmarks
- 83
- active results
- 14,315
Cross-benchmark view
Overall LLM leaderboard
The index uses 4 benchmark cohorts with comparable documented test setups. It converts the comparison-model scores into rank percentiles and averages them with equal weight. Only the 26 models with results in every cohort receive an overall rank. Models without complete coverage are omitted, and missing data is not treated as a penalty.
This is a transparent orientation index, not an additional benchmark. It does not include price, speed, context length, or results outside the listed 4 benchmark cohorts.
Source entries with incompatible documented configurations are excluded from the composite instead of being reduced to their strongest result.
Current leaders
Which LLMs lead the most important benchmarks?
These five charts keep the benchmarks separate. Each ranking uses one published leaderboard cohort, one task, and one metric. Models without an unambiguous published result in that cohort are omitted instead of receiving an estimated score.
Overall preference
LMArena Text, Style Control
Style-controlled rating from anonymous pairwise comparisons.
The published leaderboard documents model-specific test settings. 6 comparison models have more than one incompatible setup. They are omitted from this ranking instead of selecting the best result.
Reasoning
GPQA Diamond
Difficult scientific questions from the GPQA Diamond dataset.
Coding
Vibe Code Bench v1.1
Practical programming tasks from Vibe Code Bench v1.1.
The published leaderboard documents model-specific test settings. 1 comparison model has more than one incompatible setup. It is omitted from this ranking instead of selecting the best result.
Agents
Terminal-Bench 2.1
Multi-step tasks completed in a real terminal environment.
Finance
Finance Agent (v2)
Research and analysis tasks from Finance Agent v2.
A first-place result applies only to the named benchmark. It is not an overall quality score and says nothing by itself about price, speed, context length, or your own workload.
Decision criteria
What this LLM comparison lets you check
- Benchmarks
- Shared tasks and versions for reasoning, coding, agents, research, and other capabilities.
- API pricing
- Input, output, cache-read, and cache-write pricing where the provider publishes it.
- Context
- Published context windows and maximum output limits without imprecise rounding.
- Architecture
- Known parameters, active parameters, dense or MoE architecture, and open-weight status.
- Availability
- API access, lifecycle status, release date, and documented web-search support.
- Sources
- A retrieval or verification date and a direct source for each volatile result.
Need a broader directory? Browse the open-source LLM directory. For workload-specific cost estimates, use the API cost calculator.
Quick start
The 20 most popular model comparisons
Here you'll find the 20 matchups that matter right now. Compare general data, prices, benchmarks, and more.
- Muse Spark 1.3 vs. Qwen 3.8 Max 0902
- GPT-5.6 Sol vs. Muse Spark 1.3
- Gemini 3.8 Flash vs. Muse Spark 1.3
- Muse Spark 1.3 vs. Claude Opus 5
- Qwen 3.8 Max 0902 vs. DeepSeek-V4-Pro
- Claude Fable 5.1 vs. Claude Mythos 5.1
- Claude Fable 5.1 vs. GPT-5.6 Sol
- Claude Opus 5 vs. Claude Fable 5
- Claude Opus 5 vs. GPT-6 Astra
- GPT-5.6 Sol vs. Gemini 3.8 Flash
- DeepSeek-V4-Pro vs. Kimi K3
- Grok 4.6 vs. Gemini 3.8 Flash
- GLM-5.2 vs. MiMo-V2.5-Pro
- Claude Opus 5 vs. Gemini 3.8 Flash
- Claude Opus 5 vs. Grok 4.6
- GPT-5.6 Sol vs. DeepSeek-V4-Pro
- Claude Sonnet 5 vs. GPT-5.6 Terra
- Claude Opus 5 vs. Kimi K3
- GPT-5.6 Sol vs. Grok 4.6
- GPT-5.6 Sol vs. Kimi K3
Model selection
The 36 most relevant models
The database focuses on current frontier models and relevant lower-cost options. Superseded generations are excluded.
For the full technical profile of every model, browse the AI model directory.
Meta
DeepSeek
Alibaba
Mistral AI
Xiaomi
Historical reference
A comparison from ChatGPT's early model generations
The current comparison matrix intentionally focuses on supported models. This historical reference keeps the documented differences between GPT-3.5 and GPT-4 available for existing applications and migration context.
GPT-3.5 vs. GPT-4: What's the Difference?Frequently asked questions about the LLM comparison
How the benchmark matching, winners, and data freshness work.
Changelog
GPT-6 Astra added
- Added GPT-6 Astra with official specifications, pricing, context window, availability, and source links
- Expanded the comparison to 36 models and 630 head-to-head matchups
Muse Spark 1.3 and Qwen 3.8 Max 0902 added
- Added Muse Spark 1.3 with current pricing, context, availability, and source links
- Updated Qwen 3.8 Max to the 0902 snapshot and added it to the comparison roster
- Expanded the comparison to 35 models and 595 head-to-head matchups
Gemini 3.8 Flash added
- Added Gemini 3.8 Flash with official specifications, pricing, knowledge cutoff, and source links
- Expanded the comparison to 33 models and 528 head-to-head matchups
Claude 5.1 and GLM-5.3 added
- Added GLM-5.3 with official specifications, pricing, availability, and source links
- Expanded the comparison to 32 models and 496 head-to-head matchups, including Claude Fable 5.1 and Mythos 5.1
GLM-5.3-Flash and cleaner leaderboards
- Added GLM-5.3-Flash with official specifications, pricing, and source links
- Expanded the comparison to 29 models and 406 head-to-head matchups
- Removed models without a published score from leaderboard rows
Complete and cross-benchmark leaderboards
- Added all 28 comparison models to every individual benchmark leaderboard
- Added an orientation index from four setup-comparable benchmark cohorts
- Excluded ambiguous model variants and kept missing results visible without estimation
Related matchups, provider icons, and expanded FAQ
- Expanded each LLM matchup page to 14 related matchups and placed the hub button below the list
- Added local provider icons to all 28 model badges without placeholders
- Expanded the FAQ with benchmark, pricing, parameter, and radar methodology details
Benchmark leaderboards and capability profiles
- Added five separate leaderboards for overall preference, reasoning, coding, agents, and finance
- Added evidence-based capability radar charts to the head-to-head comparison pages
- Kept every ranking within one published cohort, task, metric, and source snapshot
FAQ, changelog, and internal links
- Added an FAQ about benchmark coverage, methodology, and database freshness
- Added this changelog so material updates remain visible
- Added the hub and all matchup pages to navigation, the sitemap, and related comparisons
Initial release
- Published 28 current models and 378 head-to-head matchups
- Combined shared benchmarks, API pricing, specifications, and primary sources
- Added the comparison picker and ten prioritized matchups