Compare LLMs without mixing incompatible test setups
Compare 50 leading language models across 1,225 head-to-head matchups. Every result includes shared benchmark cohorts, current API pricing, key specifications, and source links.
- models
- 50
- head-to-head matchups
- 1,225
- active benchmarks
- 89
- active results
- 14,680
Cross-benchmark view
Overall LLM leaderboard
The index uses 10 benchmark cohorts with comparable documented test setups. It converts the scores into rank percentiles and averages them with equal weight over the cohorts a model appears in. The 86 models measured in at least 4 of them receive an overall rank, and each row states how many carry it. Ranks based on fewer than half of the cohorts are marked as provisional, because they usually move once more results are published. Models below that threshold are omitted, and missing data is not treated as a penalty.
Show all 86 ranked modelsShow the top 20 only
This is a transparent orientation index, not an additional benchmark. It does not include price, speed, context length, or results outside the listed 10 benchmark cohorts.
Source entries with incompatible documented configurations are excluded from the composite instead of being reduced to their strongest result.
Current leaders
Which LLMs lead the most important benchmarks?
These five charts keep the benchmarks separate. Each ranking uses one published leaderboard cohort, one task, and one metric. Models without an unambiguous published result in that cohort are omitted instead of receiving an estimated score.
Overall preference
LMArena Text, Style Control
Style-controlled rating from anonymous pairwise comparisons.
The published leaderboard documents model-specific test settings. 6 comparison models have more than one incompatible setup. They are omitted from this ranking instead of selecting the best result.
Reasoning
GPQA Diamond
Difficult scientific questions from the GPQA Diamond dataset.
Coding
Vibe Code Bench v1.1
Practical programming tasks from Vibe Code Bench v1.1.
The published leaderboard documents model-specific test settings. 1 comparison model has more than one incompatible setup. It is omitted from this ranking instead of selecting the best result.
Agents
Terminal-Bench 2.1
Multi-step tasks completed in a real terminal environment.
Finance
Finance Agent (v2)
Research and analysis tasks from Finance Agent v2.
A first-place result applies only to the named benchmark. It is not an overall quality score and says nothing by itself about price, speed, context length, or your own workload.
Decision criteria
What this LLM comparison lets you check
- Benchmarks
- Shared tasks and versions for reasoning, coding, agents, research, and other capabilities.
- API pricing
- Input, output, cache-read, and cache-write pricing where the provider publishes it.
- Context
- Published context windows and maximum output limits without imprecise rounding.
- Architecture
- Known parameters, active parameters, dense or MoE architecture, and open-weight status.
- Availability
- API access, lifecycle status, release date, and documented web-search support.
- Sources
- A retrieval or verification date and a direct source for each volatile result.
Need a broader directory? Browse the open-source LLM directory. For workload-specific cost estimates, use the API cost calculator.
Quick start
The 20 most popular model comparisons
Here you'll find the 20 matchups that matter right now. Compare general data, prices, benchmarks, and more.
- GPT-6.1 Sol vs. GPT-6 Sol
- GPT-6.1 Sol vs. Claude Sonnet 5.5
- Muse Spark 1.3 vs. Qwen 3.8 Max 0902
- GPT-5.6 Sol vs. Muse Spark 1.3
- Gemini 3.8 Flash vs. Muse Spark 1.3
- Muse Spark 1.3 vs. Claude Opus 5
- Qwen 3.8 Max 0902 vs. DeepSeek-V4-Pro
- Claude Fable 5.1 vs. Claude Mythos 5.1
- Claude Fable 5.1 vs. GPT-5.6 Sol
- Claude Opus 5.5 vs. Claude Fable 5.1
- Claude Opus 5 vs. GPT-6 Astra
- GPT-5.6 Sol vs. Gemini 3.8 Flash
- DeepSeek-V4-Pro vs. Kimi K3
- Grok 4.6 vs. Gemini 3.8 Flash
- GLM-5.2 vs. MiMo-V2.5-Pro
- Claude Opus 5 vs. Gemini 3.8 Flash
- Claude Opus 5 vs. Grok 4.6
- GPT-5.6 Sol vs. DeepSeek-V4-Pro
- Claude Sonnet 5.5 vs. GPT-5.6 Terra
- Claude Opus 5 vs. Kimi K3
Model selection
The 50 most relevant models
The selection includes current frontier models, lower-cost options, and selected earlier generations. You can compare Claude Opus 4.5 to 4.8, Sonnet 4.5 and 4.6, and Haiku 4.5 with newer models.
For the full technical profile of every model, browse the AI model directory.
Anthropic
OpenAI
Meta
Alibaba
Mistral AI
Xiaomi
Historical reference
A comparison from ChatGPT's early model generations
The current comparison matrix intentionally focuses on supported models. This historical reference keeps the documented differences between GPT-3.5 and GPT-4 available for existing applications and migration context.
GPT-3.5 vs. GPT-4: What's the Difference?Frequently asked questions about the LLM comparison
How the benchmark matching, winners, and data freshness work.
Changelog
Claude Sonnet 5.5 added
- Added Claude Sonnet 5.5 with official specifications, pricing, context window, knowledge cutoff, and source links
- Expanded the comparison to 50 models and 1,225 head-to-head matchups, including GPT-6.1 Sol, Grok 4.7, and DeepSeek-V4.1-Flash
Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna added
- Added Claude Opus 5.5, GPT-6 Sol, and GPT-6 Luna with official specifications, pricing, context windows, knowledge cutoffs, and source links
- Expanded the comparison to 46 models and 1,035 head-to-head matchups
46-model roster and provider-balanced hosting prices
- Expanded the comparison to 46 models and 1,035 head-to-head matchups
- Switched hosting-market tariffs to the daily provider-balanced OpenRouter observation, so profiles and comparison pages show the same figure as the price index
- Refreshed deprecation and lifecycle status against current provider pages
GPT-6 Astra added
- Added GPT-6 Astra with official specifications, pricing, context window, availability, and source links
- Expanded the comparison to 36 models and 630 head-to-head matchups
Muse Spark 1.3 and Qwen 3.8 Max 0902 added
- Added Muse Spark 1.3 with current pricing, context, availability, and source links
- Updated Qwen 3.8 Max to the 0902 snapshot and added it to the comparison roster
- Expanded the comparison to 35 models and 595 head-to-head matchups
Gemini 3.8 Flash added
- Added Gemini 3.8 Flash with official specifications, pricing, knowledge cutoff, and source links
- Expanded the comparison to 33 models and 528 head-to-head matchups
Claude 5.1 and GLM-5.3 added
- Added GLM-5.3 with official specifications, pricing, availability, and source links
- Expanded the comparison to 32 models and 496 head-to-head matchups, including Claude Fable 5.1 and Mythos 5.1
GLM-5.3-Flash and cleaner leaderboards
- Added GLM-5.3-Flash with official specifications, pricing, and source links
- Expanded the comparison to 29 models and 406 head-to-head matchups
- Removed models without a published score from leaderboard rows
Complete and cross-benchmark leaderboards
- Added all 28 comparison models to every individual benchmark leaderboard
- Added an orientation index from four setup-comparable benchmark cohorts
- Excluded ambiguous model variants and kept missing results visible without estimation
Related matchups, provider icons, and expanded FAQ
- Expanded each LLM matchup page to 14 related matchups and placed the hub button below the list
- Added local provider icons to all 28 model badges without placeholders
- Expanded the FAQ with benchmark, pricing, parameter, and radar methodology details
Benchmark leaderboards and capability profiles
- Added five separate leaderboards for overall preference, reasoning, coding, agents, and finance
- Added evidence-based capability radar charts to the head-to-head comparison pages
- Kept every ranking within one published cohort, task, metric, and source snapshot
FAQ, changelog, and internal links
- Added an FAQ about benchmark coverage, methodology, and database freshness
- Added this changelog so material updates remain visible
- Added the hub and all matchup pages to navigation, the sitemap, and related comparisons
Initial release
- Published 28 current models and 378 head-to-head matchups
- Combined shared benchmarks, API pricing, specifications, and primary sources
- Added the comparison picker and ten prioritized matchups