Skip to main content

Gradually LLM Index

Compare LLMs without mixing incompatible test setups

Compare 28 leading language models across 378 head-to-head matchups. Every result includes shared benchmark cohorts, current API pricing, key specifications, and source links.

models
28
head-to-head matchups
378
active benchmarks
80
active results
14,135

Short answer

Which LLM is the best?

There is no single best LLM for every workload. A coding agent, a low-cost extraction pipeline, and a long-context research task reward different strengths.

Choose the two models you are considering. The result page shows only shared measurements and keeps pricing, context, architecture, and availability separate from benchmark scores.

Decision criteria

What this LLM comparison lets you check

Benchmarks
Shared tasks and versions for reasoning, coding, agents, research, and other capabilities.
API pricing
Input, output, cache-read, and cache-write pricing where the provider publishes it.
Context
Published context windows and maximum output limits without imprecise rounding.
Architecture
Known parameters, active parameters, dense or MoE architecture, and open-weight status.
Availability
API access, lifecycle status, release date, and documented web-search support.
Sources
A retrieval or verification date and a direct source for each volatile result.

Need a broader directory? Browse the open-source LLM directory. For workload-specific cost estimates, use the API cost calculator.

Model selection

The 28 most relevant models

The database focuses on current frontier models and relevant lower-cost options. Superseded generations are excluded.

Anthropic

Claude Opus 5Claude Fable 5Claude Sonnet 5Claude Mythos 5

OpenAI

GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGPT-5.5GPT-5.4GPT-5.4 mini

Google

Gemini 3.7 FlashGemini 3.6 FlashGemini 3.5 FlashGemini 3.1 Pro PreviewGemini 3.5 Flash-Lite

xAI

Grok 4.6Grok 4.5Grok 4.3

DeepSeek

DeepSeek-V4-ProDeepSeek-V4-Flash

Moonshot AI

Kimi K3Kimi K2.6

Z.ai

GLM-5.2GLM-5.1

Mistral AI

Mistral Large 3Mistral Medium 3.5

Xiaomi

MiMo-V2.5MiMo-V2.5-Pro

Methodology

A comparison is only as sound as its cohort

Central database

28 sources, 14,135 active results, and one consistent model schema.

Matched test setup

Results are paired only when benchmark version, task, metric, cohort, and documented setup align.

Missing data stays missing

If two models have no shared measurement series, the gap stays visible. We do not estimate scores.

Evidence, not one overall ranking

Benchmarks, pricing, and specifications sit side by side. The right model depends on your workload.

Database retrieved on August 16, 2026. Each model price has its own verification date and linked primary source.

FAQ

Questions about the LLM comparison

What is an LLM comparison?

It compares large language models using matched benchmark results, pricing, specifications, availability, and source dates.

Why do some matchups show few benchmarks?

We do not combine different benchmark versions, tasks, metrics, cohorts, reasoning levels, or agent setups into one score.

Are benchmark winners always better?

No. A measured lead applies to the documented task and setup. Cost, latency, context, and tool access can change the practical choice.

How current is the database?

The benchmark snapshot was retrieved on August 16, 2026. Volatile model facts carry their own source and verification dates.