Benchmarks
GPQA Diamond and SWE-bench Verified test different abilities. Results from different configurations and setups cannot be freely combined.
From 1950 to 2026
How did a research question become an everyday tool? Explore the milestones of AI, compare how models have developed, and examine the evidence behind forecasts.
Editorial review
From the foundations to today’s models. Search for a topic or filter by capability, developer, and year. Forecasts have their own view.
Anthropic releases Claude Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5, with a 1 million-token context window, a June 2026 knowledge cutoff, and unchanged pricing of $2 input and $10 output per 1 million tokens.
Anthropic replaces Claude Opus 5 with Opus 5.5 as its new leading model, with a 1 million-token context window, a June 2026 knowledge cutoff, and pricing of $4 input and $20 output per 1 million tokens, about 40% cheaper than its predecessor.
OpenAI releases GPT-6 Sol and GPT-6 Luna as successors to the GPT-5.6 Sol and Luna tiers, each with a 1,050,000-token context window. Sol costs $2/$10 per 1 million tokens, Luna $0.10/$0.50. GPT-6 Astra remains OpenAI's best model overall.
OpenAI introduces GPT-6 Astra as the flagship of its GPT-6 family. According to the provider, it offers a 1.05 million-token context window and has been rolling out in phases since September.
Google releases Gemini 3.8 Flash as a stable multimodal model. It offers a 1,048,576-token context window, up to 65,536 output tokens, and improves coding and agentic tasks according to Google.
Google introduces Gemini 3.8 Flash Cyber for vulnerability discovery and automated patching. Access remains limited to vetted defenders in the Fairwind program.
Meta releases Muse Spark 1.3 for agentic and coding tasks. Meta reports about 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2.
Alibaba releases a new Qwen 3.8 Max snapshot. The model retains a 1 million-token context window while improving coding, long-horizon autonomous development, agent collaboration, and visual analysis.
Anthropic releases Fable 5.1 as its generally available frontier model and Mythos 5.1 as a restricted Project Glasswing variant. Both offer a 1 million token context window, 128,000 output tokens, and a June 2026 knowledge cutoff.
Alibaba releases Qwen 3.8 Flash as a multimodal API model with a native 1-million-token context for deep thinking, visual understanding, coding, and agentic workflows.
Alibaba releases Qwen 3.8 Flash Next as a multimodal MoE under the Qwen Community License 1.0. The model activates 6 billion parameters per token and natively supports a 262,144-token context.
Z.ai releases GLM-5.3 as a new API model for agentic tasks. The model extends the GLM 5 line for professional applications.
GPQA Diamond tests graduate-level science knowledge. SWE-bench Verified tests scoped software tasks. These lines show documented frontier results from different configurations and evaluations, not a normalized ranking or a universal AI score.
GPQA Diamond over time (Loading chart)
SWE-bench Verified over time (Loading chart)
Parameter counts and context windows document technical records. Both help compare models, but neither alone predicts how reliably a model will perform on your task.
Parameter growth (log) (Loading chart)
Context-window records (log) (Loading chart)
Price and performance of current models (Loading chart)
Each cell counts sourced, non-forecast entries in this curated timeline per month. It is not a complete market census.
Entries per month (Loading chart)
A model can lead one measurement and trail another. That is why every graphic names its protocol, source and period.
GPQA Diamond and SWE-bench Verified test different abilities. Results from different configurations and setups cannot be freely combined.
Many current providers do not publish parameter counts. The curve therefore includes only models with a confirmed or estimated total parameter count. Those values are not directly comparable throughout.
A large context window describes the maximum input length. It guarantees neither correct answers nor consistent quality across the whole context.
The scatter plot compares current API list prices with one benchmark. Discounts, caching and other tariffs are outside this view.
Chatbot market share by web traffic (Loading chart)
The maps show separate datasets. The model map counts only models tracked in the Gradually database, grouped by developer location.
Use the arrow keys to rotate the globe. Press Home to restore the initial view.
| Country | Value |
|---|---|
| Canada | 4 models |
| China | 72 models |
| France | 18 models |
| United Arab Emirates | 2 models |
| United States | 135 models |
The measured task series ends in 2025. Its continuation makes its assumptions visible. Expert forecasts also answer different questions, so they sit side by side instead of becoming one falsely precise number.
Selected published METR values through late 2025. The horizon describes tasks a model completes with 50% success, measured in human working time. The shaded area extrapolates two observed doubling rates. It is neither a confidence interval nor a promise about real projects. Logarithmic time axis.
These points use different definitions and were published in different years. They do not form a consensus or a shared probability interval. A scenario describes one possible course of events, not a guaranteed outcome.
The key questions about the development of AI models and tools.