Skip to main content

From 1950 to 2026

The History of AI as a Timeline

How did a research question become an everyday tool? Explore the milestones of AI, compare how models have developed, and examine the evidence behind forecasts.

Editorial review

The milestones in a timeline

From the foundations to today’s models. Search for a topic or filter by capability, developer, and year. Forecasts have their own view.

Showing 12 of 139 entries
Modality
License & type
Anthropic

Claude Sonnet 5.5: the faster complement to Opus 5.5

Anthropic releases Claude Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5, with a 1 million-token context window, a June 2026 knowledge cutoff, and unchanged pricing of $2 input and $10 output per 1 million tokens.

Agents & ToolsProprietaryModel
Source
Anthropic

Claude Opus 5.5: new leading model, lower price

Anthropic replaces Claude Opus 5 with Opus 5.5 as its new leading model, with a 1 million-token context window, a June 2026 knowledge cutoff, and pricing of $4 input and $20 output per 1 million tokens, about 40% cheaper than its predecessor.

Agents & ToolsProprietaryModel
Source
OpenAI

GPT-6 Sol and Luna: cheaper tiers below Astra

OpenAI releases GPT-6 Sol and GPT-6 Luna as successors to the GPT-5.6 Sol and Luna tiers, each with a 1,050,000-token context window. Sol costs $2/$10 per 1 million tokens, Luna $0.10/$0.50. GPT-6 Astra remains OpenAI's best model overall.

Agents & ToolsProprietaryModel
Source
OpenAI

GPT-6 Astra: OpenAI's new flagship model

OpenAI introduces GPT-6 Astra as the flagship of its GPT-6 family. According to the provider, it offers a 1.05 million-token context window and has been rolling out in phases since September.

Agents & ToolsProprietaryModel
Source
Google

Gemini 3.8 Flash: stronger coding and agents

Google releases Gemini 3.8 Flash as a stable multimodal model. It offers a 1,048,576-token context window, up to 65,536 output tokens, and improves coding and agentic tasks according to Google.

MultimodalProprietaryModel
Source
Google

Gemini 3.8 Flash Cyber: defensive cybersecurity

Google introduces Gemini 3.8 Flash Cyber for vulnerability discovery and automated patching. Access remains limited to vetted defenders in the Fairwind program.

Agents & ToolsProprietaryModel
Source
Meta

Muse Spark 1.3: longer agentic tasks with fewer steps

Meta releases Muse Spark 1.3 for agentic and coding tasks. Meta reports about 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2.

MultimodalProprietaryModel
Source
Alibaba

Qwen 3.8 Max 0902: stronger coding and agent teams

Alibaba releases a new Qwen 3.8 Max snapshot. The model retains a 1 million-token context window while improving coding, long-horizon autonomous development, agent collaboration, and visual analysis.

MultimodalProprietaryModel
Source
Anthropic

Claude Fable 5.1 and Mythos 5.1: stronger agents

Anthropic releases Fable 5.1 as its generally available frontier model and Mythos 5.1 as a restricted Project Glasswing variant. Both offer a 1 million token context window, 128,000 output tokens, and a June 2026 knowledge cutoff.

Agents & ToolsProprietaryModel
Source
Alibaba

Qwen 3.8 Flash: multimodal API for coding and agents

Alibaba releases Qwen 3.8 Flash as a multimodal API model with a native 1-million-token context for deep thinking, visual understanding, coding, and agentic workflows.

MultimodalProprietaryModel
Source
Alibaba

Qwen 3.8 Flash Next: open multimodal weights

Alibaba releases Qwen 3.8 Flash Next as a multimodal MoE under the Qwen Community License 1.0. The model activates 6 billion parameters per token and natively supports a 262,144-token context.

MultimodalOpen WeightsModel
Source
Z.ai

GLM-5.3: Z.ai introduces a new agent

Z.ai releases GLM-5.3 as a new API model for agentic tasks. The model extends the GLM 5 line for professional applications.

Agents & ToolsProprietaryModel
Source

Measured progress, not one intelligence scale

GPQA Diamond tests graduate-level science knowledge. SWE-bench Verified tests scoped software tasks. These lines show documented frontier results from different configurations and evaluations, not a normalized ranking or a universal AI score.

GPQA Diamond over time (Loading chart)

SWE-bench Verified over time (Loading chart)

Bigger, longer and more efficient

Parameter counts and context windows document technical records. Both help compare models, but neither alone predicts how reliably a model will perform on your task.

Parameter growth (log) (Loading chart)

Context-window records (log) (Loading chart)

Price and performance of current models (Loading chart)

The pace accelerates

Each cell counts sourced, non-forecast entries in this curated timeline per month. It is not a complete market census.

Entries per month (Loading chart)

How to read the data

A model can lead one measurement and trail another. That is why every graphic names its protocol, source and period.

Benchmarks

GPQA Diamond and SWE-bench Verified test different abilities. Results from different configurations and setups cannot be freely combined.

Parameters

Many current providers do not publish parameter counts. The curve therefore includes only models with a confirmed or estimated total parameter count. Those values are not directly comparable throughout.

Context windows

A large context window describes the maximum input length. It guarantees neither correct answers nor consistent quality across the whole context.

Price and performance

The scatter plot compares current API list prices with one benchmark. Discounts, caching and other tariffs are outside this view.

Who leads chatbot web traffic

Chatbot market share by web traffic (Loading chart)

AI on the world map

The maps show separate datasets. The model map counts only models tracked in the Gradually database, grouped by developer location.

Models in the Gradually directory by developer location
Loading globe…

Use the arrow keys to rotate the globe. Press Home to restore the initial view.

2 models135 models
Data as a table
CountryValue
Canada4 models
China72 models
France18 models
United Arab Emirates2 models
United States135 models
Source: Gradually directory (Sep 25, 2026)
gradually.ai

What the data suggests, and what remains open

The measured task series ends in 2025. Its continuation makes its assumptions visible. Expert forecasts also answer different questions, so they sit side by side instead of becoming one falsely precise number.

AI task horizon, from seconds to hours
measured (METR)projection (doubling every 2.9 to 6.5 months)
1 minute1 hour1 work day1 work week1 work monthSep 25, 2026GPT-4 0314Claude Opus 4.52023202620292032

Selected published METR values through late 2025. The horizon describes tasks a model completes with 50% success, measured in human working time. The shaded area extrapolates two observed doubling rates. It is neither a confidence interval nor a promise about real projects. Logarithmic time axis.

Source: METR, Time Horizon 1.1 (2026)
gradually.ai
When will general AI arrive? Estimates diverge widely
ScenarioCommunity/marketsResearcher survey
AI 2027
Possible scenario, not a prediction
Source dated
2027
Metaculus
First general AI, community forecast on retrieval date
Source dated
2031
Grace et al. 2024
HLMI, 50% (10% by 2027); n = 2,778 researchers
Source dated
2047
2030203520402045

These points use different definitions and were published in different years. They do not form a consensus or a shared probability interval. A scenario describes one possible course of events, not a guaranteed outcome.

AI 2027, Grace et al. 2024, Metaculus
gradually.ai

Frequently asked questions about the history of AI

The key questions about the development of AI models and tools.

Changelog

The latest updates and improvements to our History of AI
v1.6September 30, 2026

New model released September 28

  • Added Claude Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5, succeeding Sonnet 5
v1.5September 25, 2026

Timeline and evidence

  • Added historical foundations and a separate forecast view
  • Activated charts, the globe, and filters on the Astro page
  • Clarified sources, dates, and comparison limits
v1.4September 22, 2026

New models released September 22

  • Added Claude Opus 5.5 as Anthropic's new leading model, succeeding Opus 5
  • Added GPT-6 Sol and GPT-6 Luna as cheaper successors to the GPT-5.6 tiers
v1.3September 3, 2026

New models released September 2

  • Added Muse Spark 1.3 as a new multimodal agentic and coding model
  • Added Qwen 3.8 Max 0902 as a new snapshot with improved coding and agent collaboration
  • Added Gemini 3.8 Flash Cyber as a restricted cybersecurity variant for vetted defenders
  • Added Qwen 3.8 Flash and Qwen 3.8 Flash Next as an API model and open-weight model
v1.2August 23, 2026

July and August leading-model updates

  • Added Claude Sonnet 5, Opus 5, and Kimi K3 to the July and August frontier releases
  • Added Grok 4.5 and Grok 4.6 as new agentic models
  • Added new Gemini Flash variants from July and August to the timeline
  • Added Qwen 3.8 as an open-weights milestone
v1.1June 28, 2026

Prehistory & forecasts

  • Added the deep-learning prehistory: milestones from AlexNet (2012) to AlphaGo and WaveNet (2016)
  • New forecast section with trend extrapolation (METR task horizon, Epoch compute) and sources
  • Forecasts for 2027-2040 as marked entries in the timeline, with uncertainty ranges
v1.0June 21, 2026

Initial release

  • Interactive timeline of generative-AI milestones from 2017 to 2026
  • Filter by modality, license, developer, and year, plus timeline and grid views
  • Stats dashboard with charts on the intelligence explosion (GPQA, SWE-bench, parameters, context)
  • Benchmark scores for older models individually sourced