Skip to main content
Language modelActive

Mercury 2.5

Inception

Released
September 8, 2026
Data date
October 3, 2026

Mercury 2.5 expands the context window to 260K tokens, compared with 128K in Mercury 2. Inception retains diffusion-based generation and reports 1,107 tokens per second on NVIDIA GPUs for the new version. That speed comes from the provider’s measurement.

Inception revised training using customer feedback and production failure cases. Mercury 2.5 generates text and is available through the Inception API, Baseten, and OpenRouter.

Specifications and access

SpecificationValue and source
Model class
Diffusion language modelSource
Context window
260K tokens according to InceptionSource
Speed
1,107 tokens per second according to the providerSource
Capabilities
Reasoning, parallel tool calls, and schema-aligned JSONSource
Access
Inception API, Baseten, and OpenRouterSource