Language modelActive
Mercury 2
Inception Labs
- Released
- February 24, 2026
- Data date
- October 3, 2026
Mercury 2 generates answers by refining text in parallel rather than through autoregressive token-by-token decoding. Inception Labs applies a diffusion approach, more commonly associated with image generation, to a reasoning language model.
The provider specifies a 128K-token context window. It is available through the Inception API and Mercury Chat, with Mercury 2.5 later expanding context to 260K tokens.
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
AA-Omniscience Index (Show measurement, test conditions, and source)
- Source value
- -50.71666666666667
- Score
- -50.71666666666667
- Metric
- AA-Omniscience Index
- Unit
- points
- Benchmark version
- AA-Omniscience (current evaluator version not exposed on model page)
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; Omniscience index is a -100 to 100 index over the evaluator's six domains. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AA-Omniscience Accuracy (Show measurement, test conditions, and source)
- Source value
- 0.21183333333333335
- Score
- 0.21183333333333335
- Metric
- AA-Omniscience Accuracy
- Unit
- ratio
- Benchmark version
- AA-Omniscience (current evaluator version not exposed on model page)
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; accuracy is the correct-answer proportion for the Omniscience evaluator. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AA-Omniscience Hallucination Rate (Show measurement, test conditions, and source)
- Source value
- 0.9122436032987947
- Score
- 0.9122436032987947
- Metric
- AA-Omniscience Hallucination Rate
- Unit
- ratio
- Benchmark version
- AA-Omniscience (current evaluator version not exposed on model page)
- Category
- source-specific
- Direction
- Lower is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; hallucination rate is the evaluator's hallucinated-answer proportion. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
GDPval-AA (Show measurement, test conditions, and source)
- Source value
- 497.66
- Score
- 497.66
- Metric
- GDPval-AA
- Unit
- elo
- Benchmark version
- GDPval-AA v2.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; GDPval-AA v2.1 reports Elo for the 220-task evaluator, anchored to DeepSeek V4.1 Flash (max) at 1600. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
τ²-Bench Telecom (Show measurement, test conditions, and source)
- Source value
- 0.707602339181287
- Score
- 0.707602339181287
- Metric
- τ²-Bench Telecom
- Unit
- ratio
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; the model page does not expose a version label for this τ²-Bench Telecom result. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
τ²-Bench Banking (Show measurement, test conditions, and source)
- Source value
- 0.0948453608247423
- Score
- 0.0948453608247423
- Metric
- τ²-Bench Banking
- Unit
- ratio
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; the model page does not expose a version label for this τ²-Bench Banking result. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Terminal-Bench Hard (Show measurement, test conditions, and source)
- Source value
- 0.265151515151515
- Score
- 0.265151515151515
- Metric
- Terminal-Bench Hard
- Unit
- ratio
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; the model page does not expose a version label for the Terminal-Bench Hard result. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Terminal-Bench 2.1 (Show measurement, test conditions, and source)
- Source value
- 0.273408239700375
- Score
- 0.273408239700375
- Metric
- Terminal-Bench 2.1
- Unit
- ratio
- Benchmark version
- Terminal-Bench 2.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; pass@1 proportion for the exact Mercury 2 High configuration. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Terminal-Bench 4.0 (Show measurement, test conditions, and source)
- Source value
- 0
- Score
- 0
- Metric
- Terminal-Bench 4.0
- Unit
- ratio
- Benchmark version
- Terminal-Bench 4.0
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; pass@1 proportion for the exact Mercury 2 High configuration. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
SciCode (Show measurement, test conditions, and source)
- Source value
- 0.377314814814815
- Score
- 0.377314814814815
- Metric
- SciCode
- Unit
- ratio
- Benchmark version
- SciCode current AA evaluator (version not exposed on model page)
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result for the current SciCode evaluator; the evaluator page describes 288 test-set subproblems. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
AA-LCR (Show measurement, test conditions, and source)
- Source value
- 0.436666666666667
- Score
- 0.436666666666667
- Metric
- AA-LCR
- Unit
- ratio
- Benchmark version
- AA-LCR v1.1
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; AA-LCR v1.1 reports average pass rate across its long-context reasoning questions. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
IFBench (Show measurement, test conditions, and source)
- Source value
- 0.697959183673469
- Score
- 0.697959183673469
- Metric
- IFBench
- Unit
- ratio
- Category
- source-specific
- Direction
- Higher is better
- Source type
- independent-evaluator
- Evaluator
- Artificial Analysis
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Independent evaluation. Artificial Analysis exact model-page result; the model page does not expose an IFBench version label. Serving configuration: reasoning_effort=high.
- Context
- Source-specific observation; it is not a shared comparison cohort.
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 4 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Inception Labs Mercury 2 announcement (retrieved October 3, 2026; October 4, 2026) · Inception Labs Mercury 2 announcement · Editorial description reviewed October 4, 2026 |
| Additional source | Inception original publication dates (retrieved October 3, 2026) |
| Additional source | Artificial Analysis (retrieved October 4, 2026) |
| Additional source | Model metadata (retrieved October 4, 2026) |
| Additional source | Model metadata (retrieved October 4, 2026) |