Language modelEnterprise access
Mercury Voice
Inception
- Released
- September 29, 2026
- Data date
- October 3, 2026
Mercury Voice targets the time to first answer token in voice agents. In its own tests on customer-service prompts, Inception reports a median of 320 milliseconds after reasoning. Despite its name, this diffusion LLM generates text; other components handle speech recognition and audio output.
Inception offers the checkpoint through its own API for enterprise customers. It handles 128K context tokens and up to 50K output tokens.
Specifications and access
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
average quality (Show measurement, test conditions, and source)
- Source value
- 70.5
- Score
- 70.5
- Metric
- average quality reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
tau3 Telecom (Show measurement, test conditions, and source)
- Source value
- 77.2
- Score
- 77.2
- Metric
- tau3 Telecom reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
tau3 Airline (Show measurement, test conditions, and source)
- Source value
- 78
- Score
- 78
- Metric
- tau3 Airline reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
tau3 Retail (Show measurement, test conditions, and source)
- Source value
- 69.3
- Score
- 69.3
- Metric
- tau3 Retail reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
IFBench (Show measurement, test conditions, and source)
- Source value
- 67.4
- Score
- 67.4
- Metric
- IFBench reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
BFCL v4 multi-turn (Show measurement, test conditions, and source)
- Source value
- 60.5
- Score
- 60.5
- Metric
- BFCL v4 multi-turn reported score
- Unit
- No unit provided
- Category
- source-specific
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- inceptionlabs.ai
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Values read from the official image-only table. Quality composite and latency use Inception's voice-agent prompts/configuration.
- Context
- Source-specific observation; it is not a shared comparison cohort. The source table does not state a unit or scale for this numeric score.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Inception Mercury Voice announcement (retrieved October 3, 2026; October 4, 2026) · Inception Mercury Voice announcement · Editorial description reviewed October 4, 2026 |
| Additional source | Inception model blog (retrieved October 3, 2026) · Inception model blog · Editorial description reviewed October 4, 2026 |