Mercury 2.5
Inception
- Released
- September 8, 2026
- Data date
- October 3, 2026
Mercury 2.5 expands the context window to 260K tokens, compared with 128K in Mercury 2. Inception retains diffusion-based generation and reports 1,107 tokens per second on NVIDIA GPUs for the new version. That speed comes from the provider’s measurement.
Inception revised training using customer feedback and production failure cases. Mercury 2.5 generates text and is available through the Inception API, Baseten, and OpenRouter.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Diffusion language modelSource |
| Context window | 260K tokens according to InceptionSource |
| Speed | 1,107 tokens per second according to the providerSource |
| Capabilities | Reasoning, parallel tool calls, and schema-aligned JSONSource |
| Access | Inception API, Baseten, and OpenRouterSource |
Page 1 of 2
Published benchmarks
Coding
Code Migration
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-code-migration:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 57
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 997.717. Cost per task in USD: 0.627223. Source ID: vals-ai. Evidence capture ID: vals-code-migration-2026-09-29-f0b317d5bed3. Source payload hash: f0b317d5bed3a2654a9ea39762d5e57f4649306adce3e6103e845ee1b9b94add. Parser version: vals-astro-v2. Imported at: 2026-09-29
IOI
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-ioi:2:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 28
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 242.94. Cost per task in USD: 0.160874. Source ID: vals-ai. Evidence capture ID: vals-ioi-2026-09-29-14c4f2dcc872. Source payload hash: 14c4f2dcc872d5e07bce234aebe4932202931d7f5b05f9bcd02e7695b5664750. Parser version: vals-astro-v2. Imported at: 2026-09-29
CyberBench v1.1
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-cyberbench:1.1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 29
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 1139.168. Cost per task in USD: 0.12462. Source ID: vals-ai. Evidence capture ID: vals-cyber-2026-09-29-36ddb5f67bc8. Source payload hash: 36ddb5f67bc8725a7ff9d1cc1349f89d29e0e907ad313eb5ac277b42b3d275e1. Parser version: vals-astro-v2. Imported at: 2026-09-29
Finance
EMB
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-emb:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 54
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 301.681. Cost per task in USD: 0.419811. Source ID: vals-ai. Evidence capture ID: vals-emb-2026-09-29-90535b2fd4dc. Source payload hash: 90535b2fd4dc213540eac0aff8ae769e34b6b1ba54e1ed6d0a51dcfe307d5598. Parser version: vals-astro-v2. Imported at: 2026-09-29
Finance Agent (v2)
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-finance-agent-v2:2:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 55
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 95.056. Cost per task in USD: 0.106441. Source ID: vals-ai. Evidence capture ID: vals-fabv2-2026-09-29-addcd124c670. Source payload hash: addcd124c6706893659c1ae26698bf6ffa29c91018ee95b98ab15b9d4b6edcdb. Parser version: vals-astro-v2. Imported at: 2026-09-29
Legal
Harvey's Legal Agent Benchmark
Overall · Task fully resolved
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-legal-agent-benchmark:1:overall
- Metric
- task resolution rate
- Scale
- 0-100
- Participants
- 56
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 125.317. Cost per task in USD: 0.192614. Source ID: vals-ai. Evidence capture ID: vals-hlab-2026-09-29-d6fb708e5bcc. Source payload hash: d6fb708e5bccec56c7a5e5ad97918d9cc47c5ebefb88af4a1a5771493dadba2f. Parser version: vals-astro-v2. Imported at: 2026-09-29
LegalBench
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-legal-bench:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 78
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 3.908. Cost per task in USD: 0.000787. Source ID: vals-ai. Evidence capture ID: vals-legal_bench-2026-09-29-c1213019b552. Source payload hash: c1213019b55234a262ae86fd4ebfe291530f315c988a4ee32fd501c13777f85f. Parser version: vals-astro-v2. Imported at: 2026-09-29
Legal Research Bench
Overall · All-pass
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-legal-research:1:overall
- Metric
- all-pass rate
- Scale
- 0-100
- Participants
- 55
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 45.996. Cost per task in USD: 0.058112. Source ID: vals-ai. Evidence capture ID: vals-legal_research-2026-09-29-b15acd492114. Source payload hash: b15acd492114174e151eff7ff17a38d547d5a1e93d2d598c153066542bbc56b7. Parser version: vals-astro-v2. Imported at: 2026-09-29
Health
MedCode
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-medcode:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 64
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-26. Latency in seconds: 7.93. Cost per task in USD: 0.004535. Source ID: vals-ai. Evidence capture ID: vals-medcode-2026-09-29-7062229b5ab4. Source payload hash: 7062229b5ab4a5e16f014a78fd550f9f34e7147be6be796b610ab1ca3de055ab. Parser version: vals-astro-v2. Imported at: 2026-09-29
MedScribe
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-medscribe:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 65
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-26. Latency in seconds: 9.992. Cost per task in USD: 0.004476. Source ID: vals-ai. Evidence capture ID: vals-medscribe-2026-09-29-6533d95c8055. Source payload hash: 6533d95c80553d2dd26927f7269ab6454c589d2181f1858c662a4f493b5685c7. Parser version: vals-astro-v2. Imported at: 2026-09-29
ProofBench
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-proof-bench:1.1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 34
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-28. Latency in seconds: 201.725. Cost per task in USD: 0.044105. Source ID: vals-ai. Evidence capture ID: vals-proof_bench-2026-09-29-95d190db7be5. Source payload hash: 95d190db7be508012399269386113bf7df88cf108a8647fae6d361d15c60126b. Parser version: vals-astro-v2. Imported at: 2026-09-29
SkillsBench
Overall
higher is betterAs of September 29, 2026Scale: 0% - 100%
Methodology and context
- Comparison cohort
- vals-skillsbench:1:overall
- Metric
- accuracy
- Scale
- 0-100
- Participants
- 30
- Methodology and context
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Harness: OpenHands. Latency in seconds: 49.364. Cost per task in USD: 0.101286. Source ID: vals-ai. Evidence capture ID: vals-skillsbench-2026-09-29-38ec42553ceb. Source payload hash: 38ec42553ceb08e4efe7336723e87e67325977245b113ef7b86401dc72b0481a. Parser version: vals-astro-v2. Imported at: 2026-09-29
Recorded measurement details
Every observation retains its source value and published test conditions.
Code Migration: Overall (Show measurement, test conditions, and source)
- Source value
- 4.494
- Score
- 4.494
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- coding
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-code-migration:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 997.717. Cost per task in USD: 0.627223. Source ID: vals-ai. Evidence capture ID: vals-code-migration-2026-09-29-f0b317d5bed3. Source payload hash: f0b317d5bed3a2654a9ea39762d5e57f4649306adce3e6103e845ee1b9b94add. Parser version: vals-astro-v2. Imported at: 2026-09-29
CyberBench v1.1: Overall (Show measurement, test conditions, and source)
- Source value
- 45.476
- Score
- 45.476
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1.1
- Category
- cybersecurity
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-cyberbench:1.1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 1139.168. Cost per task in USD: 0.12462. Source ID: vals-ai. Evidence capture ID: vals-cyber-2026-09-29-36ddb5f67bc8. Source payload hash: 36ddb5f67bc8725a7ff9d1cc1349f89d29e0e907ad313eb5ac277b42b3d275e1. Parser version: vals-astro-v2. Imported at: 2026-09-29
EMB: Overall (Show measurement, test conditions, and source)
- Source value
- 8.823
- Score
- 8.823
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- finance
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-emb:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 301.681. Cost per task in USD: 0.419811. Source ID: vals-ai. Evidence capture ID: vals-emb-2026-09-29-90535b2fd4dc. Source payload hash: 90535b2fd4dc213540eac0aff8ae769e34b6b1ba54e1ed6d0a51dcfe307d5598. Parser version: vals-astro-v2. Imported at: 2026-09-29
Finance Agent (v2): Overall (Show measurement, test conditions, and source)
- Source value
- 18.556
- Score
- 18.556
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 2
- Category
- finance
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-finance-agent-v2:2:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 95.056. Cost per task in USD: 0.106441. Source ID: vals-ai. Evidence capture ID: vals-fabv2-2026-09-29-addcd124c670. Source payload hash: addcd124c6706893659c1ae26698bf6ffa29c91018ee95b98ab15b9d4b6edcdb. Parser version: vals-astro-v2. Imported at: 2026-09-29
IOI: Overall (Show measurement, test conditions, and source)
- Source value
- 2.444
- Score
- 2.444
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 2
- Category
- coding
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-ioi:2:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 242.94. Cost per task in USD: 0.160874. Source ID: vals-ai. Evidence capture ID: vals-ioi-2026-09-29-14c4f2dcc872. Source payload hash: 14c4f2dcc872d5e07bce234aebe4932202931d7f5b05f9bcd02e7695b5664750. Parser version: vals-astro-v2. Imported at: 2026-09-29
Harvey's Legal Agent Benchmark: Overall · Task fully resolved (Show measurement, test conditions, and source)
- Source value
- 0
- Score
- 0
- Metric
- task resolution rate
- Unit
- %
- Benchmark version
- 1
- Category
- legal
- Task
- overall
- Task label
- Overall · Task fully resolved
- Direction
- Higher is better
- Comparison cohort
- vals-legal-agent-benchmark:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 125.317. Cost per task in USD: 0.192614. Source ID: vals-ai. Evidence capture ID: vals-hlab-2026-09-29-d6fb708e5bcc. Source payload hash: d6fb708e5bccec56c7a5e5ad97918d9cc47c5ebefb88af4a1a5771493dadba2f. Parser version: vals-astro-v2. Imported at: 2026-09-29
LegalBench: Overall (Show measurement, test conditions, and source)
- Source value
- 83.06
- Score
- 83.06
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- legal
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-legal-bench:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Latency in seconds: 3.908. Cost per task in USD: 0.000787. Source ID: vals-ai. Evidence capture ID: vals-legal_bench-2026-09-29-c1213019b552. Source payload hash: c1213019b55234a262ae86fd4ebfe291530f315c988a4ee32fd501c13777f85f. Parser version: vals-astro-v2. Imported at: 2026-09-29
Legal Research Bench: Overall · All-pass (Show measurement, test conditions, and source)
- Source value
- 4.327
- Score
- 4.327
- Metric
- all-pass rate
- Unit
- %
- Benchmark version
- 1
- Category
- legal
- Task
- overall
- Task label
- Overall · All-pass
- Direction
- Higher is better
- Comparison cohort
- vals-legal-research:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-27. Latency in seconds: 45.996. Cost per task in USD: 0.058112. Source ID: vals-ai. Evidence capture ID: vals-legal_research-2026-09-29-b15acd492114. Source payload hash: b15acd492114174e151eff7ff17a38d547d5a1e93d2d598c153066542bbc56b7. Parser version: vals-astro-v2. Imported at: 2026-09-29
MedCode: Overall (Show measurement, test conditions, and source)
- Source value
- 31.326
- Score
- 31.326
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- health
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-medcode:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-26. Latency in seconds: 7.93. Cost per task in USD: 0.004535. Source ID: vals-ai. Evidence capture ID: vals-medcode-2026-09-29-7062229b5ab4. Source payload hash: 7062229b5ab4a5e16f014a78fd550f9f34e7147be6be796b610ab1ca3de055ab. Parser version: vals-astro-v2. Imported at: 2026-09-29
MedScribe: Overall (Show measurement, test conditions, and source)
- Source value
- 55.093
- Score
- 55.093
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- health
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-medscribe:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-26. Latency in seconds: 9.992. Cost per task in USD: 0.004476. Source ID: vals-ai. Evidence capture ID: vals-medscribe-2026-09-29-6533d95c8055. Source payload hash: 6533d95c80553d2dd26927f7269ab6454c589d2181f1858c662a4f493b5685c7. Parser version: vals-astro-v2. Imported at: 2026-09-29
ProofBench: Overall (Show measurement, test conditions, and source)
- Source value
- 3
- Score
- 3
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1.1
- Category
- math
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-proof-bench:1.1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: private. Leaderboard date: 2026-09-28. Latency in seconds: 201.725. Cost per task in USD: 0.044105. Source ID: vals-ai. Evidence capture ID: vals-proof_bench-2026-09-29-95d190db7be5. Source payload hash: 95d190db7be508012399269386113bf7df88cf108a8647fae6d361d15c60126b. Parser version: vals-astro-v2. Imported at: 2026-09-29
SkillsBench: Overall (Show measurement, test conditions, and source)
- Source value
- 18.061
- Score
- 18.061
- Metric
- accuracy
- Unit
- %
- Benchmark version
- 1
- Category
- agentic
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-skillsbench:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: inception/mercury-2.5. Source model name: Mercury 2.5. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-27. Harness: OpenHands. Latency in seconds: 49.364. Cost per task in USD: 0.101286. Source ID: vals-ai. Evidence capture ID: vals-skillsbench-2026-09-29-38ec42553ceb. Source payload hash: 38ec42553ceb08e4efe7336723e87e67325977245b113ef7b86401dc72b0481a. Parser version: vals-astro-v2. Imported at: 2026-09-29
- Comparison methodology
- Harness: OpenHands
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.