Cartesia Ink-2
Cartesia
- Released
- July 9, 2026
- Data date
- October 3, 2026
- Model class
- Access
Cartesia Ink-2 is a streaming ASR model for voice agents. Cartesia reports about 0.1 seconds from the end of speech to the final transcript (TTFT), not to the first partial result. It adds semantic endpointing and structured entities to transcription.
The ink-2 endpoint is available through the Cartesia API and Play. Its launch named English, while a newer product page also lists Spanish, French, Hindi, and Japanese. The official sources do not give a consistent language list. Billing uses three credits per second of audio, with the dollar price depending on the plan.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Streaming speech recognitionSource |
| Access | Cartesia API and PlaySource |
| Input | AudioSource |
| Output | Transcript and structured entitiesSource |
| API model ID | ink-2Source |
| Languages | English at launch according to the announcement; the current Ink product page also lists Spanish, French, Hindi, and Japanese. Language-coverage claims are not consistent.Source |
| Time to final transcript | Provider-reported 0.1 seconds from the end of speech to the final transcript (TTFT), not to the first partial resultSource |
| Features | Semantic endpointing and structured entity recognitionSource |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
VoiceCodeBench: CTEM (Show measurement, test conditions, and source)
- Source value
- 89.541
- Score
- 89.541
- Metric
- canonical token/entity match
- Unit
- %
- Benchmark version
- 1
- Category
- general
- Task
- ctem
- Task label
- CTEM
- Direction
- Higher is better
- Comparison cohort
- vals-voice-code-bench:1:ctem
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: cartesia/ink-2. Source model name: Ink 2. Provider reported by source: Cartesia. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-24. Latency in seconds: 69.197. Standard error: 0. Source ID: vals-ai. Evidence capture ID: vals-voice-code-bench-2026-09-29-7d87ad468e81. Source payload hash: 7d87ad468e8136922daec73432384e9d40b0f52e64cc6e6a0e73db7d6e7c264c. Parser version: vals-astro-v2. Imported at: 2026-09-29
VoiceCodeBench: Overall (Show measurement, test conditions, and source)
- Source value
- 62
- Score
- 62
- Metric
- task success rate
- Unit
- %
- Benchmark version
- 1
- Category
- general
- Task
- overall
- Task label
- Overall
- Direction
- Higher is better
- Comparison cohort
- vals-voice-code-bench:1:overall
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: cartesia/ink-2. Source model name: Ink 2. Provider reported by source: Cartesia. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-24. Latency in seconds: 69.197. Standard error: 0. Source ID: vals-ai. Evidence capture ID: vals-voice-code-bench-2026-09-29-7d87ad468e81. Source payload hash: 7d87ad468e8136922daec73432384e9d40b0f52e64cc6e6a0e73db7d6e7c264c. Parser version: vals-astro-v2. Imported at: 2026-09-29
Endpointing precision (Show measurement, test conditions, and source)
- Source value
- 89
- Score
- 89
- Metric
- precision
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Internal July 2026 endpointing benchmark
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
Endpointing recall (Show measurement, test conditions, and source)
- Source value
- 97
- Score
- 97
- Metric
- recall
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Internal July 2026 endpointing benchmark
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
Endpointing F1 (Show measurement, test conditions, and source)
- Source value
- 93
- Score
- 93
- Metric
- F1
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Internal July 2026 endpointing benchmark
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
VoiceCodeBench: WER (Show measurement, test conditions, and source)
- Source value
- 9.249
- Score
- 9.249
- Metric
- word error rate
- Unit
- %
- Benchmark version
- 1
- Category
- general
- Task
- wer
- Task label
- WER
- Direction
- Lower is better
- Comparison cohort
- vals-voice-code-bench:1:wer
- Source type
- independent-evaluator
- Evaluator
- Vals AI
- Status
- active
- Retrieved at
- 2026-09-29
- Methodology
- Source model ID: cartesia/ink-2. Source model name: Ink 2. Provider reported by source: Cartesia. Evaluator: Vals AI. Dataset type: public. Leaderboard date: 2026-09-24. Latency in seconds: 69.197. Standard error: 0. Source ID: vals-ai. Evidence capture ID: vals-voice-code-bench-2026-09-29-7d87ad468e81. Source payload hash: 7d87ad468e8136922daec73432384e9d40b0f52e64cc6e6a0e73db7d6e7c264c. Parser version: vals-astro-v2. Imported at: 2026-09-29
AA-AgentTalk3 WER (Show measurement, test conditions, and source)
- Source value
- 3.4
- Score
- 3.4
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Cartesia report of AA leaderboard, June 2026
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
VoxPopuli WER (Show measurement, test conditions, and source)
- Source value
- 2.4
- Score
- 2.4
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Cartesia report of AA leaderboard, June 2026
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
Earnings22 WER (Show measurement, test conditions, and source)
- Source value
- 5.3
- Score
- 5.3
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Cartesia report of AA leaderboard, June 2026
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
WER across 14 English accents (Show measurement, test conditions, and source)
- Source value
- 8
- Score
- 8
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Cartesia report; competitors Ink 8, Flux 10, Scribe 12, Universal 3.5 13
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
WER (Show measurement, test conditions, and source)
- Source value
- 6.5
- Score
- 6.5
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Cartesia Ink-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 36 real calls; July 2026 internal comparison
- Context
- Cartesia's reported values and dated Vals observations use distinct metrics and cohorts; no cross-source rank is inferred.
Published prices
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
- Price
- 3 credits per second of audio; the dollar price and overage charges depend on the planSource
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Unresolved | benchmark-cohorts Original wording Cartesia's reported WER and endpointing results, current Vals WAcc/CTEM/TSR rows, and the dated 2026-09-29 Vals WER/CTEM/TSR capture use distinct metrics or cohorts. They remain separate and are not overwritten or ranked together. |
| Additional source | Cartesia Ink-2 announcement (retrieved October 3, 2026; October 4, 2026) · Cartesia Ink-2 announcement · Editorial description reviewed October 3, 2026 |
| Additional source | Cartesia Ink product page (retrieved October 3, 2026) · Cartesia Ink product page · Editorial description reviewed October 3, 2026 |
| Additional source | Cartesia pricing (retrieved October 3, 2026) · Cartesia pricing · Editorial description reviewed October 3, 2026 |
| Additional source | Vals AI (retrieved September 29, 2026) |