Skip to main content
Audio modelActive

Cartesia Ink-2

Cartesia

Released
July 9, 2026
Data date
October 3, 2026

Cartesia Ink-2 is a streaming ASR model for voice agents. Cartesia reports about 0.1 seconds from the end of speech to the final transcript (TTFT), not to the first partial result. It adds semantic endpointing and structured entities to transcription.

The ink-2 endpoint is available through the Cartesia API and Play. Its launch named English, while a newer product page also lists Spanish, French, Hindi, and Japanese. The official sources do not give a consistent language list. Billing uses three credits per second of audio, with the dollar price depending on the plan.

Specifications and access

SpecificationValue and source
Model class
Streaming speech recognitionSource
Access
Cartesia API and PlaySource
Input
AudioSource
Output
Transcript and structured entitiesSource
API model ID
ink-2Source
Languages
English at launch according to the announcement; the current Ink product page also lists Spanish, French, Hindi, and Japanese. Language-coverage claims are not consistent.Source
Time to final transcript
Provider-reported 0.1 seconds from the end of speech to the final transcript (TTFT), not to the first partial resultSource
Features
Semantic endpointing and structured entity recognitionSource