Skip to main content
Audio modelPreview

MAI-Voice-2.1-Flash

Microsoft

Released
October 1, 2026
Data date
October 3, 2026

MAI-Voice-2.1-Flash is the low-latency variant of Microsoft’s current speech synthesis line. It supports the same 23 languages and 26 locales and can generate up to 45 seconds of audio in one generation.

Microsoft cites about 150 milliseconds end to end and a price of $15 per 1M characters. The Foundry API identifier is mai-voice-2-1-flash. Access is a public preview without an SLA, and the weights are not public.

Specifications and access

SpecificationValue and source
Model class
Low-latency speech synthesisSource
Access
Microsoft Foundry APISource
Input
TextSource
Output
AudioSource
API model ID
mai-voice-2-1-flashSource
Preview terms
Public preview without an SLASource
Languages
23 languages and 26 localesSource
Latency
About 150 milliseconds end to endSource
Generation limit
Up to 45 seconds of audioSource