Skip to main content
Audio modelPreview

MAI-Voice-2.1

Microsoft

Released
October 1, 2026
Data date
October 3, 2026
Model class

MAI-Voice-2.1 turns text into natural-sounding speech across 23 languages and 26 locales. A short reference recording can guide the voice, with Microsoft documenting consent guardrails for cloning.

Access is provided through Microsoft Foundry as a public preview without an SLA. Microsoft lists a price of $22 per 1M characters. MAI-Voice-2.1 is for generated audio, not transcription or general audio understanding.

Specifications and access

SpecificationValue and source
Model class
Speech synthesisSource
Access
Microsoft Foundry APISource
Input
TextSource
Output
AudioSource
API model ID
mai-voice-2-1Source
Preview terms
Public preview without an SLASource
Languages
23 languages and 26 localesSource
Voice cloning
A few seconds of reference audio with consent guardrailsSource