Skip to main content
Audio modelActive

MAI-Voice-2

Microsoft

Released
June 2, 2026
Data date
October 3, 2026
Model class
Languages

MAI-Voice-2 is Microsoft’s earlier Foundry speech synthesis model with 15 supported languages. A 5 to 60 second reference recording can guide voice cloning, with Microsoft describing consent guardrails.

Microsoft released MAI-Voice-2 on June 2, 2026. Its Foundry model identifier for audio output is mai-voice-2.

Specifications and access

SpecificationValue and source
Model class
Speech synthesisSource
Access
Microsoft Foundry APISource
Input
Text and reference audioSource
Output
AudioSource
API model ID
mai-voice-2Source
Languages
15 languagesSource
Voice cloning
A 5 to 60 second reference recording with consent guardrailsSource