Audio modelPreview
MAI-Voice-2.1
Microsoft
- Released
- October 1, 2026
- Data date
- October 3, 2026
- Model class
- Languages
- Access
MAI-Voice-2.1 turns text into natural-sounding speech across 23 languages and 26 locales. A short reference recording can guide the voice, with Microsoft documenting consent guardrails for cloning.
Access is provided through Microsoft Foundry as a public preview without an SLA. Microsoft lists a price of $22 per 1M characters. MAI-Voice-2.1 is for generated audio, not transcription or general audio understanding.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Speech synthesisSource |
| Access | Microsoft Foundry APISource |
| Input | TextSource |
| Output | AudioSource |
| API model ID | mai-voice-2-1Source |
| Preview terms | Public preview without an SLASource |
| Languages | 23 languages and 26 localesSource |
| Voice cloning | A few seconds of reference audio with consent guardrailsSource |
Published prices
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
- Price
- $22 per 1M charactersSource
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Not published | quality-benchmarks Original wording Only an operational latency figure was found; no exact Voice-2.1 quality/MOS benchmark. |
| Additional source | Microsoft MAI Voice preview documentation (retrieved October 3, 2026) · Microsoft MAI Voice preview documentation · Editorial description reviewed October 3, 2026 |
| Additional source | Microsoft MAI Voice announcement (retrieved October 3, 2026) · Microsoft MAI Voice announcement · Editorial description reviewed October 3, 2026 |
| Additional source | Microsoft MAI-Voice-2.1 model page (retrieved October 3, 2026; October 4, 2026) · Microsoft MAI-Voice-2.1 model page · Editorial description reviewed October 3, 2026 |