Audio modelActive
MAI-Voice-2
Microsoft
- Released
- June 2, 2026
- Data date
- October 3, 2026
- Model class
- Languages
- Access
MAI-Voice-2 is Microsoft’s earlier Foundry speech synthesis model with 15 supported languages. A 5 to 60 second reference recording can guide voice cloning, with Microsoft describing consent guardrails.
Microsoft released MAI-Voice-2 on June 2, 2026. Its Foundry model identifier for audio output is mai-voice-2.
Specifications and access
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
win rate versus Voice 1 (Show measurement, test conditions, and source)
- Source value
- 72.1
- Score
- 72.1
- Metric
- win rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Higher is better
- Source type
- vendor-reported
- Evaluator
- Microsoft MAI-Voice-2 announcement
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 2,500 listening tests; 2,222 responses; 11 languages; preference test, not MOS
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Microsoft MAI Voice 2 announcement (retrieved October 3, 2026; October 4, 2026) · Microsoft MAI Voice 2 announcement · Editorial description reviewed October 3, 2026 |
| Additional source | Microsoft MAI Voice 2 Flash announcement (retrieved October 3, 2026) · Microsoft MAI Voice 2 Flash announcement · Editorial description reviewed October 3, 2026 |