Skip to main content
Audio modelPreview

MAI-Transcribe-2

Microsoft

Released
September 3, 2026
Data date
October 3, 2026
Languages

MAI-Transcribe-2 is Microsoft’s file-oriented speech-to-text endpoint for 60 languages. Speaker attribution, timestamps, and domain biasing extend the spoken words with structured transcript information.

The API identifier is mai-transcribe-2. Foundry access is a public preview without an SLA, with an introductory price of $0.10 per hour of audio. Live input uses the separate streaming variant with its own identifier.

Specifications and access

SpecificationValue and source
Model class
Speech recognitionSource
Access
Microsoft Foundry APISource
Input
Recorded audioSource
Output
Transcript with timestamps and speaker attributionSource
API model ID
mai-transcribe-2Source
Preview terms
Public preview without an SLASource
Languages
60 languagesSource
Features
Speaker diarization, timestamps, and domain biasingSource