Skip to main content
Audio modelPreview

MAI-Transcribe-1.5

Microsoft

Released
June 2, 2026
Data date
October 3, 2026
Languages

MAI-Transcribe-1.5 is Microsoft’s speech-to-text model for 43 languages. Keyword biasing helps with domain terms, and Microsoft cites up to five times the speed of comparable models.

The endpoint accepts audio and returns text. Foundry access is a public preview without an SLA. Microsoft lists $0.36 per hour of audio. It is intended for transcription, not speech synthesis or audio-to-audio dialogue.

Specifications and access

SpecificationValue and source
Model class
Speech recognitionSource
Access
Microsoft Foundry APISource
Input
AudioSource
Output
TranscriptSource
API model ID
mai-transcribe-1.5Source
Preview terms
Public preview without an SLASource
Languages
43 languagesSource
Features
Keyword biasing and up to five times the speed of comparable modelsSource