MAI-Transcribe-2-Streaming
Microsoft
- Released
- October 1, 2026
- Data date
- October 3, 2026
- Model class
- Access
MAI-Transcribe-2-Streaming accepts live audio and returns real-time transcripts. Microsoft describes continuous language detection across 60 languages, so language changes do not need to wait for the end of a recording.
The variant is available as a Microsoft Foundry preview under mai-transcribe-2-streaming. Its announced introductory price is $0.54 per hour of audio through the end of 2026.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Streaming speech recognitionSource |
| Access | Microsoft Foundry APISource |
| Input | Live audioSource |
| Output | Streaming transcriptSource |
| API model ID | mai-transcribe-2-streamingSource |
| Languages | 60 languages with continuous language detectionSource |
| Streaming | Real-time transcription for ongoing audio inputSource |
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
Artificial Analysis final transcription WER (Show measurement, test conditions, and source)
- Source value
- 2.5
- Score
- 2.5
- Metric
- word error rate
- Unit
- %
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- MAI-Transcribe-2-Streaming Model Card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. Final-transcription AA-WER Streaming after detected End-of-Speech; September 28, 2026 leaderboard snapshot in Microsoft's model card, page 2. Published evaluator method: AA-AgentTalk 50%, VoxPopuli-Cleaned-AA 25%, Earnings22-Cleaned-AA 25%; WER duration-weighted within each dataset.
- Context
- Microsoft reproduces an Artificial Analysis result; the evaluator's own model record was not retrieved directly. Exact serving configuration, sample count for this run, benchmark version in the memo, and run date are unspecified. Snapshot date is not run date; 60 supported languages do not establish test language coverage.
Published prices
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
- Introductory price
- $0.54 per hour of audio through the end of 2026Source
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 4 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Not published | evaluation-protocol Original wording The model card and evaluator methodology establish final-transcription WER, snapshot date, and dataset weighting. Exact serving configuration, run-specific sample count, run date, and a directly retrieved model record remain unavailable. |
| Additional source | Microsoft streaming transcription announcement (retrieved October 3, 2026; October 4, 2026) · Microsoft streaming transcription announcement · Editorial description reviewed October 3, 2026 |
| Additional source | Microsoft MAI-Transcribe-2 model page (retrieved October 3, 2026; October 4, 2026) · Microsoft MAI-Transcribe-2 model page · Editorial description reviewed October 3, 2026 |
| Additional source | MAI-Transcribe-2-Streaming Model Card (retrieved October 4, 2026) |
| Additional source | Model metadata (retrieved October 4, 2026) |