Qwen3.8-Omni-Flash-Realtime
Alibaba
- Released
- September 21, 2026
- Data date
- October 3, 2026
- Model class
Qwen3.8-Omni-Flash-Realtime connects live text, audio, and video input to text and audio output. Its realtime documentation lists WebSocket, WebRTC, and AOQ transports; function calling and MCP form the documented action path.
The variant has a 196,608-token input context and up to 65,536 output tokens. Rates differ between Beijing and Singapore. Speech output is billed for both audio and its corresponding text. No separate new weights checkpoint is published.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Omnimodal real-time modelSource |
| Access | Alibaba Cloud Model Studio APISource |
| Input | Text, audio, and videoSource |
| Output | Text and audioSource |
| API model ID | qwen3.8-omni-flash-realtimeSource |
| Audio history | Up to 100 audio turns or 600 seconds; video up to 50 turns or 240 secondsSource |
| Transport | WebSocket, WebRTC, or AOQSource |
| Classification | Real-time API variant without separately published new weightsSource |
| Singapore rate limit | 60 requests and 2M tokens per minuteSource |
Page 1 of 4
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
SEED content stability, Chinese (Show measurement, test conditions, and source)
- Source value
- 0.74
- Score
- 0.74
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; SEED subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SEED content stability, English (Show measurement, test conditions, and source)
- Source value
- 0.89
- Score
- 0.89
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; SEED subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SEED content stability, hard (Show measurement, test conditions, and source)
- Source value
- 5.29
- Score
- 5.29
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; SEED subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SpeechSuperClue content stability, multilingual (Show measurement, test conditions, and source)
- Source value
- 3.34
- Score
- 3.34
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; multilingual / cross-lingual subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SpeechSuperClue content stability, cross-lingual (Show measurement, test conditions, and source)
- Source value
- 4
- Score
- 4
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; multilingual / cross-lingual subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SwanBench-Speech content stability, Chinese (Show measurement, test conditions, and source)
- Source value
- 1.96
- Score
- 1.96
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; SwanBench-Speech subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SwanBench-Speech content stability, English (Show measurement, test conditions, and source)
- Source value
- 2.37
- Score
- 2.37
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; SwanBench-Speech subsets
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
LongSpeechGeneration content stability, Chinese (Show measurement, test conditions, and source)
- Source value
- 1.74
- Score
- 1.74
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; long-horizon generation
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
LongSpeechGeneration content stability, English (Show measurement, test conditions, and source)
- Source value
- 1.59
- Score
- 1.59
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 11
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. zero-shot voice cloning; long-horizon generation
- Context
- Provider-reported zero-shot voice-cloning result; it is not an independently reproduced cohort.
SEED Custom Voice content stability, Chinese (Show measurement, test conditions, and source)
- Source value
- 0.83
- Score
- 0.83
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 12
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. custom-voice evaluation; table notes define WER/CER lower, similarity/accuracy/win rate higher
- Context
- Provider-reported Custom Voice result; it is not an independently reproduced cohort.
SEED Custom Voice content stability, English (Show measurement, test conditions, and source)
- Source value
- 0.93
- Score
- 0.93
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 12
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. custom-voice evaluation; table notes define WER/CER lower, similarity/accuracy/win rate higher
- Context
- Provider-reported Custom Voice result; it is not an independently reproduced cohort.
SEED Custom Voice content stability, hard (Show measurement, test conditions, and source)
- Source value
- 5.22
- Score
- 5.22
- Metric
- content-stability WER/CER
- Unit
- %
- Benchmark version
- arXiv:2609.25611v1, Table 12
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Qwen3.8-Omni technical report
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. custom-voice evaluation; table notes define WER/CER lower, similarity/accuracy/win rate higher
- Context
- Provider-reported Custom Voice result; it is not an independently reproduced cohort.
Page 1 of 4
Published prices
Prices apply to the stated unit. Resolution, output length, and provider can change the cost.
- China (Beijing) pricing
- Per 1M tokens in CNY. Audio input 6, audio output 12; text/image/video input 1.5, text output 4.5Source
- Singapore pricing
- Per 1M tokens in CNY. Audio input 6.781, audio output 13.636; text/image/video input 1.677, text output 5.104Source
- Speech billing
- Speech output is billed as audio and its corresponding textSource
- Free quota
- 1M tokens in China (Beijing), valid for 90 days; not automatically applicable to other regionsSource
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Not found in the inspected sources | independent-verification Original wording No independent reproduction of the exact realtime checkpoint's Table 10-12 results was found in the checked sources. |
| Additional source | Alibaba Cloud Qwen3.8 Omni Realtime documentation (retrieved October 3, 2026) · Alibaba Cloud Qwen3.8 Omni Realtime documentation · Editorial description reviewed October 3, 2026 |
| Additional source | Alibaba Cloud newly released models (retrieved October 3, 2026) · Alibaba Cloud newly released models · Editorial description reviewed October 3, 2026 |
| Additional source | Alibaba Cloud regional pricing and limits (retrieved October 3, 2026) · Alibaba Cloud regional Realtime pricing · Editorial description reviewed October 3, 2026 |
| Additional source | Qwen3.8-Omni technical report (retrieved October 4, 2026) |