Audio modelOpen weights
Confucius4-R2T2
NetEase Youdao
- Released
- -
- Data date
- October 3, 2026
- Model class
- Access
Confucius4-R2T2 is a streaming ASR model for Chinese and English. It builds on Qwen3-ASR-1.7B, keeps an append-only stable transcript prefix, and allows decoding chunks from 80 milliseconds to two seconds.
The model card reports roughly 200 to 600 milliseconds of average latency. The weights run locally under the NetEase Model Use License Agreement.
Specifications and access
| Specification | Value and source |
|---|---|
| Model class | Streaming speech recognitionSource |
| Access | Local weightsSource |
| Input | Chinese and English audioSource |
| Output | Stable streaming transcriptSource |
| Checkpoint | netease-youdao/Confucius4-R2T2Source |
| License | NetEase Model Use License AgreementSource |
| Base model | Built on Qwen/Qwen3-ASR-1.7BSource |
| Latency | About 200 to 600 milliseconds on averageSource |
| Decoding | Configurable chunks from 80 milliseconds to 2 secondsSource |
Page 1 of 2
Measurements without matching peer values
These measurements have no matching peer values under the same test conditions. Their original values and sources remain available here.
AMI WER (Show measurement, test conditions, and source)
- Source value
- 11.37
- Score
- 11.37
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
GigaSpeech clean WER (Show measurement, test conditions, and source)
- Source value
- 9.6
- Score
- 9.6
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
LibriSpeech clean WER (Show measurement, test conditions, and source)
- Source value
- 2.13
- Score
- 2.13
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
LibriSpeech other WER (Show measurement, test conditions, and source)
- Source value
- 4.88
- Score
- 4.88
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
SPGISpeech WER (Show measurement, test conditions, and source)
- Source value
- 3
- Score
- 3
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
VoxPopuli WER (Show measurement, test conditions, and source)
- Source value
- 3.07
- Score
- 3.07
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
Earnings22 WER (Show measurement, test conditions, and source)
- Source value
- 9.36
- Score
- 9.36
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
TED-LIUM WER (Show measurement, test conditions, and source)
- Source value
- 3.34
- Score
- 3.34
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
EN-RealSI WER (Show measurement, test conditions, and source)
- Source value
- 8.4
- Score
- 8.4
- Metric
- word error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration; true streaming, append-only
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
Wenet-net CER (Show measurement, test conditions, and source)
- Source value
- 5.87
- Score
- 5.87
- Metric
- character error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
Wenet-meeting CER (Show measurement, test conditions, and source)
- Source value
- 7.27
- Score
- 7.27
- Metric
- character error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
SPEECHIO-06 CER (Show measurement, test conditions, and source)
- Source value
- 7.3
- Score
- 7.3
- Metric
- character error rate
- Unit
- %
- Benchmark version
- source version not stated
- Category
- audio
- Direction
- Lower is better
- Source type
- vendor-reported
- Evaluator
- Confucius4-R2T2 model card
- Status
- active
- Retrieved at
- 2026-10-04
- Methodology
- Vendor-reported result. 160 ms representative streaming configuration
- Context
- Source-specific observation; it does not establish a compatible cross-model cohort.
Page 1 of 2
Sources and data date
Every statement links to its underlying documentation or leaderboard.
| Type | Evidence and data date |
|---|---|
| Research status | Research date October 4, 2026. 1 source URLs checked. This documents the inspected sources, not an exhaustive inventory of every publication. |
| Additional source | Confucius4-R2T2 model card (retrieved October 3, 2026; October 4, 2026) · Confucius4-R2T2 model card · Editorial description reviewed October 3, 2026 |