Skip to main content
Audio modelOpen weights

Confucius4-R2T2

NetEase Youdao

Released
-
Data date
October 3, 2026

Confucius4-R2T2 is a streaming ASR model for Chinese and English. It builds on Qwen3-ASR-1.7B, keeps an append-only stable transcript prefix, and allows decoding chunks from 80 milliseconds to two seconds.

The model card reports roughly 200 to 600 milliseconds of average latency. The weights run locally under the NetEase Model Use License Agreement.

Specifications and access

SpecificationValue and source
Model class
Streaming speech recognitionSource
Access
Local weightsSource
Input
Chinese and English audioSource
Output
Stable streaming transcriptSource
Checkpoint
netease-youdao/Confucius4-R2T2Source
License
NetEase Model Use License AgreementSource
Base model
Built on Qwen/Qwen3-ASR-1.7BSource
Latency
About 200 to 600 milliseconds on averageSource
Decoding
Configurable chunks from 80 milliseconds to 2 secondsSource