Skip to main content
Audio modelOpen weights

MOSS-Audio-4B-Thinking

OpenMOSS Team

Released
April 13, 2026
Data date
October 3, 2026

MOSS-Audio-4B-Thinking combines audio understanding with text reasoning. The roughly 4.6-billion-parameter model accepts audio and text and can answer questions, caption recordings, or produce timestamped transcripts.

The output remains text. Despite its audio input, MOSS-Audio-4B-Thinking is not a speech synthesizer and does not produce a spoken answer. The weights are available locally under Apache-2.0.

Specifications and access

SpecificationValue and source
Model class
Audio understanding and text reasoningSource
Access
Local weightsSource
Input
Audio and textSource
Output
TextSource
Checkpoint
OpenMOSS-Team/MOSS-Audio-4B-ThinkingSource
License
Apache-2.0Source
Tasks
ASR, audio question answering, audio captioning, and timestamped ASRSource
Reasoning
Generates text answers with a Qwen3-4B backbone, not spoken audio responsesSource
Repository
Official OpenMOSS Audio repositorySource