Skip to main content
Language modelOpen weights

Ling-3.0-flash-VL

InclusionAI

Released
-
Data date
October 3, 2026

Ling-3.0-flash-VL adds native image and video input to Ling 3.0 Flash. VideoRoPE represents spatial positions and temporal order, supporting questions about events in longer videos. InclusionAI adds a visual encoder to the language backbone for this purpose.

The MoE model activates 5.5 of its 124 billion parameters per token and handles up to 256K context tokens. Its local weights use MIT.

Specifications and access

SpecificationValue and source
Model ID
inclusionAI/Ling-3.0-flash-VLSource
Model class
Multimodal language model with a vision encoderSource
Parameters
124 billion total, 5.5 billion activeSource
Context window
Up to 256K tokens according to the model cardSource
Architecture
ViT vision encoder, MLP projector, VideoRoPE, and sparse MoE backboneSource
Input
Text, images, and videoSource
License