Skip to main content
Specialist modelOpen weights

LensVLM-9B

Apple

Released
-
Data date
October 3, 2026

LensVLM-9B is Apple’s vision-language model for long documents. It scans compressed visual representations of text and expands only the pages relevant to the current question. Visual processing therefore stays selective instead of loading every page at full resolution.

The 9-billion-parameter checkpoint runs locally with Transformers, vLLM, or SGLang under Apple’s Machine Learning Research Model License. The released demo supports 5x, 10x, and 15x compression. Its model card lists no hosted Inference Provider.

Specifications and access

SpecificationValue and source
Model ID
apple/LensVLM-9BSource
Model class
Vision-language model for selective context expansionSource
Parameters
9 billionSource
Compression
5x, 10x, or 15x input compressionSource
Access
Local execution, no hosted Inference Provider on Hugging FaceSource
License
Apple Machine Learning Research Model LicenseSource