SmolVLM2-500M-Video-Instruct
SmolVLM2-500M-Video is a lightweight multimodal model designed to analyze video content. The model processes videos, images, and text inputs to generate text outputs - whether answering questions about media files, comparing visual content, or transcribing text from images. Despite its compact size, requiring only 1.8GB of GPU RAM for video inference, it delivers robust performance on complex multimodal tasks. This e...
Params
510 M
Context
—
Downloads 30d
1.3 M
Likes
178
Download history
daily snapshots · 56 days
▲ 256 K in the last 30 days (17.0%)
1.5 M1.3 M
Aug 23Sep 2Sep 12Sep 21
1.5 M995 K
Jul 28Aug 15Sep 3Sep 21
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f32 | 2.0 GB | 2.8 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- SmolVLMForConditionalGeneration
- Parameters
- 510 M
- Tensor type
- F32
- Vocabulary
- 49,280
- Licence
- apache-2.0
- First seen on the Hub
- 2025-02-11
- Training datasets
- HuggingFaceM4/the_cauldron, HuggingFaceM4/Docmatix, lmms-lab/LLaVA-OneVision-Data, lmms-lab/M4-Instruct-Data, HuggingFaceFV/finevideo, MAmmoTH-VL/MAmmoTH-VL-Instruct-12M, lmms-lab/LLaVA-Video-178K, orrzohar/Video-STaR, Mutonix/Vript, TIGER-Lab/VISTA-400K, Enxin/MovieChat-1K_train, ShareGPT4Video/ShareGPT4Video
- Added to our catalog
- 2026-07-28
Compare with any image-text-to-text model
Popular comparisons