SmolVLM-256M-Instruct
SmolVLM-256M is the smallest multimodal model in the world. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightweight architecture makes it suitable for on-device applications while maintaining strong performance on multimodal tasks. It can run inference on one i...
Params
260 M
Context
—
Downloads 30d
872 K
Likes
398
Download history
daily snapshots · 56 days
▲ 0 in the last 30 days (0.0%)
872 K872 K
Aug 23Sep 2Sep 12Sep 21
1.1 M872 K
Jul 28Aug 15Sep 3Sep 21
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | bf16 | 0.5 GB | 1.1 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Idefics3ForConditionalGeneration
- Parameters
- 260 M
- Tensor type
- BF16
- Vocabulary
- 49,280
- Licence
- apache-2.0
- First seen on the Hub
- 2025-01-17
- Base model
- SmolLM2-135M-Instruct
- Training datasets
- HuggingFaceM4/the_cauldron, HuggingFaceM4/Docmatix
- Added to our catalog
- 2026-07-28
Family
Base model and the most-downloaded derivatives in the catalog.
Compare with any image-text-to-text model