HuggingFaceTB / image-text-to-text updated 1 year ago

SmolVLM-256M-Instruct

SmolVLM-256M is the smallest multimodal model in the world. It accepts arbitrary sequences of image and text inputs to produce text outputs. It's designed for efficiency. SmolVLM can answer questions about images, describe visual content, or transcribe text. Its lightweight architecture makes it suitable for on-device applications while maintaining strong performance on multimodal tasks. It can run inference on one i...

Params
260 M
Context
Downloads 30d
1.1 M
Likes
395
Commercial use: allowed apache-2.0 Not gated SAFETENSORS 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
1.1 M1.0 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 0.5 GB 1.1 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Idefics3ForConditionalGeneration
Parameters
260 M
Tensor type
BF16
Vocabulary
49,280
Licence
apache-2.0
First seen on the Hub
2025-01-17
Base model
SmolLM2-135M-Instruct
Training datasets
HuggingFaceM4/the_cauldron, HuggingFaceM4/Docmatix
Added to our catalog
2026-07-28

Family

Base model and the most-downloaded derivatives in the catalog.