Updated 2026-08-06 · ranked by real download data

Fastest-growing VLMs

Ranked nightly from our download snapshots of the Hugging Face catalog. Every entry shows its licence and what hardware it realistically needs.

#ModelParamsContextCommercial use30dMin VRAM
01 Qwen3.6-35B-A3B-NVFP4 unsloth · image-text-to-text 24.6 B ✓ apache-2.0 1.8 M from 58.4 GB
02 Qwen3.5-122B-A10B Qwen · image-text-to-text 125.1 B ✓ apache-2.0 2.2 M from 294.5 GB
03 surya-ocr-2 datalab-to · image-text-to-text 690 M ⚠ openrail 1.2 M from 2.1 GB
04 Qwen2.5-VL-32B-Instruct Qwen · image-text-to-text 33.5 B 128 K ✓ apache-2.0 1.1 M from 79.1 GB
05 gemma-3-27b-it-int4-awq gaunernst · image-text-to-text 27.4 B ⚠ gemma 1.2 M from 65.0 GB
06 Qwen3.6-27B-NVFP4 unsloth · image-text-to-text 21.2 B ✓ apache-2.0 3.4 M from 50.4 GB
07 SmolVLM2-500M-Video-Instruct HuggingFaceTB · image-text-to-text 510 M ✓ apache-2.0 1.2 M from 1.7 GB
08 Qwen3.5-9B-AWQ QuantTrio · image-text-to-text 9.7 B ✓ apache-2.0 1.3 M from 23.2 GB
09 Qwen3-VL-8B-Instruct-FP8 Qwen · image-text-to-text 8.8 B ✓ apache-2.0 3.2 M from 21.1 GB
10 Qwen3.6-35B-A3B-FP8 Qwen · image-text-to-text 36.0 B ✓ apache-2.0 8.9 M from 85.0 GB
11 Qwen3.6-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.0 M from 65.8 GB
12 moondream2 vikhyatk · image-text-to-text 1.9 B ✓ apache-2.0 2.6 M from 5.0 GB
13 Qwen3.5-9B Qwen · image-text-to-text 9.7 B ✓ apache-2.0 12.4 M from 23.2 GB
14 Qwen3.6-27B-AWQ-INT4 cyankiwi · image-text-to-text 29.3 B ✓ apache-2.0 2.6 M from 69.4 GB
15 Qwen3.5-122B-A10B-FP8 Qwen · image-text-to-text 125.1 B ✓ apache-2.0 1.3 M from 294.5 GB
16 chandra-ocr-2 datalab-to · image-text-to-text 5.3 B ⚠ openrail 3.2 M from 13.0 GB
17 Qwen3.5-2B Qwen · image-text-to-text 2.3 B ✓ apache-2.0 2.6 M from 5.8 GB
18 Qwen3.5-0.8B Qwen · image-text-to-text 870 M ✓ apache-2.0 3.1 M from 2.5 GB
19 Qwen3.6-27B-FP8 Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.8 M from 65.8 GB
20 deepseek-vl2-tiny deepseek-ai · image-text-to-text 3.4 B other 1.3 M from 8.4 GB
21 diffusiongemma-26B-A4B-it google · image-text-to-text 25.8 B ✓ apache-2.0 2.0 M from 61.2 GB
22 gemma-4-31B-it-qat-w4a16-ct google · image-text-to-text 33.6 B ✓ apache-2.0 2.3 M from 79.5 GB
23 GLM-OCR zai-org · image-text-to-text 1.3 B ✓ mit 3.8 M from 3.6 GB
24 Qwen2.5-VL-3B-Instruct Qwen · image-text-to-text 3.8 B 128 K unknown 8.2 M from 9.3 GB
25 Qwen3.5-4B Qwen · image-text-to-text 4.7 B ✓ apache-2.0 6.7 M from 11.5 GB
Membership and ranking refresh nightly after the snapshot run. VRAM is an estimate for the smallest quantization at 8K context — methodology.