Updated 2026-09-20 · ranked by real download data

Fastest-growing VLMs

Ranked nightly from our download snapshots of the Hugging Face catalog. Every entry shows its licence and what hardware it realistically needs.

#ModelParamsContextCommercial use30dMin VRAM
01 GLM-5.3-Flash zai-org · image-text-to-text 321.3 B ✓ mit 2.9 M from 755.6 GB
02 Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF cdiamond · image-text-to-text ✓ apache-2.0 4.0 M
03 Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF DavidAU · image-text-to-text ✓ apache-2.0 1.5 M
04 Qwen3.6-27B-MTP-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.2 M
05 Qwen3.8-Flash-Next-GGUF unsloth · image-text-to-text other 1.5 M
06 Qwen3.5-2B Qwen · image-text-to-text 2.3 B ✓ apache-2.0 4.8 M from 5.8 GB
07 Qwen3.6-35B-A3B-MTP-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.0 M
08 Qwen3-VL-8B-Instruct Qwen · image-text-to-text 8.8 B ✓ apache-2.0 19.5 M from 21.1 GB
09 moondream2 vikhyatk · image-text-to-text 1.9 B ✓ apache-2.0 1.9 M from 5.0 GB
10 gemma-4-26B-A4B-it google · image-text-to-text 26.5 B ✓ apache-2.0 9.9 M from 62.9 GB
11 Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF HauhauCS · image-text-to-text ✓ apache-2.0 2.3 M
12 Huihui-Qwen3.8-27B-abliterated-GGUF huihui-ai · image-text-to-text 27.0 B ✓ apache-2.0 2.8 M from 64.0 GB
13 Qwen3.8-27B-MLX-5bit lmstudio-community · image-text-to-text 5.5 B ✓ apache-2.0 4.5 M from 13.4 GB
14 Florence-2-base microsoft · image-text-to-text 230 M ✓ mit 3.0 M from 1.0 GB
15 Qwen3.8-27B-MLX-6bit lmstudio-community · image-text-to-text 6.4 B ✓ apache-2.0 4.5 M from 15.4 GB
16 gemma-3-4b-it google · image-text-to-text 4.3 B ⚠ gemma 1.8 M from 10.6 GB
17 medgemma-4b-it google · image-text-to-text 4.3 B other 1.1 M from 10.6 GB
18 gemma-4-31B-it google · image-text-to-text 32.7 B ✓ apache-2.0 9.1 M from 77.3 GB
19 XYZAILab_XYZ-Aquila-mini-GGUF bartowski · image-text-to-text ✓ apache-2.0 878 K
20 surya-ocr-2 datalab-to · image-text-to-text 690 M ⚠ openrail 1.2 M from 2.1 GB
21 endless-frontier_BigBang-v1-GGUF bartowski · image-text-to-text ✓ apache-2.0 984 K
22 Qwen3-VL-2B-Instruct Qwen · image-text-to-text 2.1 B ✓ apache-2.0 3.0 M from 5.5 GB
23 Qwen2-VL-7B-Instruct-AWQ Qwen · image-text-to-text 8.3 B 33 K ✓ apache-2.0 1.6 M from 20.0 GB
Membership and ranking refresh nightly after the snapshot run. VRAM is an estimate for the smallest quantization at 8K context — methodology.