1,011 models · refreshed nightly

Image text to text models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 Qwen3.5-9B Qwen · image-text-to-text 9.7 B ✓ apache-2.0 12.4 M from 23.2 GB
02 gemma-4-31B-it google · image-text-to-text 32.7 B ✓ apache-2.0 11.9 M from 74.2 GB
03 gemma-4-26B-A4B-it google · image-text-to-text 26.5 B ✓ apache-2.0 11.7 M from 61.3 GB
04 Qwen2.5-VL-7B-Instruct Qwen · image-text-to-text 8.3 B 128 K ✓ apache-2.0 9.3 M from 20.0 GB
05 Qwen3.6-35B-A3B-FP8 Qwen · image-text-to-text 36.0 B ✓ apache-2.0 8.9 M from 47.1 GB
06 Qwen2.5-VL-3B-Instruct Qwen · image-text-to-text 3.8 B 128 K unknown 8.2 M from 9.3 GB
07 Qwen3.6-27B-FP8 Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.8 M from 38.6 GB
08 Qwen3.6-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.0 M from 65.8 GB
09 Qwen3.5-4B Qwen · image-text-to-text 4.7 B ✓ apache-2.0 6.7 M from 11.5 GB
10 Qwen3.6-35B-A3B Qwen · image-text-to-text 36.0 B ✓ apache-2.0 5.9 M from 85.0 GB
11 Qwen3-VL-8B-Instruct Qwen · image-text-to-text 8.8 B ✓ apache-2.0 4.9 M from 21.1 GB
12 gemma-4-31B-it-FP8-block RedHatAI · image-text-to-text 31.3 B ✓ apache-2.0 4.8 M from 41.8 GB
13 Qwen3-VL-4B-Instruct Qwen · image-text-to-text 4.4 B ✓ apache-2.0 4.0 M from 10.9 GB
14 GLM-OCR zai-org · image-text-to-text 1.3 B ✓ mit 3.8 M from 3.6 GB
15 gemma-4-26B-A4B-it-AWQ-4bit cyankiwi · image-text-to-text 26.6 B ✓ apache-2.0 3.7 M from 23.4 GB
16 Qwen3.6-27B-NVFP4 unsloth · image-text-to-text 21.2 B ✓ apache-2.0 3.4 M from 29.4 GB
17 Qwen3-VL-8B-Instruct-FP8 Qwen · image-text-to-text 8.8 B ✓ apache-2.0 3.2 M from 13.5 GB
18 chandra-ocr-2 datalab-to · image-text-to-text 5.3 B ⚠ openrail 3.2 M from 12.9 GB
19 llava-1.5-7b-hf llava-hf · image-text-to-text 7.1 B ⚠ llama2 3.1 M from 17.1 GB
20 Qwen3.5-0.8B Qwen · image-text-to-text 870 M ✓ apache-2.0 3.1 M from 2.6 GB
21 Qwen2-VL-2B-Instruct Qwen · image-text-to-text 2.2 B 33 K ✓ apache-2.0 3.0 M from 5.7 GB
22 Qwen3-VL-32B-Instruct Qwen · image-text-to-text 33.4 B ✓ apache-2.0 2.8 M from 78.9 GB
23 Florence-2-base microsoft · image-text-to-text 230 M ✓ mit 2.8 M from 1.0 GB
24 Unlimited-OCR baidu · image-text-to-text 3.3 B 33 K ✓ mit 2.7 M from 8.3 GB
25 Qwen3.5-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 2.7 M from 65.8 GB
26 Qwen3.5-2B Qwen · image-text-to-text 2.3 B ✓ apache-2.0 2.6 M from 5.8 GB
27 moondream2 vikhyatk · image-text-to-text 1.9 B ✓ apache-2.0 2.6 M from 5.0 GB
28 Qwen3.6-27B-AWQ-INT4 cyankiwi · image-text-to-text 29.3 B ✓ apache-2.0 2.6 M from 27.4 GB
29 Qwen3.5-35B-A3B Qwen · image-text-to-text 36.0 B ✓ apache-2.0 2.6 M from 85.0 GB
30 DeepSeek-OCR deepseek-ai · image-text-to-text 3.3 B 8 K ✓ mit 2.3 M from 8.3 GB
31 gemma-4-31B-it-qat-w4a16-ct google · image-text-to-text 33.6 B ✓ apache-2.0 2.3 M from 31.1 GB
32 Qwen3.5-122B-A10B Qwen · image-text-to-text 125.1 B ✓ apache-2.0 2.2 M from 294.5 GB
33 Qwen3-VL-2B-Instruct Qwen · image-text-to-text 2.1 B ✓ apache-2.0 2.2 M from 5.5 GB
34 DeepSeek-OCR-2 deepseek-ai · image-text-to-text 3.4 B 8 K ✓ apache-2.0 2.1 M from 8.5 GB
35 diffusiongemma-26B-A4B-it google · image-text-to-text 25.8 B ✓ apache-2.0 2.0 M from 61.2 GB
36 Qwen2.5-VL-7B-Instruct-AWQ Qwen · image-text-to-text 8.3 B 128 K ✓ apache-2.0 2.0 M from 9.4 GB
37 Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive HauhauCS · image-text-to-text ✓ apache-2.0 1.9 M from 13.3 GB
38 Qwen3.6-35B-A3B-NVFP4 unsloth · image-text-to-text 24.6 B ✓ apache-2.0 1.8 M from 33.3 GB
39 gemma-3-4b-it google · image-text-to-text · gated 4.3 B ⚠ gemma 1.8 M from 10.6 GB
40 Qwen2-VL-7B-Instruct-AWQ Qwen · image-text-to-text 8.3 B 33 K ✓ apache-2.0 1.8 M from 9.4 GB
41 Qwen3-VL-235B-A22B-Instruct Qwen · image-text-to-text 235.7 B ✓ apache-2.0 1.8 M from 554.3 GB
42 Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF DavidAU · image-text-to-text ✓ apache-2.0 1.6 M from 13.8 GB
43 Qwen2-VL-7B-Instruct Qwen · image-text-to-text 8.3 B 33 K ✓ apache-2.0 1.5 M from 20.0 GB
44 InternVL2-2B OpenGVLab · image-text-to-text 2.2 B ✓ mit 1.5 M from 5.7 GB
45 gemma-4-26B-A4B-it-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.4 M from 11.4 GB
46 Qwen2.5-VL-32B-Instruct-AWQ Qwen · image-text-to-text 33.5 B 128 K ✓ apache-2.0 1.4 M from 28.3 GB
47 Llama-3.1-Nemotron-Nano-VL-8B-V1 nvidia · image-text-to-text 8.7 B other 1.3 M from 21.0 GB
48 Qwen3.5-9B-AWQ QuantTrio · image-text-to-text 9.7 B ✓ apache-2.0 1.3 M from 15.6 GB
49 Phi-3.5-vision-instruct microsoft · image-text-to-text 4.2 B 131 K ✓ mit 1.3 M from 10.2 GB
50 gemma-3-12b-it google · image-text-to-text · gated 12.2 B ⚠ gemma 1.3 M from 29.1 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.