1,054 models · refreshed nightly

Image text to text models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
01 Qwen3-VL-8B-Instruct Qwen · image-text-to-text 8.8 B ✓ apache-2.0 19.5 M from 21.1 GB
02 gemma-4-26B-A4B-it google · image-text-to-text 26.5 B ✓ apache-2.0 9.9 M from 61.3 GB
03 Qwen3.6-35B-A3B-FP8 Qwen · image-text-to-text 36.0 B ✓ apache-2.0 9.8 M from 47.1 GB
04 Qwen3.5-9B Qwen · image-text-to-text 9.7 B ✓ apache-2.0 9.3 M from 23.2 GB
05 gemma-4-31B-it google · image-text-to-text 32.7 B ✓ apache-2.0 9.1 M from 74.2 GB
06 Qwen3.8-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.4 M from 65.8 GB
07 Qwen3.8-27B-FP8 Qwen · image-text-to-text 27.8 B ✓ apache-2.0 7.1 M from 38.6 GB
08 Qwen3.5-4B Qwen · image-text-to-text 4.7 B ✓ apache-2.0 6.9 M from 11.5 GB
09 Qwen2.5-VL-7B-Instruct Qwen · image-text-to-text 8.3 B 128 K ✓ apache-2.0 6.9 M from 20.0 GB
10 Qwen3.6-27B-FP8 Qwen · image-text-to-text 27.8 B ✓ apache-2.0 5.8 M from 38.6 GB
11 Qwen3.5-2B Qwen · image-text-to-text 2.3 B ✓ apache-2.0 4.8 M from 5.8 GB
12 Qwen3.8-27B-MLX-4bit lmstudio-community · image-text-to-text 4.7 B ✓ apache-2.0 4.7 M from 18.9 GB
13 Qwen3.8-27B-MLX-8bit lmstudio-community · image-text-to-text 8.0 B ✓ apache-2.0 4.5 M from 34.2 GB
14 Qwen3.8-27B-MLX-6bit lmstudio-community · image-text-to-text 6.4 B ✓ apache-2.0 4.5 M from 26.5 GB
15 Qwen3.8-27B-MLX-5bit lmstudio-community · image-text-to-text 5.5 B ✓ apache-2.0 4.5 M from 22.7 GB
16 Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF cdiamond · image-text-to-text ✓ apache-2.0 4.0 M from 19.3 GB
17 Qwen3-VL-4B-Instruct Qwen · image-text-to-text 4.4 B ✓ apache-2.0 3.7 M from 10.9 GB
18 Qwen3.6-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 3.6 M from 65.8 GB
19 Qwen3.6-35B-A3B Qwen · image-text-to-text 36.0 B ✓ apache-2.0 3.3 M from 85.0 GB
20 Qwen3-VL-2B-Instruct Qwen · image-text-to-text 2.1 B ✓ apache-2.0 3.0 M from 5.5 GB
21 Florence-2-base microsoft · image-text-to-text 230 M ✓ mit 3.0 M from 1.0 GB
22 GLM-5.3-Flash zai-org · image-text-to-text 321.3 B ✓ mit 2.9 M from 409.9 GB
23 Huihui-Qwen3.8-27B-abliterated-GGUF huihui-ai · image-text-to-text 27.0 B ✓ apache-2.0 2.8 M from 16.5 GB
24 chandra-ocr-2 datalab-to · image-text-to-text 5.3 B ⚠ openrail 2.7 M from 12.9 GB
25 Qwen3-VL-8B-Instruct-FP8 Qwen · image-text-to-text 8.8 B ✓ apache-2.0 2.5 M from 13.5 GB
26 Qwen2.5-VL-3B-Instruct Qwen · image-text-to-text 3.8 B 128 K unknown 2.5 M from 9.3 GB
27 Qwen3.5-0.8B Qwen · image-text-to-text 870 M ✓ apache-2.0 2.4 M from 2.6 GB
28 DeepSeek-OCR deepseek-ai · image-text-to-text 3.3 B 8 K ✓ mit 2.3 M from 8.3 GB
29 Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF HauhauCS · image-text-to-text ✓ apache-2.0 2.3 M from 1.5 GB
30 Unlimited-OCR baidu · image-text-to-text 3.3 B 33 K ✓ mit 2.2 M from 8.3 GB
31 Kimi-K3 moonshotai · image-text-to-text 2,779.9 B other 2.1 M from 2,134.5 GB
32 Qwen3.8-27B-NVFP4 RadixArk · image-text-to-text 18.2 B ✓ apache-2.0 2.1 M from 27.3 GB
33 Qwen2-VL-2B-Instruct Qwen · image-text-to-text 2.2 B 33 K ✓ apache-2.0 2.1 M from 5.7 GB
34 Qwen2.5-VL-7B-Instruct-AWQ Qwen · image-text-to-text 8.3 B 128 K ✓ apache-2.0 2.0 M from 9.4 GB
35 GLM-OCR zai-org · image-text-to-text 1.3 B ✓ mit 2.0 M from 3.6 GB
36 Qwen3.5-35B-A3B Qwen · image-text-to-text 36.0 B ✓ apache-2.0 1.9 M from 85.0 GB
37 Gemma-4-E4B-Uncensored-HauhauCS-Aggressive HauhauCS · image-text-to-text ⚠ gemma 1.9 M from 5.4 GB
38 moondream2 vikhyatk · image-text-to-text 1.9 B ✓ apache-2.0 1.9 M from 5.0 GB
39 Qwen3.5-27B Qwen · image-text-to-text 27.8 B ✓ apache-2.0 1.9 M from 65.8 GB
40 gemma-3-4b-it google · image-text-to-text · gated 4.3 B ⚠ gemma 1.8 M from 10.6 GB
41 llava-1.5-7b-hf llava-hf · image-text-to-text 7.1 B ⚠ llama2 1.8 M from 17.1 GB
42 gemma-4-26B-A4B-it-AWQ-4bit cyankiwi · image-text-to-text 25.8 B ✓ apache-2.0 1.7 M from 23.3 GB
43 Qwen3.6-27B-NVFP4 unsloth · image-text-to-text 21.2 B ✓ apache-2.0 1.7 M from 29.4 GB
44 Qwen2-VL-7B-Instruct-AWQ Qwen · image-text-to-text 8.3 B 33 K ✓ apache-2.0 1.6 M from 9.4 GB
45 Qwen3.5-35B-A3B-FP8 Qwen · image-text-to-text 36.0 B ✓ apache-2.0 1.6 M from 47.1 GB
46 Qwen3.5-9B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.5 M from 5.2 GB
47 Qwen3.8-Flash-Next-GGUF unsloth · image-text-to-text other 1.5 M from 3.6 GB
48 Qwen3.8-27B-AWQ-INT4 cyankiwi · image-text-to-text 27.8 B ✓ apache-2.0 1.5 M from 27.8 GB
49 Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF DavidAU · image-text-to-text ✓ apache-2.0 1.5 M from 5.9 GB
50 Inkling-Small-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.4 M from 54.0 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.