1,054 models · refreshed nightly

Image text to text models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
51 Qwen3.6-35B-A3B-NVFP4 unsloth · image-text-to-text 24.6 B ✓ apache-2.0 1.3 M from 33.3 GB
52 Unlimited-OCR-AWQ sahilchachra · image-text-to-text 3.4 B 33 K ✓ mit 1.3 M from 4.1 GB
53 Qwen3.6-35B-A3B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.3 M from 11.6 GB
54 SmolVLM2-500M-Video-Instruct HuggingFaceTB · image-text-to-text 510 M ✓ apache-2.0 1.3 M from 2.8 GB
55 Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF DavidAU · image-text-to-text ✓ apache-2.0 1.3 M from 13.8 GB
56 gemma-3-27b-it-int4-awq gaunernst · image-text-to-text 27.4 B ⚠ gemma 1.3 M from 24.9 GB
57 surya-ocr-2 datalab-to · image-text-to-text 690 M ⚠ openrail 1.2 M from 2.1 GB
58 Qwen3.8-27B-NVFP4 Inferact · image-text-to-text 17.6 B ✓ apache-2.0 1.2 M from 32.2 GB
59 Qwen3.8-27B-GGUF ggml-org · image-text-to-text 27.0 B ✓ apache-2.0 1.2 M from 6.4 GB
60 Qwen3.6-27B-MTP-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.2 M from 14.3 GB
61 Qwen3.8-27B-GSQ-RCO-GGUF ISTA-DASLab · image-text-to-text ✓ apache-2.0 1.2 M from 1.5 GB
62 Qwen2.5-VL-32B-Instruct Qwen · image-text-to-text 33.5 B 128 K ✓ apache-2.0 1.1 M from 80.6 GB
63 medgemma-4b-it google · image-text-to-text · gated 4.3 B other 1.1 M from 10.6 GB
64 gemma-4-31B-it-NVFP4 RedHatAI · image-text-to-text 32.7 B ✓ apache-2.0 1.1 M from 31.0 GB
65 surya-ocr-2-gguf datalab-to · image-text-to-text ⚠ openrail 1.1 M from 1.9 GB
66 Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V10-GGUF LuffyTheFox · image-text-to-text ✓ apache-2.0 1.1 M from 29.9 GB
67 Qwen3.6-27B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.1 M from 14.1 GB
68 Muse-Glimmer-30B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.1 M from 12.3 GB
69 Qwen2.5-VL-32B-Instruct-AWQ Qwen · image-text-to-text 33.5 B 128 K ✓ apache-2.0 1.1 M from 28.3 GB
70 Qwen3.6-35B-A3B-MTP-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.0 M from 13.0 GB
71 Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V11-GGUF LuffyTheFox · image-text-to-text ✓ apache-2.0 1.0 M from 29.9 GB
72 Qwen3.5-122B-A10B Qwen · image-text-to-text 125.1 B ✓ apache-2.0 1.0 M from 294.5 GB
73 Qwen3.5-9B-AWQ QuantTrio · image-text-to-text 9.7 B ✓ apache-2.0 1.0 M from 15.6 GB
74 Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V9-GGUF LuffyTheFox · image-text-to-text ✓ apache-2.0 998 K from 29.8 GB
75 endless-frontier_BigBang-v1-GGUF bartowski · image-text-to-text ✓ apache-2.0 984 K from 11.8 GB
76 Qwen3.6-35B-A3B-NVFP4 RedHatAI · image-text-to-text 34.7 B ✓ apache-2.0 979 K from 33.2 GB
77 Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive HauhauCS · image-text-to-text ✓ apache-2.0 974 K from 13.3 GB
78 Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V12-GGUF LuffyTheFox · image-text-to-text ✓ apache-2.0 965 K from 29.9 GB
79 Qwen3.5-122B-A10B-FP8 Qwen · image-text-to-text 125.1 B ✓ apache-2.0 954 K from 159.1 GB
80 gemma-4-26B-A4B-it-GGUF unsloth · image-text-to-text ✓ apache-2.0 952 K from 11.4 GB
81 Qwen3.6-27B-AWQ-INT4 cyankiwi · image-text-to-text 29.3 B ✓ apache-2.0 941 K from 27.4 GB
82 gemma-4-26B-A4B-it-FP8-dynamic RedHatAI · image-text-to-text 26.5 B ✓ apache-2.0 941 K from 36.0 GB
83 Qwen3.5-4B-GGUF unsloth · image-text-to-text ✓ apache-2.0 905 K from 2.8 GB
84 DeepSeek-OCR-2 deepseek-ai · image-text-to-text 3.4 B 8 K ✓ apache-2.0 889 K from 8.5 GB
85 XYZAILab_XYZ-Aquila-mini-GGUF bartowski · image-text-to-text ✓ apache-2.0 878 K from 11.3 GB
86 Tiel-Coder-35B-A3B-GGUF-MTP peculiar-ragdoll · image-text-to-text 35.0 B ✓ mit 853 K from 19.7 GB
87 Rax-4.5 raxcore-dev · image-text-to-text 2.3 B ✓ apache-2.0 846 K from 5.8 GB
88 gemma-4-12b-it-GGUF unsloth · image-text-to-text 12.0 B ✓ apache-2.0 844 K from 7.9 GB
89 Qwen2-VL-7B-Instruct Qwen · image-text-to-text 8.3 B 33 K ✓ apache-2.0 841 K from 20.0 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.