1,011 models · refreshed nightly

Image text to text models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
51 Qwen3.5-122B-A10B-FP8 Qwen · image-text-to-text 125.1 B ✓ apache-2.0 1.3 M from 159.1 GB
52 Qwen3.6-35B-A3B-AWQ-4bit cyankiwi · image-text-to-text 36.0 B ✓ apache-2.0 1.3 M from 33.4 GB
53 deepseek-vl2-tiny deepseek-ai · image-text-to-text 3.4 B other 1.3 M from 8.4 GB
54 gemma-3-27b-it-int4-awq gaunernst · image-text-to-text 27.4 B ⚠ gemma 1.2 M from 24.9 GB
55 gemma-4-31B-it-NVFP4 RedHatAI · image-text-to-text 32.7 B ✓ apache-2.0 1.2 M from 31.0 GB
56 gemma-4-26B-A4B-it-FP8-dynamic RedHatAI · image-text-to-text 26.6 B ✓ apache-2.0 1.2 M from 36.0 GB
57 Qwen3.6-27B-MTP-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.2 M from 14.3 GB
58 Qwen3.5-4B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.2 M from 2.8 GB
59 SmolVLM2-500M-Video-Instruct HuggingFaceTB · image-text-to-text 510 M ✓ apache-2.0 1.2 M from 2.8 GB
60 surya-ocr-2 datalab-to · image-text-to-text 690 M ⚠ openrail 1.2 M from 2.1 GB
61 Qwen3.6-35B-A3B-AWQ QuantTrio · image-text-to-text 36.0 B ✓ apache-2.0 1.1 M from 33.9 GB
62 Kimi-K3 moonshotai · image-text-to-text 2,779.9 B other 1.1 M from 2,134.5 GB
63 Qwen2.5-VL-32B-Instruct Qwen · image-text-to-text 33.5 B 128 K ✓ apache-2.0 1.1 M from 80.6 GB
64 SmolVLM-256M-Instruct HuggingFaceTB · image-text-to-text 260 M ✓ apache-2.0 1.1 M from 1.1 GB
65 llava-onevision-qwen2-0.5b-ov-hf llava-hf · image-text-to-text 890 M ✓ apache-2.0 1.1 M from 2.6 GB
66 Qwen3.5-9B-GGUF unsloth · image-text-to-text ✓ apache-2.0 1.0 M from 5.2 GB
67 Qwen3.6-27B-MLX-8bit lmstudio-community · image-text-to-text 8.0 B ✓ apache-2.0 983 K from 34.2 GB
68 Kimi-K2.5 moonshotai · image-text-to-text 1,058.6 B other 966 K from 814.0 GB
69 Qwen3.6-27B-MLX-4bit lmstudio-community · image-text-to-text 4.7 B ✓ apache-2.0 942 K from 18.9 GB
70 MiniCPM-V-4.6 openbmb · image-text-to-text 1.3 B ✓ apache-2.0 942 K from 3.6 GB
71 Gemma-4-E4B-Uncensored-HauhauCS-Aggressive HauhauCS · image-text-to-text ⚠ gemma 937 K from 5.4 GB
72 Qwen3.5-35B-A3B-FP8 Qwen · image-text-to-text 36.0 B ✓ apache-2.0 924 K from 47.1 GB
73 LocateAnything-3B nvidia · image-text-to-text 3.8 B other 923 K from 9.5 GB
74 Qwythos-9B-Claude-Mythos-5-1M-GGUF empero-ai · image-text-to-text ✓ apache-2.0 922 K from 7.0 GB
75 InternVL2-1B OpenGVLab · image-text-to-text 940 M ✓ mit 920 K from 2.7 GB
76 gemma-4-26B-A4B-it-QAT-MLX-4bit lmstudio-community · image-text-to-text 4.6 B ✓ apache-2.0 914 K from 18.4 GB
77 Qwen3.6-27B-GPTQ-Pro-4bit groxaxo · image-text-to-text 27.4 B ✓ apache-2.0 909 K from 25.2 GB
78 UI-TARS-1.5-7B ByteDance-Seed · image-text-to-text 8.3 B 128 K ✓ apache-2.0 909 K from 38.2 GB
79 Qwen3.6-27B-MLX-6bit lmstudio-community · image-text-to-text 6.4 B ✓ apache-2.0 907 K from 26.5 GB
80 gemma-4-26B-A4B-it-NVFP4 RedHatAI · image-text-to-text 15.1 B ✓ apache-2.0 897 K from 20.8 GB
81 Qwen3.6-35B-A3B-GGUF unsloth · image-text-to-text ✓ apache-2.0 895 K from 11.6 GB
82 Qwen3.6-27B-MLX-5bit lmstudio-community · image-text-to-text 5.5 B ✓ apache-2.0 892 K from 22.7 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.