1,011 models · refreshed nightly

Text generation models

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dMin VRAM
201 Qwen3Guard-Gen-4B Qwen · text-generation 4.4 B 33 K ✓ apache-2.0 440 K from 10.9 GB
202 Hy3-GGUF AngelSlim · text-generation ✓ apache-2.0 439 K from 101.4 GB
203 Qwen2.5-72B-Instruct Qwen · text-generation 72.7 B 33 K other 435 K from 171.4 GB
204 Nemotron-Labs-Diffusion-8B-Base nvidia · text-generation 8.5 B 4 K other 434 K from 20.5 GB
205 Olmo-3-7B-Instruct allenai · text-generation 7.3 B 66 K ✓ apache-2.0 434 K from 17.7 GB
206 DeepSeek-V4-Flash-0731 deepseek-ai · text-generation 304.2 B 1.0 M ✓ mit 433 K from 229.7 GB
207 GLM-4.5-Air zai-org · text-generation 110.5 B 131 K ✓ mit 433 K from 260.1 GB
208 Qwen2.5-14B-Instruct unsloth · text-generation 14.8 B 33 K ✓ apache-2.0 429 K from 35.2 GB
209 Zamba2-1.2B-instruct Zyphra · text-generation 1.2 B 4 K ✓ apache-2.0 429 K from 6.0 GB
210 LLaDA-8B-Instruct GSAI-ML · text-generation 8.0 B ✓ mit 429 K from 19.3 GB
211 japanese-gpt-neox-small rinna · text-generation 200 M 2 K ✓ mit 428 K from 1.3 GB
212 llama-3.3-70b-instruct-awq casperhansen · text-generation 70.6 B 131 K ⚠ llama3.3 426 K from 54.8 GB
213 saiga_llama3_8b IlyaGusev · text-generation 8.0 B 8 K other 423 K from 19.4 GB
214 gemma-2-9b-it google · text-generation · gated 9.2 B ⚠ gemma 417 K from 22.2 GB
215 Qwen3-14B-FP8 Qwen · text-generation 14.8 B 41 K ✓ apache-2.0 412 K from 20.7 GB
216 Meta-Llama-3.1-8B-Instruct unsloth · text-generation 8.0 B 131 K ⚠ llama3.1 410 K from 19.4 GB
217 Qwen2.5-Coder-14B-Instruct-AWQ Qwen · text-generation 14.8 B 33 K ✓ apache-2.0 406 K from 13.7 GB
218 Vikhr-Nemo-12B-Instruct-R-21-09-24 Vikhrmodels · text-generation 12.3 B 1.0 M ✓ apache-2.0 396 K from 30.8 GB
219 DeepSeek-R1-Distill-Llama-8B deepseek-ai · text-generation 8.0 B 131 K ✓ mit 395 K from 19.4 GB
220 EXAONE-3.5-7.8B-Instruct-AWQ LGAI-EXAONE · text-generation 7.8 B 33 K other 389 K from 7.5 GB
221 MiMo-V2.5 XiaomiMiMo · text-generation 310.8 B 1.0 M ✓ mit 389 K from 394.4 GB
222 Ornith-1.0-397B deepreinforce-ai · text-generation 396.8 B ✓ mit 387 K from 933.0 GB
223 Ornith-1.0-397B ornith-ai · text-generation 396.8 B ✓ mit 387 K from 933.0 GB
224 Laguna-S-2.1-NVFP4 poolside · text-generation 117.6 B 1.0 M openmdw-1.1 386 K from 127.8 GB
225 Ilama-3.2-1B hmellor · text-generation 1.2 B 131 K unknown 381 K from 6.1 GB
226 mistral-7b-v0.3-bnb-4bit unsloth · text-generation 7.5 B 33 K ✓ apache-2.0 380 K from 6.2 GB
227 GLM-5.2-AWQ-INT4 cyankiwi · text-generation 753.3 B 1.0 M ✓ mit 378 K from 635.1 GB
228 gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF yuxinlu1 · text-generation ✓ apache-2.0 373 K from 1.4 GB
229 Fanar-1-9B-Instruct QCRI · text-generation 8.8 B 4 K ✓ apache-2.0 364 K from 21.1 GB
230 NVIDIA-Nemotron-Nano-12B-v2 nvidia · text-generation 12.3 B other 364 K from 29.4 GB
231 Apertus-8B-Instruct-2509 swiss-ai · text-generation 8.1 B 66 K ✓ apache-2.0 361 K from 19.4 GB
232 Qwen3.5-397B-A17B-NVFP4 nvidia · text-generation ✓ apache-2.0 359 K from 276.8 GB
233 CodeLlama-7b-hf codellama · text-generation 6.7 B 16 K ⚠ llama2 356 K from 16.3 GB
234 gpt-oss-20b-MXFP4-Q8 mlx-community · text-generation 20.9 B 131 K ✓ apache-2.0 355 K from 16.9 GB
235 GLM-4.5-Air unsloth · text-generation 110.5 B 131 K ✓ mit 352 K from 260.1 GB
236 gpt-neo-125m EleutherAI · text-generation 150 M 2 K ✓ mit 350 K from 1.1 GB
237 gemma-3-270m-it google · text-generation · gated 270 M ⚠ gemma 346 K from 1.1 GB
238 Kimi-K2.6-NVFP4 nvidia · text-generation other 346 K from 655.2 GB
239 Qwen3-4B-Thinking-2507 Qwen · text-generation 4.0 B 262 K ✓ apache-2.0 341 K from 10.0 GB
240 MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF GnLOLot · text-generation ✓ apache-2.0 340 K from 1.3 GB
241 Qwen3-Next-80B-A3B-Instruct-FP8 Qwen · text-generation 81.3 B 262 K ✓ apache-2.0 339 K from 103.0 GB
242 Apertus-70B-Instruct-2509-quantized.w4a16 RedHatAI · text-generation 11.3 B 66 K ✓ apache-2.0 338 K from 46.0 GB
243 Llama-2-7b-hf NousResearch · text-generation 6.7 B 4 K unknown 335 K from 16.3 GB
244 granite-4.1-3b ibm-granite · text-generation 3.4 B 131 K ✓ apache-2.0 329 K from 8.5 GB
245 Qwen3-Coder-480B-A35B-Instruct-FP8 Qwen · text-generation 480.2 B 262 K ✓ apache-2.0 327 K from 602.9 GB
246 DeepSeek-R1-0528-NVFP4-v2 nvidia · text-generation 393.6 B 164 K ✓ mit 325 K from 514.2 GB
247 Mistral-7B-Instruct-v0.2-AWQ TheBloke · text-generation 7.2 B 33 K ✓ apache-2.0 323 K from 6.2 GB
248 QwQ-32B Qwen · text-generation 32.8 B 41 K ✓ apache-2.0 318 K from 77.5 GB
249 Jan-v3.5-4B-gguf janhq · text-generation ✓ apache-2.0 316 K from 2.8 GB
250 NVIDIA-Nemotron-Nano-9B-v2 nvidia · text-generation 8.9 B 131 K other 313 K from 21.4 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.