1,011 models · refreshed nightly
Text generation models
Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 151 | DeepSeek-R1-Distill-Qwen-1.5B | 1.8 B | 131 K | ✓ mit | 619 K | from 4.7 GB |
| 152 | OLMo-2-0425-1B | 1.5 B | 4 K | ✓ apache-2.0 | 618 K | from 7.3 GB |
| 153 | NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 | 31.6 B | 262 K | other | 613 K | from 41.2 GB |
| 154 | Qwen2-7B-Instruct | 7.6 B | 33 K | ✓ apache-2.0 | 606 K | from 18.4 GB |
| 155 | Qwen2.5-Coder-7B-Instruct-AWQ | 7.6 B | 33 K | ✓ apache-2.0 | 604 K | from 7.9 GB |
| 156 | VLM2Vec-Full | 4.2 B | 131 K | ✓ apache-2.0 | 599 K | from 10.2 GB |
| 157 | deepseek-coder-7b-instruct-v1.5 | 6.9 B | 4 K | other | 599 K | from 16.7 GB |
| 158 | Qwen3-30B-A3B-Instruct-2507-FP8 | 30.5 B | 262 K | ✓ apache-2.0 | 588 K | from 39.4 GB |
| 159 | LFM2.5-1.2B-Instruct | 1.2 B | 128 K | other | 584 K | from 3.3 GB |
| 160 | Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit | 5.3 B | 262 K | ✓ apache-2.0 | 583 K | from 21.2 GB |
| 161 | gemma-2-2b-it | 2.6 B | — | ⚠ gemma | 576 K | from 6.6 GB |
| 162 | gpt2-medium | 380 M | — | ✓ mit | 572 K | from 2.2 GB |
| 163 | gpt-oss-20b-GGUF | — | 131 K | ✓ apache-2.0 | 560 K | from 13.1 GB |
| 164 | Hermes-3-Llama-3.1-8B | 8.0 B | 131 K | ⚠ llama3 | 558 K | from 19.4 GB |
| 165 | DeepSeek-Coder-V2-Lite-Instruct | 15.7 B | 164 K | other | 557 K | from 37.4 GB |
| 166 | MiniMax-M3-NVFP4 | 246.6 B | — | other | 557 K | from 312.6 GB |
| 167 | Qwen3-Coder-Next | 79.7 B | 262 K | ✓ apache-2.0 | 554 K | from 187.7 GB |
| 168 | deepseek-coder-6.7b-instruct | 6.7 B | 16 K | other | 553 K | from 16.3 GB |
| 169 | Llama-3.2-3B | 3.2 B | — | ⚠ llama3.2 | 552 K | from 8.0 GB |
| 170 | MiMo-7B-RL | 7.8 B | 33 K | ✓ mit | 552 K | from 18.9 GB |
| 171 | Phi-3-mini-4k-instruct | 3.8 B | 4 K | ✓ mit | 544 K | from 9.5 GB |
| 172 | macbert4csc-base-chinese | 100 M | 512 | ✓ apache-2.0 | 542 K | from 1.0 GB |
| 173 | Qwen2.5-Coder-7B-Instruct-GPTQ-Int4 | 7.6 B | 33 K | ✓ apache-2.0 | 536 K | from 7.8 GB |
| 174 | Qwen-AgentWorld-35B-A3B-GGUF | — | — | ✓ apache-2.0 | 534 K | from 13.1 GB |
| 175 | Qwen2.5-Coder-7B-Instruct-AWQ | 7.6 B | 33 K | ✓ apache-2.0 | 522 K | from 7.8 GB |
| 176 | DeepSeek-V4-Flash-DSpark | 165.3 B | 1.0 M | ✓ mit | 519 K | from 208.9 GB |
| 177 | EXAONE-3.5-7.8B-Instruct | 7.8 B | 33 K | other | 518 K | from 36.1 GB |
| 178 | DeepSeek-R1-Distill-Llama-70B | 70.6 B | 131 K | ✓ mit | 512 K | from 166.3 GB |
| 179 | gemma-2-9b-it-AWQ-INT4 | 9.2 B | 8 K | ⚠ gemma | 504 K | from 8.7 GB |
| 180 | DeepSeek-V2-Lite | 15.7 B | 164 K | other | 503 K | from 37.4 GB |
| 181 | t5gemma-s-s-prefixlm | 310 M | — | ⚠ gemma | 499 K | from 1.2 GB |
| 182 | SmolLM2-360M-Instruct | 360 M | 8 K | ✓ apache-2.0 | 494 K | from 1.4 GB |
| 183 | Llama-3.1-405B-FP8 | 405.9 B | — | ⚠ llama3.1 | 493 K | from 597.3 GB |
| 184 | GLM-5.1-FP8 | 753.9 B | 203 K | ✓ mit | 493 K | from 945.4 GB |
| 185 | Qwen2.5-32B-Instruct-GPTQ-Int4 | 32.8 B | 33 K | ✓ apache-2.0 | 492 K | from 26.7 GB |
| 186 | tiny-mixtral | 250 M | 131 K | unknown | 490 K | from 1.6 GB |
| 187 | GLM-4.7-Flash-AWQ-4bit | 32.1 B | 203 K | ✓ mit | 489 K | from 27.5 GB |
| 188 | Meta-Llama-3.1-8B-Instruct-FP8 | 8.0 B | 131 K | ⚠ llama3.1 | 479 K | from 11.7 GB |
| 189 | NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 560.5 B | 262 K | other | 469 K | from 1,317.7 GB |
| 190 | typhoon2.5-qwen3-4b | 4.0 B | 262 K | ✓ apache-2.0 | 468 K | from 10.0 GB |
| 191 | Bielik-11B-v3.0-Instruct | 11.2 B | — | ✓ apache-2.0 | 464 K | from 26.7 GB |
| 192 | Llama-3.3-70B-Instruct | 70.6 B | — | ⚠ llama3.3 | 464 K | from 166.3 GB |
| 193 | Qwen3-4B-AWQ | 4.0 B | 41 K | ✓ apache-2.0 | 463 K | from 4.0 GB |
| 194 | Qwen3-30B-A3B-abliterated | 30.5 B | 41 K | ✓ apache-2.0 | 459 K | from 206.6 GB |
| 195 | Bonsai-27B-mlx-1bit | 1.7 B | — | ✓ apache-2.0 | 457 K | from 6.4 GB |
| 196 | Phi-4-mini-instruct | 3.8 B | 131 K | ✓ mit | 455 K | from 9.5 GB |
| 197 | Ternary-Bonsai-27B-mlx-2bit | 2.6 B | — | ✓ apache-2.0 | 453 K | from 10.2 GB |
| 198 | xlnet-base-cased | — | — | ✓ mit | 452 K | — |
| 199 | Qwen3-8B-Base | 8.2 B | 33 K | ✓ apache-2.0 | 451 K | from 19.7 GB |
| 200 | Qwen2.5-3B-Instruct-unsloth-bnb-4bit | 3.2 B | 33 K | ✓ apache-2.0 | 449 K | from 3.6 GB |
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.