1,011 models · refreshed nightly
Text generation models
Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 51 | Qwen3-Coder-Next-FP8 | 79.7 B | 262 K | ✓ apache-2.0 | 2.2 M | from 100.9 GB |
| 52 | Qwen-72B | 72.3 B | 33 K | other | 2.2 M | from 170.4 GB |
| 53 | gemma-3-270m | 270 M | — | ⚠ gemma | 2.1 M | from 1.1 GB |
| 54 | GLM-4.7-Flash | 31.2 B | 203 K | ✓ mit | 2.1 M | from 73.9 GB |
| 55 | Meta-Llama-3-8B | 8.0 B | — | ⚠ llama3 | 2.1 M | from 19.4 GB |
| 56 | Qwen2.5-Coder-7B-Instruct | 7.6 B | 33 K | ✓ apache-2.0 | 2.0 M | from 18.4 GB |
| 57 | SmolLM2-135M-Instruct | 130 M | 8 K | ✓ apache-2.0 | 2.0 M | from 0.8 GB |
| 58 | Qwen3-30B-A3B-Instruct-2507 | 30.5 B | 262 K | ✓ apache-2.0 | 1.9 M | from 72.3 GB |
| 59 | Qwen3.6-27B-NVFP4 | 18.2 B | — | ✓ apache-2.0 | 1.9 M | from 27.3 GB |
| 60 | PowerMoE-3b | 3.4 B | 4 K | ✓ apache-2.0 | 1.8 M | from 15.9 GB |
| 61 | SmolLM2-135M | 130 M | 8 K | ✓ apache-2.0 | 1.8 M | from 0.8 GB |
| 62 | Meta-Llama-3-8B-Instruct | 8.0 B | — | ⚠ llama3 | 1.8 M | from 19.4 GB |
| 63 | GLM-5.2-NVFP4 | 381.0 B | 1.0 M | ✓ mit | 1.8 M | from 569.0 GB |
| 64 | Llama-3.2-1B | 1.2 B | — | ⚠ llama3.2 | 1.8 M | from 3.4 GB |
| 65 | Qwen2.5-32B-Instruct | 32.8 B | 33 K | ✓ apache-2.0 | 1.7 M | from 77.5 GB |
| 66 | Qwen2.5-Coder-32B-Instruct-AWQ | 32.8 B | 33 K | ✓ apache-2.0 | 1.7 M | from 26.7 GB |
| 67 | diffusiongemma-26B-A4B-it-NVFP4 | 14.4 B | — | ✓ apache-2.0 | 1.7 M | from 23.4 GB |
| 68 | DeepSeek-R1-0528-Qwen3-8B | 8.2 B | 131 K | ✓ mit | 1.7 M | from 19.7 GB |
| 69 | DeepSeek-V4-Pro | 1,598.8 B | 1.0 M | ✓ mit | 1.6 M | from 1,191.5 GB |
| 70 | Qwen3-0.6B-FP8 | 750 M | 41 K | ✓ apache-2.0 | 1.6 M | from 1.8 GB |
| 71 | Llama-3.1-8B | 8.0 B | — | ⚠ llama3.1 | 1.5 M | from 19.4 GB |
| 72 | OpenELM-1_1B-Instruct | 1.1 B | — | apple-amlr | 1.5 M | from 3.0 GB |
| 73 | SmolLM-1.7B-Instruct-quantized.w4a16 | 1.8 B | 2 K | ✓ apache-2.0 | 1.5 M | from 2.7 GB |
| 74 | Llama-3.2-1B-Instruct-FP8-dynamic | 1.5 B | 131 K | ⚠ llama3.2 | 1.5 M | from 3.0 GB |
| 75 | gpt2-large | 810 M | — | ✓ mit | 1.5 M | from 4.2 GB |
| 76 | Gemma-4-26B-A4B-NVFP4 | 14.4 B | — | ✓ apache-2.0 | 1.5 M | from 23.3 GB |
| 77 | Qwen3-Coder-30B-A3B-Instruct | 30.5 B | 262 K | ✓ apache-2.0 | 1.5 M | from 72.3 GB |
| 78 | Llama-3.2-3B-Instruct | 3.2 B | — | ⚠ llama3.2 | 1.5 M | from 8.0 GB |
| 79 | Qwen3-Coder-30B-A3B-Instruct-FP8 | 30.5 B | 262 K | ✓ apache-2.0 | 1.4 M | from 39.4 GB |
| 80 | OTel-LLM-8B-A1B-IT | — | 128 K | ✓ apache-2.0 | 1.4 M | — |
| 81 | DeepSeek-V2-Lite-Chat | 15.7 B | 164 K | other | 1.4 M | from 37.4 GB |
| 82 | Mistral-7B-Instruct-v0.2 | 7.2 B | 33 K | ✓ apache-2.0 | 1.3 M | from 17.5 GB |
| 83 | Qwen2.5-72B-Instruct-AWQ | 73.0 B | 33 K | other | 1.3 M | from 57.2 GB |
| 84 | h2ovl-mississippi-800m | 830 M | — | ✓ apache-2.0 | 1.3 M | from 2.4 GB |
| 85 | h2ovl-mississippi-2b | 2.2 B | — | ✓ apache-2.0 | 1.3 M | from 5.6 GB |
| 86 | DeepSeek-V3.2 | 685.4 B | 164 K | ✓ mit | 1.2 M | from 861.7 GB |
| 87 | DeepSeek-V3 | 684.5 B | 164 K | unknown | 1.2 M | from 860.6 GB |
| 88 | Qwen2.5-Coder-32B-Instruct | 32.8 B | 33 K | ✓ apache-2.0 | 1.2 M | from 77.5 GB |
| 89 | falcon-7b | 7.2 B | — | ✓ apache-2.0 | 1.2 M | from 17.5 GB |
| 90 | bloomz-560m | 560 M | — | ⚠ bigscience-bloom-rail-1.0 | 1.2 M | from 1.8 GB |
| 91 | tiny-gpt2 | — | — | unknown | 1.2 M | — |
| 92 | Qwen3-4B-Instruct-2507-FP8 | 4.4 B | 262 K | ✓ apache-2.0 | 1.2 M | from 6.9 GB |
| 93 | Qwen3-VL-30B-A3B-Instruct-AWQ | 31.1 B | — | ✓ apache-2.0 | 1.2 M | from 24.8 GB |
| 94 | GLM-5-FP8 | 753.9 B | 203 K | ✓ mit | 1.2 M | from 945.4 GB |
| 95 | Qwen3-8B-AWQ | 8.2 B | 41 K | ✓ apache-2.0 | 1.1 M | from 8.4 GB |
| 96 | DeepSeek-V4-Flash-NVFP4 | 166.7 B | 1.0 M | ✓ mit | 1.1 M | from 210.6 GB |
| 97 | Phi-3.5-mini-instruct | 3.8 B | 131 K | ✓ mit | 1.1 M | from 9.5 GB |
| 98 | TinyLlama-1.1B-Chat-v0.3-GPTQ | 1.1 B | 2 K | ✓ apache-2.0 | 1.1 M | from 1.5 GB |
| 99 | Qwen2.5-Coder-7B | 7.6 B | 33 K | ✓ apache-2.0 | 1.0 M | from 18.4 GB |
| 100 | Qwen2-0.5B | 490 M | 131 K | ✓ apache-2.0 | 998 K | from 1.7 GB |
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.