Updated 2026-08-06 · ranked by real download data
Trending text-generation
Ranked nightly from our download snapshots of the Hugging Face catalog. Every entry shows its licence and what hardware it realistically needs.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 01 | Qwen3-Coder-30B-A3B-Instruct-GGUF | — | — | ✓ apache-2.0 | 4.7 M | — |
| 02 | Laguna-S-2.1-NVFP4 | 117.6 B | 1.0 M | openmdw-1.1 | 386 K | from 276.8 GB |
| 03 | Qwen3Guard-Gen-4B | 4.4 B | 33 K | ✓ apache-2.0 | 440 K | from 10.9 GB |
| 04 | GLM-5.2 | 753.3 B | 1.0 M | ✓ mit | 2.2 M | from 1,770.8 GB |
| 05 | Kimi-K2.7-Code-NVFP4 | — | — | other | 836 K | — |
| 06 | gemma-2-9b-it-AWQ-INT4 | 9.2 B | 8 K | ⚠ gemma | 504 K | from 22.2 GB |
| 07 | llama-3.3-70b-instruct-awq | 70.6 B | 131 K | ⚠ llama3.3 | 426 K | from 166.3 GB |
| 08 | Qwen-72B | 72.3 B | 33 K | other | 2.2 M | from 170.4 GB |
| 09 | Olmo-3-7B-Instruct | 7.3 B | 66 K | ✓ apache-2.0 | 434 K | from 17.7 GB |
| 10 | Apertus-70B-Instruct-2509-quantized.w4a16 | 11.3 B | 66 K | ✓ apache-2.0 | 338 K | from 27.1 GB |
| 11 | CodeLlama-7b-hf | 6.7 B | 16 K | ⚠ llama2 | 356 K | from 16.3 GB |
| 12 | MiniCPM5-1B | 1.1 B | 131 K | ✓ apache-2.0 | 926 K | from 3.0 GB |
| 13 | Qwen3-Coder-Next-FP8-dynamic | 79.8 B | 262 K | ✓ apache-2.0 | 864 K | from 187.9 GB |
| 14 | bloomz-560m | 560 M | — | ⚠ bigscience-bloom-rail-1.0 | 1.2 M | from 1.8 GB |
| 15 | Qwen3.6-35B-A3B-abliterated-v4 | 34.7 B | 262 K | ✓ apache-2.0 | 916 K | from 82.0 GB |
| 16 | DeepSeek-R1-Distill-Qwen-14B | 14.8 B | 131 K | ✓ mit | 766 K | from 35.2 GB |
| 17 | Qwen2.5-Coder-14B-Instruct-GPTQ-Int4 | 14.8 B | 33 K | ✓ apache-2.0 | 237 K | from 35.2 GB |
| 18 | gemma-3-1b-it | 1.0 B | — | ⚠ gemma | 4.7 M | from 2.9 GB |
| 19 | Mistral-7B-Instruct-v0.2-AWQ | 7.2 B | 33 K | ✓ apache-2.0 | 323 K | from 17.5 GB |
| 20 | NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | 560.5 B | 262 K | other | 469 K | from 1,317.7 GB |
| 21 | Qwen3-8B-AWQ | 8.2 B | 41 K | ✓ apache-2.0 | 1.1 M | from 19.7 GB |
| 22 | Qwen3.6-27B-Text-NVFP4-MTP | 16.7 B | — | ✓ apache-2.0 | 747 K | from 39.7 GB |
| 23 | GLM-5.2-AWQ-INT4 | 753.3 B | 1.0 M | ✓ mit | 378 K | from 1,770.8 GB |
| 24 | deepseek-v4-gguf | — | — | ✓ mit | 875 K | — |
| 25 | t5gemma-s-s-prefixlm | 310 M | — | ⚠ gemma | 499 K | from 1.2 GB |
Membership and ranking refresh nightly after the snapshot run. VRAM is an estimate for the smallest quantization at 8K context — methodology.