Updated 2026-08-06 · ranked by real download data
Models that fit in 12 GB
Ranked nightly from our download snapshots of the Hugging Face catalog. Every entry shows its licence and what hardware it realistically needs.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 01 | all-MiniLM-L6-v2 | 20 M | 512 | ✓ apache-2.0 | 257.8 M | from 0.6 GB |
| 02 | bge-small-en-v1.5 | 30 M | 512 | ✓ mit | 72.4 M | from 0.7 GB |
| 03 | paraphrase-multilingual-MiniLM-L12-v2 | 120 M | 512 | ✓ apache-2.0 | 59.4 M | from 1.0 GB |
| 04 | Qwen3-0.6B | 750 M | 41 K | ✓ apache-2.0 | 29.7 M | from 2.3 GB |
| 05 | all-mpnet-base-v2 | 110 M | 512 | ✓ apache-2.0 | 25.2 M | from 1.0 GB |
| 06 | t5-small | 60 M | — | ✓ apache-2.0 | 24.7 M | from 0.8 GB |
| 07 | bge-reranker-v2-m3 | 570 M | 8 K | ✓ apache-2.0 | 19.2 M | from 3.1 GB |
| 08 | mobilenetv3_small_100.lamb_in1k | 0 M | — | ✓ apache-2.0 | 18.5 M | from 0.5 GB |
| 09 | nomic-embed-text-v1.5 | 140 M | 2 K | ✓ apache-2.0 | 15.7 M | from 1.1 GB |
| 10 | multilingual-e5-small | 120 M | 512 | ✓ mit | 15.6 M | from 1.0 GB |
| 11 | Qwen2.5-1.5B-Instruct | 1.5 B | 33 K | ✓ apache-2.0 | 14.0 M | from 4.1 GB |
| 12 | gpt2 | 140 M | — | ✓ mit | 13.7 M | from 1.1 GB |
| 13 | bge-large-en-v1.5 | 340 M | 512 | ✓ mit | 13.1 M | from 2.0 GB |
| 14 | paraphrase-multilingual-mpnet-base-v2 | 280 M | 512 | ✓ apache-2.0 | 11.5 M | from 1.8 GB |
| 15 | Llama-3.2-1B-Instruct | 1.2 B | — | ⚠ llama3.2 | 10.4 M | from 3.4 GB |
| 16 | bge-base-en-v1.5 | 110 M | 512 | ✓ mit | 10.0 M | from 1.0 GB |
| 17 | Qwen3-Embedding-0.6B | 600 M | 33 K | ✓ apache-2.0 | 9.8 M | from 1.9 GB |
| 18 | whisper-large-v3-turbo | 810 M | — | ✓ mit | 8.7 M | from 2.4 GB |
| 19 | Qwen2.5-VL-3B-Instruct | 3.8 B | 128 K | unknown | 8.2 M | from 9.3 GB |
| 20 | Qwen3-1.7B | 2.0 B | 41 K | ✓ apache-2.0 | 7.8 M | from 5.3 GB |
| 21 | multilingual-e5-large | 560 M | 512 | ✓ mit | 7.5 M | from 3.0 GB |
| 22 | multilingual-e5-base | 280 M | 512 | ✓ mit | 7.2 M | from 1.8 GB |
| 23 | clap-htsat-fused | 150 M | — | ✓ apache-2.0 | 6.9 M | from 1.2 GB |
| 24 | Qwen3.5-4B | 4.7 B | — | ✓ apache-2.0 | 6.7 M | from 11.5 GB |
| 25 | Qwen2.5-0.5B-Instruct | 490 M | 33 K | ✓ apache-2.0 | 6.6 M | from 1.7 GB |
Membership and ranking refresh nightly after the snapshot run. VRAM is an estimate for the smallest quantization at 8K context — methodology.