Updated 2026-08-06 · ranked by real download data

Models that fit in 12 GB

Ranked nightly from our download snapshots of the Hugging Face catalog. Every entry shows its licence and what hardware it realistically needs.

#ModelParamsContextCommercial use30dMin VRAM
01 all-MiniLM-L6-v2 sentence-transformers · sentence-similarity 20 M 512 ✓ apache-2.0 257.8 M from 0.6 GB
02 bge-small-en-v1.5 BAAI · feature-extraction 30 M 512 ✓ mit 72.4 M from 0.7 GB
03 paraphrase-multilingual-MiniLM-L12-v2 sentence-transformers · sentence-similarity 120 M 512 ✓ apache-2.0 59.4 M from 1.0 GB
04 Qwen3-0.6B Qwen · text-generation 750 M 41 K ✓ apache-2.0 29.7 M from 2.3 GB
05 all-mpnet-base-v2 sentence-transformers · sentence-similarity 110 M 512 ✓ apache-2.0 25.2 M from 1.0 GB
06 t5-small google-t5 · translation 60 M ✓ apache-2.0 24.7 M from 0.8 GB
07 bge-reranker-v2-m3 BAAI · text-classification 570 M 8 K ✓ apache-2.0 19.2 M from 3.1 GB
08 mobilenetv3_small_100.lamb_in1k timm · image-classification 0 M ✓ apache-2.0 18.5 M from 0.5 GB
09 nomic-embed-text-v1.5 nomic-ai · sentence-similarity 140 M 2 K ✓ apache-2.0 15.7 M from 1.1 GB
10 multilingual-e5-small intfloat · sentence-similarity 120 M 512 ✓ mit 15.6 M from 1.0 GB
11 Qwen2.5-1.5B-Instruct Qwen · text-generation 1.5 B 33 K ✓ apache-2.0 14.0 M from 4.1 GB
12 gpt2 openai-community · text-generation 140 M ✓ mit 13.7 M from 1.1 GB
13 bge-large-en-v1.5 BAAI · feature-extraction 340 M 512 ✓ mit 13.1 M from 2.0 GB
14 paraphrase-multilingual-mpnet-base-v2 sentence-transformers · sentence-similarity 280 M 512 ✓ apache-2.0 11.5 M from 1.8 GB
15 Llama-3.2-1B-Instruct meta-llama · text-generation 1.2 B ⚠ llama3.2 10.4 M from 3.4 GB
16 bge-base-en-v1.5 BAAI · feature-extraction 110 M 512 ✓ mit 10.0 M from 1.0 GB
17 Qwen3-Embedding-0.6B Qwen · feature-extraction 600 M 33 K ✓ apache-2.0 9.8 M from 1.9 GB
18 whisper-large-v3-turbo openai · ASR 810 M ✓ mit 8.7 M from 2.4 GB
19 Qwen2.5-VL-3B-Instruct Qwen · image-text-to-text 3.8 B 128 K unknown 8.2 M from 9.3 GB
20 Qwen3-1.7B Qwen · text-generation 2.0 B 41 K ✓ apache-2.0 7.8 M from 5.3 GB
21 multilingual-e5-large intfloat · feature-extraction 560 M 512 ✓ mit 7.5 M from 3.0 GB
22 multilingual-e5-base intfloat · sentence-similarity 280 M 512 ✓ mit 7.2 M from 1.8 GB
23 clap-htsat-fused laion · audio-classification 150 M ✓ apache-2.0 6.9 M from 1.2 GB
24 Qwen3.5-4B Qwen · image-text-to-text 4.7 B ✓ apache-2.0 6.7 M from 11.5 GB
25 Qwen2.5-0.5B-Instruct Qwen · text-generation 490 M 33 K ✓ apache-2.0 6.6 M from 1.7 GB
Membership and ranking refresh nightly after the snapshot run. VRAM is an estimate for the smallest quantization at 8K context — methodology.