1,011 models · refreshed nightly

Models that run on RTX 4090 · 24 GB

Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.

#ModelParamsContextCommercial use30dOn RTX 4090 · 24 GB
01 all-MiniLM-L6-v2 sentence-transformers · sentence-similarity 20 M 512 ✓ apache-2.0 257.8 M 0.6 GB
02 bge-small-en-v1.5 BAAI · feature-extraction 30 M 512 ✓ mit 72.4 M 0.7 GB
03 paraphrase-multilingual-MiniLM-L12-v2 sentence-transformers · sentence-similarity 120 M 512 ✓ apache-2.0 59.4 M 1.0 GB
04 Qwen3-0.6B Qwen · text-generation 750 M 41 K ✓ apache-2.0 29.7 M 2.3 GB
05 all-mpnet-base-v2 sentence-transformers · sentence-similarity 110 M 512 ✓ apache-2.0 25.2 M 1.0 GB
06 t5-small google-t5 · translation 60 M ✓ apache-2.0 24.7 M 0.8 GB
07 bge-reranker-v2-m3 BAAI · text-classification 570 M 8 K ✓ apache-2.0 19.2 M 3.1 GB
08 mobilenetv3_small_100.lamb_in1k timm · image-classification 0 M ✓ apache-2.0 18.5 M 0.5 GB
09 Qwen3-8B Qwen · text-generation 8.2 B 41 K ✓ apache-2.0 16.3 M 19.7 GB
10 nomic-embed-text-v1.5 nomic-ai · sentence-similarity 140 M 2 K ✓ apache-2.0 15.7 M 1.1 GB
11 multilingual-e5-small intfloat · sentence-similarity 120 M 512 ✓ mit 15.6 M 1.0 GB
12 Qwen2.5-1.5B-Instruct Qwen · text-generation 1.5 B 33 K ✓ apache-2.0 14.0 M 4.1 GB
13 gpt2 openai-community · text-generation 140 M ✓ mit 13.7 M 1.1 GB
14 bge-large-en-v1.5 BAAI · feature-extraction 340 M 512 ✓ mit 13.1 M 2.0 GB
15 Qwen3.5-9B Qwen · image-text-to-text 9.7 B ✓ apache-2.0 12.4 M 23.2 GB
16 Qwen2.5-7B-Instruct Qwen · text-generation 7.6 B 33 K ✓ apache-2.0 12.1 M 18.4 GB
17 paraphrase-multilingual-mpnet-base-v2 sentence-transformers · sentence-similarity 280 M 512 ✓ apache-2.0 11.5 M 1.8 GB
18 Llama-3.2-1B-Instruct meta-llama · text-generation · gated 1.2 B ⚠ llama3.2 10.4 M 3.4 GB
19 bge-base-en-v1.5 BAAI · feature-extraction 110 M 512 ✓ mit 10.0 M 1.0 GB
20 Qwen3-Embedding-0.6B Qwen · feature-extraction 600 M 33 K ✓ apache-2.0 9.8 M 1.9 GB
21 Qwen2.5-VL-7B-Instruct Qwen · image-text-to-text 8.3 B 128 K ✓ apache-2.0 9.3 M 20.0 GB
22 whisper-large-v3-turbo openai · ASR 810 M ✓ mit 8.7 M 2.4 GB
23 Qwen2.5-VL-3B-Instruct Qwen · image-text-to-text 3.8 B 128 K unknown 8.2 M 9.3 GB
24 Llama-3.1-8B-Instruct meta-llama · text-generation · gated 8.0 B ⚠ llama3.1 8.0 M 19.4 GB
25 Qwen3-1.7B Qwen · text-generation 2.0 B 41 K ✓ apache-2.0 7.8 M 5.3 GB
26 multilingual-e5-large intfloat · feature-extraction 560 M 512 ✓ mit 7.5 M 3.0 GB
27 multilingual-e5-base intfloat · sentence-similarity 280 M 512 ✓ mit 7.2 M 1.8 GB
28 clap-htsat-fused laion · audio-classification 150 M ✓ apache-2.0 6.9 M 1.2 GB
29 Qwen3.5-4B Qwen · image-text-to-text 4.7 B ✓ apache-2.0 6.7 M 11.5 GB
30 Qwen2.5-0.5B-Instruct Qwen · text-generation 490 M 33 K ✓ apache-2.0 6.6 M 1.7 GB
31 Qwen2.5-3B-Instruct Qwen · text-generation 3.1 B 33 K other 6.0 M 7.8 GB
32 whisper-large-v3 openai · ASR 1.5 B ✓ apache-2.0 5.5 M 10.9 GB
33 nsfw_image_detection Falconsai · image-classification 90 M ✓ apache-2.0 5.4 M 0.9 GB
34 nomic-embed-text-v1 nomic-ai · sentence-similarity 140 M 8 K ✓ apache-2.0 5.1 M 1.1 GB
35 Ornith-1.0-9B-GGUF deepreinforce-ai · text-generation ✓ mit 4.9 M 6.7 GB
36 Ornith-1.0-9B-GGUF ornith-ai · text-generation ✓ mit 4.9 M 6.7 GB
37 mxbai-embed-large-v1 mixedbread-ai · feature-extraction 340 M 512 ✓ apache-2.0 4.9 M 1.3 GB
38 Qwen3-VL-8B-Instruct Qwen · image-text-to-text 8.8 B ✓ apache-2.0 4.9 M 21.1 GB
39 whisper-base openai · ASR 70 M ✓ apache-2.0 4.8 M 0.8 GB
40 bge-small-zh-v1.5 BAAI · feature-extraction 20 M 512 ✓ mit 4.8 M 0.6 GB
41 gemma-3-1b-it google · text-generation · gated 1.0 B ⚠ gemma 4.7 M 2.8 GB
42 Qwen3-Coder-30B-A3B-Instruct-GGUF unsloth · text-generation ✓ apache-2.0 4.7 M 12.9 GB
43 vit-base-patch16-224 google · image-classification 90 M ✓ apache-2.0 4.7 M 0.9 GB
44 Qwen2.5-7B-Instruct-AWQ Qwen · text-generation 7.6 B 33 K ✓ apache-2.0 4.6 M 7.8 GB
45 Qwen3-4B Qwen · text-generation 4.0 B 41 K ✓ apache-2.0 4.6 M 10.0 GB
46 bge-reranker-base BAAI · text-classification 280 M 512 ✓ mit 4.6 M 1.8 GB
47 Prompt-Guard-86M meta-llama · text-classification · gated 280 M ⚠ llama3.1 4.5 M 1.8 GB
48 Qwen3-ASR-0.6B Qwen · ASR 940 M ✓ apache-2.0 4.3 M 2.7 GB
49 granite-4.1-8b ibm-granite · text-generation 8.8 B 131 K ✓ apache-2.0 4.2 M 21.2 GB
50 Qwen3-VL-4B-Instruct Qwen · image-text-to-text 4.4 B ✓ apache-2.0 4.0 M 10.9 GB
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.