1,011 models · refreshed nightly
Audio classification models
Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 01 | clap-htsat-fused | 150 M | — | ✓ apache-2.0 | 6.9 M | from 1.2 GB |
| 02 | ast-finetuned-audioset-10-10-0.4593 | 90 M | — | ✓ bsd-3-clause | 1.2 M | from 0.9 GB |
| 03 | wav2vec2-large-robust-12-ft-emotion-msp-dim | 170 M | — | ✗ cc-by-nc-sa-4.0 | 751 K | from 1.3 GB |
| 04 | wav2vec2-large-xlsr-53-gender-recognition-librispeech | 320 M | — | ✓ apache-2.0 | 626 K | from 1.9 GB |
| 05 | MERT-v1-330M | — | — | ✗ cc-by-nc-4.0 | 560 K | — |
| 06 | audiobox-aesthetics | 100 M | — | ✓ cc-by-4.0 | 401 K | from 1.0 GB |
| 07 | voice-gender-classifier | 20 M | — | ✓ mit | 308 K | from 0.6 GB |
| 08 | mms-lid-126 | 970 M | — | ✗ cc-by-nc-4.0 | 308 K | from 4.9 GB |
| 09 | MuQ-large-msd-iter | 330 M | — | ✗ cc-by-nc-4.0 | 292 K | from 2.0 GB |
| 10 | open-vakgyata | 60 M | — | ✗ cc-by-nc-4.0 | 278 K | from 0.8 GB |
| 11 | wav2vec2-large-robust-24-ft-age-gender | 320 M | — | ✗ cc-by-nc-sa-4.0 | 255 K | from 1.9 GB |
| 12 | hubert-large-speech-emotion-recognition-russian-dusha-finetuned | 320 M | — | ✓ apache-2.0 | 248 K | from 1.9 GB |
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.