1,054 models · refreshed nightly
Audio classification models
Every model in the catalog with its licence, estimated VRAM and daily-tracked downloads. Filters update the URL — share any view.
| # | Model | Params | Context | Commercial use | 30d | Min VRAM |
|---|---|---|---|---|---|---|
| 01 | clap-htsat-fused | 150 M | — | ✓ apache-2.0 | 8.2 M | from 1.2 GB |
| 02 | wav2vec2-large-robust-24-ft-age-gender | 320 M | — | ✗ cc-by-nc-sa-4.0 | 2.5 M | from 1.9 GB |
| 03 | ced-gguf | — | — | ✓ apache-2.0 | 922 K | from 0.6 GB |
| 04 | wav2vec2-large-robust-12-ft-emotion-msp-dim | 170 M | — | ✗ cc-by-nc-sa-4.0 | 744 K | from 1.3 GB |
| 05 | wav2vec2-large-xlsr-53-gender-recognition-librispeech | 320 M | — | ✓ apache-2.0 | 714 K | from 1.9 GB |
| 06 | ast-finetuned-audioset-10-10-0.4593 | 90 M | — | ✓ bsd-3-clause | 705 K | from 0.9 GB |
| 07 | audiobox-aesthetics | 100 M | — | ✓ cc-by-4.0 | 641 K | from 1.0 GB |
| 08 | open-vakgyata | 60 M | — | ✗ cc-by-nc-4.0 | 504 K | from 0.8 GB |
| 09 | MuQ-large-msd-iter | 330 M | — | ✗ cc-by-nc-4.0 | 295 K | from 2.0 GB |
| 10 | voice-gender-classifier | 20 M | — | ✓ mit | 291 K | from 0.6 GB |
| 11 | hubert-large-speech-emotion-recognition-russian-dusha-finetuned | 320 M | — | ✓ apache-2.0 | 219 K | from 1.9 GB |
VRAM figures are estimates for the smallest available quantization at 8K context — see /methodology. Downloads refresh nightly from the Hugging Face API.