nvidia / automatic-speech-recognition updated 16 hours ago

parakeet-ctc-1.1b

parakeet-ctc-1.1b is an ASR model that transcribes speech in lower case English alphabet. This model is jointly developed by NVIDIA NeMo and Suno.ai teams. It is an XXL version of FastConformer CTC [1] (around 1.1B parameters) model. See the model architecture section and NeMo documentation for complete architecture details.

Params
1.1 B
Context
Downloads 30d
1.9 M
Likes
52
Commercial use: allowed cc-by-4.0 Not gated GGUF · SAFETENSORS 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
1.9 M1.5 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
parakeet-ctc-1.1b.q8_0.gguf Q8_0 1.2 GB 2.0 GB ✅ Runs comfortably
model.safetensors f32 4.3 GB 5.3 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
ParakeetForCTC
Parameters
1.1 B
Tensor type
F32
Vocabulary
1,025
Licence
cc-by-4.0
First seen on the Hub
2023-12-28
Training datasets
librispeech_asr, fisher_corpus, Switchboard-1, WSJ-0, WSJ-1, National-Singapore-Corpus-Part-1, National-Singapore-Corpus-Part-6, vctk, voxpopuli, europarl, multilingual_librispeech, mozilla-foundation/common_voice_8_0, MLCommons/peoples_speech
GigaSpeech (reported)
10.27
Vox Populi (reported)
6.53
tedlium-v3 (reported)
3.54
Earnings-22 (reported)
13.69
SPGI Speech (reported)
4.2
AMI (Meetings test) (reported)
15.62
LibriSpeech (clean) (reported)
1.83
LibriSpeech (other) (reported)
3.54
Mozilla Common Voice 9.0 (reported)
9.02
Added to our catalog
2026-07-28