nguyenvulebinh / automatic-speech-recognition updated 3 years ago

wav2vec2-base-vi-vlsp2020

Our models use wav2vec2 architecture, pre-trained on 13k hours of Vietnamese youtube audio (un-label data) and fine-tuned on 250 hours labeled of VLSP ASR dataset on 16kHz sampled speech audio. You can find more description here

Params
Context
Downloads 30d
665 K
Likes
2
Commercial use: not allowed · cc-by-nc-4.0 Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
799 K500 K
Jul 28Jul 31Aug 3Aug 6

Specifications

Architecture
Wav2Vec2ForCTC
Vocabulary
98
Layers / heads
12 / 12
Licence
cc-by-nc-4.0
First seen on the Hub
2022-11-04
Training datasets
vlsp-asr-2020
Added to our catalog
2026-07-28