kredor / token-classification updated 2 years ago

punctuate-all

This is based on Oliver Guhr's work. The difference is that it is a finetuned xlm-roberta-base instead of an xlm-roberta-large and on twelve languages instead of four. The languages are: English, German, French, Spanish, Bulgarian, Italian, Polish, Dutch, Czech, Portugese, Slovak, Slovenian.

Params
Context
512
Downloads 30d
708 K
Likes
28
Commercial use: allowed mit Not gated View on Hugging Face ↗

Download history

daily snapshots · 10 days
709 K692 K
Jul 28Jul 31Aug 3Aug 6

Specifications

Architecture
XLMRobertaForTokenClassification
Context length
512
Vocabulary
250,002
Layers / heads
12 / 12
Licence
mit
First seen on the Hub
2022-04-09
Training datasets
wmt/europarl
Added to our catalog
2026-07-28