s-nlp / text-classification updated 1 year ago

roberta_toxicity_classifier

This model is trained for toxicity classification task. The dataset used for training is the merge of the English parts of the three datasets by Jigsaw (Jigsaw 2018, Jigsaw 2019, Jigsaw 2020), containing around 2 million examples. We split it into two parts and fine-tune a RoBERTa model (RoBERTa: A Robustly Optimized BERT Pretraining Approach) on it. The classifiers perform closely on the test set of the first Jigsaw...

Params
Context
512
Downloads 30d
259 K
Likes
75
Commercial use: conditional · openrail++ Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 55 days
▲ 38 K in the last 30 days (17.0%)
278 K222 K
Jul 28Aug 15Sep 2Sep 20

Specifications

Architecture
RobertaForSequenceClassification
Context length
512
Vocabulary
50,265
Layers / heads
12 / 12
Licence
openrail++
First seen on the Hub
2022-03-02
Training datasets
google/jigsaw_toxicity_pred
Added to our catalog
2026-07-28
Compare with any text-classification model