microsoft / text-classification updated 4 years ago

deberta-xlarge-mnli

DeBERTa improves the BERT and RoBERTa models using disentangled attention and enhanced mask decoder. It outperforms BERT and RoBERTa on majority of NLU tasks with 80GB training data.

Params
Context
512
Downloads 30d
349 K
Likes
23
Commercial use: allowed mit Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 50 days
▲ 461 K in the last 30 days (56.9%)
818 K218 K
Aug 2Aug 18Sep 4Sep 20

Specifications

Architecture
DebertaForSequenceClassification
Context length
512
Vocabulary
50,265
Layers / heads
48 / 16
Licence
mit
First seen on the Hub
2022-03-02
Training datasets
undisclosed
Added to our catalog
2026-08-02
Compare with any text-classification model