indobenchmark / feature-extraction updated 5 years ago

indobert-base-p1

IndoBERT is a state-of-the-art language model for Indonesian based on the BERT model. The pretrained model is trained using a masked language modeling (MLM) objective and next sentence prediction (NSP) objective.

Params
Context
512
Downloads 30d
670 K
Likes
53
Commercial use: allowed mit Not gated 1 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
969 K670 K
Jul 28Jul 31Aug 3Aug 6

Specifications

Architecture
BertModel
Context length
512
Vocabulary
50,000
Layers / heads
12 / 12
Licence
mit
First seen on the Hub
2022-03-02
Training datasets
Indo4B
Added to our catalog
2026-07-28