jinaai / feature-extraction updated 3 months ago

jina-embeddings-v3

jina-embeddings-v3 is a multilingual multi-task text embedding model designed for a variety of NLP applications. Based on the Jina-XLM-RoBERTa architecture, this model supports Rotary Position Embeddings to handle long input sequences up to 8192 tokens. Additionally, it features 5 LoRA adapters to generate task-specific embeddings efficiently.

Params
570 M
Context
8,192
Downloads 30d
3.2 M
Likes
1,152
Commercial use: not allowed · cc-by-nc-4.0 Not gated SAFETENSORS 94 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
3.2 M3.2 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 1.1 GB 1.8 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
XLMRobertaModel
Parameters
570 M
Tensor type
BF16
Context length
8,192
Vocabulary
250,002
Layers / heads
24 / 16
Licence
cc-by-nc-4.0
First seen on the Hub
2024-09-05
Training datasets
undisclosed
MTEB AFQMC (default) (reported)
43.472678264757
MTEB NQ-PL (default) (reported)
62.543
MTEB FiQA-PL (default) (reported)
37.11
MTEB Quora-PL (default) (reported)
88.23
MTEB ArguAna-PL (default) (reported)
62.731
MTEB DBPedia-PL (default) (reported)
14.939
MTEB MSMARCO-PL (default) (reported)
9.623
MTEB SCIDOCS-PL (default) (reported)
11.198
MTEB SciFact-PL (default) (reported)
71.35
MTEB HotpotQA-PL (default) (reported)
57.522
MTEB NFCorpus-PL (default) (reported)
11.477
MTEB TRECCOVID-PL (default) (reported)
72.189
Added to our catalog
2026-07-28