datalab-to / image-text-to-text updated 3 months ago

surya-ocr-2

- Accuracy - scores 83.3% on olmOCR-bench (top under 3B params) - Speed - throughput of 5 pages/s on an RTX 5090 - Multilingual - scores 87.2% on an internal benchmark set of 91 languages (more here) - Layout analysis (table, image, header, etc.) with reading order - Table recognition (rows + columns)

Params
690 M
Context
Downloads 30d
1.2 M
Likes
111
Commercial use: conditional · openrail Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 55 days
▲ 45 K in the last 30 days (3.5%)
1.5 M873 K
Jul 28Aug 15Sep 2Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 1.4 GB 2.1 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Qwen3_5ForConditionalGeneration
Parameters
690 M
Tensor type
BF16
Licence
openrail
First seen on the Hub
2026-05-14
Training datasets
undisclosed
Added to our catalog
2026-07-28
Compare with any image-text-to-text model