chandra-ocr-2
Chandra 2 is a state of the art OCR model from Datalab that outputs markdown, HTML, and JSON. It is highly accurate at extracting text from images and PDFs, while preserving layout information.
Params
5.3 B
Context
—
Downloads 30d
2.7 M
Likes
515
Download history
daily snapshots · 55 days
▲ 89 K in the last 30 days (3.4%)
2.9 M2.6 M
Aug 22Sep 1Sep 11Sep 20
3.2 M2.6 M
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | bf16 | 10.6 GB | 12.9 GB | ✅ Runs comfortably |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Architecture
- Qwen3_5ForConditionalGeneration
- Parameters
- 5.3 B
- Tensor type
- BF16
- Licence
- openrail
- First seen on the Hub
- 2026-03-16
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Compare with any image-text-to-text model