nvidia / image-text-to-text updated 8 months ago

Llama-3.1-Nemotron-Nano-VL-8B-V1

Llama Nemotron Nano VL is a leading document intelligence vision language model (VLMs) that enables the ability to query and summarize images from the physical or virtual world. Llama Nemotron Nano VL is deployable in the data center, cloud and at the edge, including Jetson Orin and laptop by AWQ 4bit quantization through TinyChat framework. We find: (1) image-text pairs are not enough, interleaved image-text is esse...

Params
8.7 B
Context
Downloads 30d
1.3 M
Likes
181
Licence: other Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 10 days
1.3 M1.3 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors bf16 17.4 GB 21.0 GB ⚠️ Tight — reduce context
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Parameters
8.7 B
Tensor type
BF16
Licence
other
First seen on the Hub
2025-06-03
Training datasets
undisclosed
Added to our catalog
2026-07-28