NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
:---:--- Total Parameters 550B (55B active) Architecture LatentMoE - Mamba-2 + MoE + Attention hybrid with Multi-Token Prediction (MTP) Context Length Up to 1M tokens Minimum GPU Requirement 4xGB200, 4xB200, 4x GB300, 4x B300, 8xH100 Supported Languages English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, and Chinese Best For Frontier reasoning, complex agentic workflows, long-con...
Params
302.8 B
Context
262,144
Downloads 30d
345 K
Likes
332
Download history
daily snapshots · 55 days
▲ 10 K in the last 30 days (2.9%)
441 K273 K
Aug 22Sep 1Sep 11Sep 20
441 K209 K
Jul 28Aug 15Sep 2Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | u8 | 352.3 GB | 433.5 GB | ❌ Won’t fit |
| model.safetensors (bf16, full) | bf16 + 256K ctx | 352.3 GB | 1,841.6 GB | ❌ Won’t fit |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Run it
copy-paste, exact tags checked against the Hub$ curl -s https://aimodelscomparison.com/api/v1/models/nvidia-nemotron-3-ultra-550b-a55b-nvfp4
{
"hf_id": "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4",
"params_b": 302.83,
"context_length": 262144,
"license": { "id": "other", "commercial": "unknown" },
"downloads_30d": 345437,
"vram_estimates": [
{ "quant": "u8", "gb": 433.5 }
],
"updated_at": "2026-08-25T01:00:42Z"
}
Specifications
- Architecture
- NemotronHForCausalLM
- Parameters
- 302.8 B
- Tensor type
- U8
- Context length
- 262,144
- Vocabulary
- 131,072
- Licence
- other
- First seen on the Hub
- 2026-06-03
- Training datasets
- nvidia/nemotron-post-training-v3, nvidia/nemotron-pre-training-datasets
- Added to our catalog
- 2026-07-28
Compare with any text-generation model
Popular comparisons
Appears in