nvidia / text-generation updated 5 months ago

MiniMax-M2.5-NVFP4

The NVIDIA MiniMax-M2.5-NVFP4 model is the quantized version of MiniMax's MiniMax-M2.5 model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA MiniMax-M2.5 NVFP4 model is quantized with Nvidia Model Optimizer.

Params
116.4 B
Context
196,608
Downloads 30d
222 K
Likes
38
Licence: other Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 39 days
▲ 311 K in the last 30 days (58.3%)
777 K222 K
Aug 14Aug 27Sep 9Sep 21

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors u8 139.9 GB 171.8 GB ❌ Won’t fit
model.safetensors (bf16, full) bf16 + 192K ctx 139.9 GB 573.2 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · curl · api/v1
$ curl -s https://aimodelscomparison.com/api/v1/models/minimax-m2-5-nvfp4
{
  "hf_id": "nvidia/MiniMax-M2.5-NVFP4",
  "params_b": 116.35,
  "context_length": 196608,
  "license": { "id": "other", "commercial": "unknown" },
  "downloads_30d": 221989,
  "vram_estimates": [
    { "quant": "u8", "gb": 171.8 }
  ],
  "updated_at": "2026-08-14T01:00:24Z"
}
est. VRAM —on RTX 4090 · 24 GBJSON API →

Specifications

Architecture
MiniMaxM2ForCausalLM
Parameters
116.4 B
Tensor type
U8
Context length
196,608
Vocabulary
200,064
Layers / heads
62 / 48
Licence
other
First seen on the Hub
2026-03-11
Base model
MiniMax-M2.5
Training datasets
undisclosed
Added to our catalog
2026-08-14

Family

Base model and the most-downloaded derivatives in the catalog.

Compare with any text-generation model