Qwen3-Next-80B-A3B-Instruct-FP8
Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next-generation foundation models Qwen3-Next.
Params
81.3 B
Context
262,144
Downloads 30d
185 K
Likes
90
Download history
daily snapshots · 56 days
▲ 93 K in the last 30 days (33.4%)
286 K185 K
Aug 23Sep 2Sep 12Sep 21
339 K185 K
Jul 28Aug 15Sep 3Sep 21
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| model.safetensors | f8_e4m3 | 82.1 GB | 103.0 GB | ❌ Won’t fit |
| model.safetensors (bf16, full) | bf16 + 256K ctx | 82.1 GB | 481.1 GB | ❌ Won’t fit |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Run it
copy-paste, exact tags checked against the Hub$ curl -s https://aimodelscomparison.com/api/v1/models/qwen3-next-80b-a3b-instruct-fp8
{
"hf_id": "Qwen/Qwen3-Next-80B-A3B-Instruct-FP8",
"params_b": 81.33,
"context_length": 262144,
"license": { "id": "apache-2.0", "commercial": "yes" },
"downloads_30d": 185311,
"vram_estimates": [
{ "quant": "f8_e4m3", "gb": 103.0 }
],
"updated_at": "2026-07-28T18:06:17Z"
}
Specifications
- Architecture
- Qwen3NextForCausalLM
- Parameters
- 81.3 B
- Tensor type
- F8_E4M3
- Context length
- 262,144
- Vocabulary
- 151,936
- Layers / heads
- 48 / 16
- Licence
- apache-2.0
- First seen on the Hub
- 2025-09-22
- Base model
- Qwen3-Next-80B-A3B-Instruct
- Training datasets
- undisclosed
- Added to our catalog
- 2026-07-28
Family
Base model and the most-downloaded derivatives in the catalog.
Compare with any text-generation model