unsloth / text-generation updated 1 year ago

Qwen3-235B-A22B-GGUF

Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants.

Params
Context
40,960
Downloads 30d
184 K
Likes
77
Commercial use: allowed apache-2.0 Not gated GGUF View on Hugging Face ↗

Download history

daily snapshots · 10 days
184 K173 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
Qwen3-235B-A22B-BF16-00003-of-00010.gguf GGUF 49.8 GB 55.2 GB ❌ Won’t fit
Qwen3-235B-A22B-Q2_K_L-00001-of-00002.gguf Q2_K_L 49.7 GB 55.2 GB ❌ Won’t fit
Qwen3-235B-A22B-Q3_K_M-00001-of-00003.gguf Q3_K_M 49.8 GB 55.3 GB ❌ Won’t fit
Qwen3-235B-A22B-Q4_1-00002-of-00003.gguf Q4 49.8 GB 55.3 GB ❌ Won’t fit
Qwen3-235B-A22B-Q2_K-00001-of-00002.gguf Q2_K 49.9 GB 55.4 GB ❌ Won’t fit
Qwen3-235B-A22B-Q4_K_M-00001-of-00003.gguf Q4_K_M 49.9 GB 55.4 GB ❌ Won’t fit
Qwen3-235B-A22B-IQ4_XS-00001-of-00003.gguf IQ4_XS 50.0 GB 55.5 GB ❌ Won’t fit
Qwen3-235B-A22B-Q3_K_S-00002-of-00003.gguf Q3_K_S 50.0 GB 55.5 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · Q4
$ ollama run qwen3-235b-a22b-gguf

# pin the quantization explicitly
$ ollama run qwen3-235b-a22b-gguf-q4
est. VRAM 55.3 GBon RTX 4090 · 24 GBJSON API →

Specifications

Architecture
Qwen3MoeForCausalLM
Context length
40,960
Vocabulary
151,936
Layers / heads
94 / 64
Licence
apache-2.0
First seen on the Hub
2025-04-28
Base model
Qwen3-235B-A22B
Training datasets
undisclosed
Added to our catalog
2026-07-28

Family

Base model and the most-downloaded derivatives in the catalog.

Sponsored · GPU cloud
Not enough VRAM?
Spin up a 24 GB L4 instance in 40 seconds. $0.44/hr.