unsloth / text-generation updated 1 month ago

GLM-5.2-GGUF

See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. You can now run GLM-5.2 in Unsloth Studio with toggles for High and Max thinking. Read our GLM-5.2 guide for analysis and instructions. See below for example of 1-bit UD-IQ1M GGUF running in Unsloth:

Params
Context
Downloads 30d
238 K
Likes
618
Commercial use: allowed mit Not gated GGUF 2 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
344 K238 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
GLM-5.2-BF16-00005-of-00033.gguf GGUF 46.8 GB 52.0 GB ❌ Won’t fit
GLM-5.2-UD-IQ2_M-00002-of-00006.gguf IQ2_M 49.2 GB 54.6 GB ❌ Won’t fit
GLM-5.2-UD-IQ2_XXS-00003-of-00006.gguf IQ2_XXS 49.1 GB 54.6 GB ❌ Won’t fit
GLM-5.2-Q8_0-00003-of-00017.gguf Q8 49.3 GB 54.7 GB ❌ Won’t fit
GLM-5.2-UD-IQ1_S-00003-of-00006.gguf IQ1_S 49.7 GB 55.2 GB ❌ Won’t fit
GLM-5.2-UD-IQ1_M-00005-of-00006.gguf IQ1_M 49.9 GB 55.4 GB ❌ Won’t fit
GLM-5.2-UD-IQ3_S-00003-of-00008.gguf IQ3_S 49.9 GB 55.4 GB ❌ Won’t fit
GLM-5.2-UD-IQ3_XXS-00002-of-00007.gguf IQ3_XXS 50.0 GB 55.5 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · GGUF
$ ollama run glm-5-2-gguf

# pin the quantization explicitly
$ ollama run glm-5-2-gguf-gguf
est. VRAM 52.0 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
mit
First seen on the Hub
2026-06-17
Base model
GLM-5.2
Training datasets
undisclosed
Added to our catalog
2026-07-28

Family

Base model and the most-downloaded derivatives in the catalog.

Sponsored · GPU cloud
Not enough VRAM?
Spin up a 24 GB L4 instance in 40 seconds. $0.44/hr.