unsloth / text-generation updated 2 weeks ago

GLM-5.3-Flash-GGUF

Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants. To run, please use our llama.cpp PR or use the Unsloth Desktop app. You can now run GLM-5.3-Flash in our Unsloth Desktop UI. See below for 1-bit GLM-5.3-Flash (Low) run inside of Unsloth Desktop:

Params
Context
Downloads 30d
630 K
Likes
416
Commercial use: allowed mit Not gated GGUF 2 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
630 K285 K
Sep 11Sep 14Sep 17Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
GLM-5.3-Flash-UD-IQ3_XXS-00002-of-00004.gguf IQ3_XXS 49.3 GB 54.7 GB ❌ Won’t fit
GLM-5.3-Flash-UD-IQ1_S-00002-of-00003.gguf IQ1_S 49.6 GB 55.1 GB ❌ Won’t fit
GLM-5.3-Flash-UD-IQ2_XXS-00002-of-00004.gguf IQ2_XXS 49.9 GB 55.4 GB ❌ Won’t fit
GLM-5.3-Flash-UD-Q2_K_XL-00003-of-00004.gguf Q2_K_XL 49.9 GB 55.4 GB ❌ Won’t fit
GLM-5.3-Flash-BF16-00005-of-00014.gguf GGUF 50.0 GB 55.5 GB ❌ Won’t fit
GLM-5.3-Flash-Q8_0-00005-of-00008.gguf Q8 50.0 GB 55.5 GB ❌ Won’t fit
GLM-5.3-Flash-UD-IQ1_M-00002-of-00003.gguf IQ1_M 50.0 GB 55.5 GB ❌ Won’t fit
GLM-5.3-Flash-UD-IQ4_XS-00002-of-00005.gguf IQ4_XS 50.0 GB 55.5 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · GGUF
$ ollama run glm-5-3-flash-gguf

# pin the quantization explicitly
$ ollama run glm-5-3-flash-gguf-gguf
est. VRAM 55.5 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
mit
First seen on the Hub
2026-08-26
Base model
GLM-5.3-Flash
Training datasets
undisclosed
Added to our catalog
2026-09-11

Family

Base model and the most-downloaded derivatives in the catalog.

Compare with any text-generation model