legraphista / text-generation updated 2 years ago

glm-4-9b-chat-IMat-GGUF

Original Model: THUDM/glm-4-9b-chat Original dtype: BF16 (bfloat16) Quantized by: https://github.com/ggerganov/llama.cpp/pull/6999 IMatrix dataset: here

Params
Context
Downloads 30d
993 K
Likes
5
Licence: other Not gated GGUF 2 languages View on Hugging Face ↗

Download history

daily snapshots · 14 days
993 K183 K
Aug 27Aug 31Sep 5Sep 9

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
glm-4-9b-chat.IQ1_S.gguf IQ1_S 3.1 GB 3.9 GB ✅ Runs comfortably
glm-4-9b-chat.IQ1_M.gguf IQ1_M 3.2 GB 4.0 GB ✅ Runs comfortably
glm-4-9b-chat.IQ2_XXS.gguf IQ2_XXS 3.4 GB 4.3 GB ✅ Runs comfortably
glm-4-9b-chat.IQ2_XS.gguf IQ2_XS 3.6 GB 4.5 GB ✅ Runs comfortably
glm-4-9b-chat.IQ2_S.gguf IQ2_S 3.8 GB 4.6 GB ✅ Runs comfortably
glm-4-9b-chat.IQ2_M.gguf IQ2_M 3.9 GB 4.8 GB ✅ Runs comfortably
glm-4-9b-chat.IQ3_M.gguf IQ3_M 4.8 GB 5.8 GB ✅ Runs comfortably
glm-4-9b-chat.BF16.gguf GGUF 18.8 GB 21.2 GB ⚠️ Tight — reduce context
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · GGUF
$ ollama run glm-4-9b-chat-imat-gguf

# pin the quantization explicitly
$ ollama run glm-4-9b-chat-imat-gguf-gguf
est. VRAM 21.2 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
other
First seen on the Hub
2024-06-20
Training datasets
undisclosed
Added to our catalog
2026-08-27
Sponsored · GPU cloud
Not enough VRAM?
Spin up a 24 GB L4 instance in 40 seconds. $0.44/hr.