bartowski / text-generation updated 1 month ago

Kwaipilot_KAT-Coder-V2.5-Dev-GGUF

- llama.cpp - ramalama - LM Studio - koboldcpp - Jan AI - Text Generation Web UI - LoLLMs - Atomic Chat

Params
Context
Downloads 30d
396 K
Likes
133
Commercial use: allowed apache-2.0 Not gated GGUF 2 languages View on Hugging Face ↗

Download history

daily snapshots · 40 days
▲ 120 K in the last 30 days (43.6%)
397 K164 K
Aug 12Aug 25Sep 7Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
Kwaipilot_KAT-Coder-V2.5-Dev-IQ2_XXS.gguf IQ2_XXS 9.8 GB 11.3 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ2_XS.gguf IQ2_XS 10.8 GB 12.4 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ2_S.gguf IQ2_S 11.0 GB 12.6 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ2_M.gguf IQ2_M 12.1 GB 13.8 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ3_XXS.gguf IQ3_XXS 14.9 GB 16.9 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ3_XS.gguf IQ3_XS 16.2 GB 18.3 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ3_M.gguf IQ3_M 16.9 GB 19.1 GB ✅ Runs comfortably
Kwaipilot_KAT-Coder-V2.5-Dev-IQ4_NL.gguf IQ4_NL 19.9 GB 22.3 GB ⚠️ Tight — reduce context
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · IQ2_M
$ ollama run kwaipilot-kat-coder-v2-5-dev-gguf

# pin the quantization explicitly
$ ollama run kwaipilot-kat-coder-v2-5-dev-gguf-iq2_m
est. VRAM 13.8 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
apache-2.0
First seen on the Hub
2026-07-23
Training datasets
undisclosed
Added to our catalog
2026-08-12
Compare with any text-generation model