unsloth / image-text-to-text updated 2 weeks ago

gemma-4-26B-A4B-it-GGUF

Jun 9 Update: Added MTP support. See our MTP Guide. Apr 11 Update: Re-download for Google's latest chat template and llama.cpp fixes. Gemma 4 can now be run and fine-tuned in Unsloth Studio. Read our guide. See all versions of Gemma 4 (GGUF, 16-bit etc.) in our collection. Example of Gemma 4 E4B (4-bit GGUF) running in Unsloth Studio with tool-calling:

Params
Context
Downloads 30d
1.4 M
Likes
1,029
Commercial use: allowed apache-2.0 Not gated GGUF View on Hugging Face ↗

Download history

daily snapshots · 10 days
1.5 M1.4 M
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
gemma-4-26B-A4B-it-UD-IQ2_XXS.gguf IQ2_XXS 9.9 GB 11.4 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-UD-IQ2_M.gguf IQ2_M 10.0 GB 11.5 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-UD-IQ3_S.gguf IQ3_S 11.3 GB 12.9 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-UD-IQ3_XXS.gguf IQ3_XXS 11.4 GB 13.1 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-UD-IQ4_NL.gguf IQ4_NL 13.6 GB 15.5 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-UD-IQ4_XS.gguf IQ4_XS 13.6 GB 15.5 GB ✅ Runs comfortably
gemma-4-26B-A4B-it-Q8_0.gguf Q8_0 26.9 GB 30.0 GB ❌ Won’t fit
gemma-4-26B-A4B-it-BF16-00001-of-00002.gguf GGUF 49.9 GB 55.4 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
Gemma4ForConditionalGeneration
Licence
apache-2.0
First seen on the Hub
2026-04-01
Base model
gemma-4-26B-A4B-it
Training datasets
undisclosed
Added to our catalog
2026-07-28

Family

Base model and the most-downloaded derivatives in the catalog.