unsloth / image-text-to-text updated 2 weeks ago

Qwen3.8-Flash-Next-GGUF

Unsloth Dynamic 3.0 achieves superior accuracy & outperforms other leading quants. MTP is now available for 1.3-1.7x Faster inference in Unsloth. Read Guide To run, please use llama.cpp or use our Unsloth Desktop app. See below for Qwen3.8-Flash-Next run in Unsloth Desktop with thinking controls:

Params
Context
Downloads 30d
1.5 M
Likes
1,003
Licence: other Not gated GGUF View on Hugging Face ↗

Download history

daily snapshots · 12 days
1.5 M936 K
Sep 9Sep 13Sep 17Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
mtp-Qwen3.8-Flash-Next-Q4_K_M.gguf Q4_K_M 2.8 GB 3.6 GB ✅ Runs comfortably
mtp-Qwen3.8-Flash-Next-Q8_0.gguf Q8_0 4.1 GB 5.1 GB ✅ Runs comfortably
Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.gguf IQ3_XXS 49.6 GB 55.0 GB ❌ Won’t fit
Qwen3.8-Flash-Next-UD-IQ4_XS-00002-of-00003.gguf IQ4_XS 49.8 GB 55.3 GB ❌ Won’t fit
Qwen3.8-Flash-Next-UD-IQ1_M-00002-of-00003.gguf IQ1_M 50.0 GB 55.5 GB ❌ Won’t fit
Qwen3.8-Flash-Next-UD-IQ1_S-00002-of-00003.gguf IQ1_S 50.0 GB 55.5 GB ❌ Won’t fit
Qwen3.8-Flash-Next-Q8_0-00003-of-00006.gguf Q8 54.4 GB 60.3 GB ❌ Won’t fit
Qwen3.8-Flash-Next-BF16-00003-of-00008.gguf GGUF 102.4 GB 113.1 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Licence
other
First seen on the Hub
2026-08-26
Training datasets
undisclosed
Added to our catalog
2026-09-09
Compare with any image-text-to-text model