AtomicChat / text-generation updated 3 weeks ago

Qwen3.8-Flash-Next-GGUF

Built from Qwen's original weights with our own importance matrix. The calibration corpora behind our builds are public. Qwen3.8-Flash-Next is the first open-weight release of the architecture behind Qwen4. These GGUFs are self-quantized from Qwen's original weights with our own importance matrix, published alongside the quants. The quants are still uploading and need a llama.cpp build with Qwen3.8-Flash-Next support...

Params
Context
Downloads 30d
285 K
Likes
148
Licence: other Not gated GGUF View on Hugging Face ↗

Download history

tracking started — chart appears after 7 days of snapshots (3 recorded)
285 K downloads in the last 30 days

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
mmproj-Qwen3.8-Flash-Next-BF16.gguf GGUF 0.9 GB 1.5 GB ✅ Runs comfortably
Qwen3.8-Flash-Next-AD-3.84bpw-IQ4_XS-M64-00002-of-00028.gguf IQ4_XS 38.4 GB 42.7 GB ❌ Won’t fit
Qwen3.8-Flash-Next-AD-4.27bpw-Q4_K_M-M64-00002-of-00033.gguf Q4_K_M 38.4 GB 42.7 GB ❌ Won’t fit
Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64-00002-of-00033.gguf Q5_K_M 54.4 GB 60.3 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · Q4_K_M
$ ollama run qwen3-8-flash-next-gguf-atomicchat

# pin the quantization explicitly
$ ollama run qwen3-8-flash-next-gguf-atomicchat-q4_k_m
est. VRAM 42.7 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
other
First seen on the Hub
2026-08-26
Training datasets
undisclosed
Added to our catalog
2026-09-19
Compare with any text-generation model