tvall43 / text-generation updated 3 months ago

Qwen3.6-14B-A3B-FableVibes-GGUF

GGUF quantizations of Qwen3.6-14B-A3B-FableVibes, a 14B MoE model fine-tuned on reasoning traces from Claude Fable 5.

Params
Context
Downloads 30d
419 K
Likes
194
Commercial use: allowed apache-2.0 Not gated GGUF 1 languages View on Hugging Face ↗

Download history

daily snapshots · 36 days
▲ 194 K in the last 30 days (86.7%)
432 K176 K
Aug 17Aug 29Sep 10Sep 21

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
Qwen3.6-14B-A3B-FableVibes-Q2_K.gguf Q2_K 5.3 GB 6.4 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-Q3_K_M.gguf Q3_K_M 6.8 GB 7.9 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-Q4_K_M.gguf Q4_K_M 8.5 GB 9.8 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-MXFP4_MOE.gguf GGUF 8.6 GB 10.0 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-Q5_K_M.gguf Q5_K_M 9.9 GB 11.3 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-Q6_K.gguf Q6_K 11.3 GB 13.0 GB ✅ Runs comfortably
Qwen3.6-14B-A3B-FableVibes-Q8_0.gguf Q8_0 14.7 GB 16.6 GB ✅ Runs comfortably
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Run it

copy-paste, exact tags checked against the Hub
~ · ollama · Q4_K_M
$ ollama run qwen3-6-14b-a3b-fablevibes-gguf

# pin the quantization explicitly
$ ollama run qwen3-6-14b-a3b-fablevibes-gguf-q4_k_m
est. VRAM 9.8 GBon RTX 4090 · 24 GBJSON API →

Specifications

Licence
apache-2.0
First seen on the Hub
2026-06-14
Training datasets
Glint-Research/Fable-5-traces, angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k, osieosie/tmax-sft-skill-tax-20260505-2.2k-combined-balanced-qwen3.6-27b-thinking, zake7749/Qwen3.6-35B-A3B-Tool-Calling, nickrosh/Evol-Instruct-Code-80k-v1
Added to our catalog
2026-08-17
Compare with any text-generation model