sugoitoolkit / translation updated 11 months ago

Sugoi-32B-Ultra-GGUF

Unleashing the full potential of the previous sugoi 32B model, Sugoi 32B Ultra. Benchmark soon.

Params
Context
Downloads 30d
256 K
Likes
5
Commercial use: allowed apache-2.0 Not gated GGUF 2 languages View on Hugging Face ↗

Download history

daily snapshots · 10 days
260 K237 K
Jul 28Jul 31Aug 3Aug 6

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
Sugoi-32B-Ultra-Q2_K.gguf Q2_K 12.3 GB 14.0 GB ✅ Runs comfortably
Sugoi-32B-Ultra-Q4_K_M.gguf Q4_K_M 19.9 GB 22.3 GB ⚠️ Tight — reduce context
Sugoi-32B-Ultra-Q8_0.gguf Q8_0 34.8 GB 38.8 GB ❌ Won’t fit
Sugoi-32B-Ultra-F16.gguf GGUF 65.5 GB 72.6 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Licence
apache-2.0
First seen on the Hub
2025-08-23
Base model
Qwen2.5-32B-Instruct
Training datasets
undisclosed
Added to our catalog
2026-07-28

Family

Base model and the most-downloaded derivatives in the catalog.