Inkling-Small-GGUF
See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. You can now run Inkling in Unsloth Studio with toggles for Thinking. Read our Inkling guide for analysis and instructions. See below for example of 1-bit UD-IQ1S GGUF running in Unsloth:
Params
—
Context
—
Downloads 30d
1.4 M
Likes
89
Download history
daily snapshots · 25 days1.4 M855 K
Aug 27Sep 4Sep 12Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 · 24 GB |
|---|---|---|---|---|
| Inkling-Small-Q8_0-00005-of-00007.gguf | Q8 | 48.6 GB | 54.0 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ2_XXS-00002-of-00003.gguf | IQ2_XXS | 49.4 GB | 54.8 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ3_XXS-00002-of-00003.gguf | IQ3_XXS | 49.4 GB | 54.8 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ2_M-00002-of-00003.gguf | IQ2_M | 49.5 GB | 54.9 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ1_M-00002-of-00003.gguf | IQ1_M | 49.7 GB | 55.1 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ1_S-00002-of-00003.gguf | IQ1_S | 49.7 GB | 55.1 GB | ❌ Won’t fit |
| Inkling-Small-UD-IQ3_S-00002-of-00004.gguf | IQ3_S | 49.7 GB | 55.2 GB | ❌ Won’t fit |
| Inkling-Small-BF16-00007-of-00011.gguf | GGUF | 50.0 GB | 55.5 GB | ❌ Won’t fit |
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.
Specifications
- Licence
- apache-2.0
- First seen on the Hub
- 2026-07-30
- Training datasets
- undisclosed
- Added to our catalog
- 2026-08-27
Compare with any image-text-to-text model