moonshotai / image-text-to-text updated 2 weeks ago

Kimi-K3

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Params
2,779.9 B
Context
Downloads 30d
2.1 M
Likes
11,433
Licence: other Not gated SAFETENSORS View on Hugging Face ↗

Download history

daily snapshots · 48 days
▲ 213 K in the last 30 days (9.1%)
2.9 M968 K
Aug 4Aug 20Sep 5Sep 20

Can you run it?

Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.

FileQuantSizeEst. VRAMVerdict on RTX 4090 · 24 GB
model.safetensors u8 1,560.9 GB 2,134.5 GB ❌ Won’t fit
Estimate: file size × 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark — how we calculate this.

Specifications

Architecture
KimiK3ForConditionalGeneration
Parameters
2,779.9 B
Tensor type
U8
Licence
other
First seen on the Hub
2026-06-13
Training datasets
undisclosed
Added to our catalog
2026-08-04
Compare with any image-text-to-text model