GLM-5.3-Flash
π Join our WeChat or Discord community. π Check out the GLM-5.3-Flash blog and GLM-5 Technical report. π Use GLM-5.3-Flash API services on Z.ai API Platform.
Params
321.3 B
Context
β
Downloads 30d
2.9 M
Likes
2,480
Download history
daily snapshots Β· 22 days2.9 M190 K
Aug 30Sep 6Sep 13Sep 20
Can you run it?
Estimated VRAM at 8K context unless noted. Pick your hardware to see the verdict per quantization.
| File | Quant | Size | Est. VRAM | Verdict on RTX 4090 Β· 24 GB |
|---|---|---|---|---|
| model.safetensors | f8_e4m3 | 328.3 GB | 409.9 GB | β Wonβt fit |
Estimate: file size Γ 1.1 + KV cache at 8K + 0.5 GB overhead. Not a benchmark β how we calculate this.
Specifications
- Architecture
- Glm5NextForConditionalGeneration
- Parameters
- 321.3 B
- Tensor type
- F8_E4M3
- Licence
- mit
- First seen on the Hub
- 2026-08-25
- Training datasets
- undisclosed
- Added to our catalog
- 2026-08-30
Family
Base model and the most-downloaded derivatives in the catalog.
Compare with any image-text-to-text model
Popular comparisons
Appears in