DeepSeek-V4-Flash vs Qwen3-235B-A22B-Instruct-2507
Specs, VRAM requirements and download trends — updated 22 September 2026.
Qwen · B
Params
235.1 B
Context
262,144
30d
197 K
apache-2.0 · commercial OK
Specification comparison
Differences are highlighted; identical values are muted.
| Specification | DeepSeek-V4-Flash | Qwen3-235B-A22B-Instruct-2507 |
|---|---|---|
| Parameters | 158.1 B | 235.1 B |
| Architecture | DeepseekV4ForCausalLM | Qwen3MoeForCausalLM |
| Context length | 1,048,576 | 262,144 |
| Licence | mit | apache-2.0 |
| Commercial use | Allowed | Allowed |
| Languages | — | — |
| Downloads 30d | 1,524,356 | 197,322 |
| Downloads all time | 12.0 M | 1.8 M |
| Quantizations on the Hub | SAFETENSORS | SAFETENSORS |
| Gated | No | No |
| First seen on the Hub | 2026-04-22 | 2025-07-21 |
Download trend
57 days of snapshots recorded — chart in 0 days.
12.0 M
DeepSeek-V4-Flash · all-time downloads
1.8 M
Qwen3-235B-A22B-Instruct-2507 · all-time downloads
VRAM side by side
On RTX 4090 · 24 GB · 8K context unless noted
Qwen3-235B-A22B-Instruct-2507 · fp16 @ 256K ctx
1,646.2 GB / 24 GB
DeepSeek-V4-Flash · fp16 @ 1024K ctx
3,211.0 GB / 24 GB