DeepSeek-R1 Distill Qwen 32B has 32.8B parameters, 64 layers and 8 KV heads of dimension 128. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 22.4 GB, so the smallest card that fits is 24 GB.
Memory by quantisation and context
| Quant | 4,096 ctx | 8,192 ctx | 32,768 ctx |
|---|---|---|---|
| Q8_0 | 37.7 GB · fits 48 GB | 38.8 GB · fits 48 GB | 45.2 GB · fits 48 GB |
| Q5_K_M | 25.8 GB · fits 32 GB | 26.9 GB · fits 32 GB | 33.3 GB · fits 48 GB |
| Q4_K_M | 21.4 GB · fits 24 GB | 22.4 GB · fits 24 GB | 28.9 GB · fits 32 GB |
Weights at this quant: 19 GB at Q4. Every extra 1,000 tokens of context adds about 0.2621 GB of KV cache at FP16. Try other settings in the VRAM calculator.
Speed by GPU, at Q4 and 8k context
| GPU | Memory | Bandwidth | Fits | Est. tokens/s | Feels like |
|---|---|---|---|---|---|
| RTX 3060 12 GB | 12 GB | 360 GB/s | no | – | does not fit |
| RTX 4060 Ti 16 GB | 16 GB | 288 GB/s | no | – | does not fit |
| RTX 4070 12 GB | 12 GB | 504 GB/s | no | – | does not fit |
| RTX 3090 24 GB | 24 GB | 936 GB/s | tight | 34 | comfortable for chat |
| RTX 4090 24 GB | 24 GB | 1,008 GB/s | tight | 37 | comfortable for chat |
| RTX 5090 32 GB | 32 GB | 1,792 GB/s | yes | 66 | faster than you can read |
| RTX 6000 Ada 48 GB | 48 GB | 960 GB/s | yes | 35 | comfortable for chat |
| L40S 48 GB | 48 GB | 864 GB/s | yes | 32 | comfortable for chat |
| RTX PRO 6000 Blackwell 96 GB | 96 GB | 1,792 GB/s | yes | 66 | faster than you can read |
| A100 80 GB | 80 GB | 2,039 GB/s | yes | 75 | faster than you can read |
| H100 SXM 80 GB | 80 GB | 3,352 GB/s | yes | 123 | faster than you can read |
Single-stream decode ceiling from memory bandwidth at 70% efficiency. Prompt processing and batching not included. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.
Notes
- Architecture values from deepseek-ai/DeepSeek-R1-Distill-Qwen-32B config.json. Verify against the model card before buying hardware for this model.
- Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.
Questions
Can I run DeepSeek-R1 Distill Qwen 32B on a 24 GB card?
Yes, at Q4_K_M and 8k context it needs about 22.4 GB, which fits a 24 GB card tightly.
How much VRAM does DeepSeek-R1 Distill Qwen 32B need at Q8?
About 38.8 GB at 8k context, or 45.2 GB at 32k. Q8 is near-lossless; use it when it fits.
How fast is DeepSeek-R1 Distill Qwen 32B on an RTX 4090?
Roughly 37 tokens per second at Q4, single stream, which is comfortable for chat. Real runtimes land within about 20% of this either way.
See how it compares in which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.