Qwen3 30B-A3B (MoE): VRAM requirements and which GPUs run it

How much VRAM Qwen3 30B-A3B (MoE) needs at Q4, Q5 and Q8, the smallest GPU that fits, and expected tokens per second on common cards.

Updated 8 September 2026 · estimates are labelled as estimates

Qwen3 30B-A3B (MoE) has 30.5B parameters with 3.3B active per token (mixture of experts), 48 layers and 4 KV heads of dimension 128. At the everyday setting, Q4_K_M and 8k context, it needs an estimated 19.7 GB, so the smallest card that fits is 24 GB.

Memory by quantisation and context

Quant4,096 ctx8,192 ctx32,768 ctx
Q8_034.5 GB · fits 48 GB34.9 GB · fits 48 GB37.3 GB · fits 48 GB
Q5_K_M23.4 GB · fits 32 GB23.8 GB · fits 32 GB26.2 GB · fits 32 GB
Q4_K_M19.3 GB · fits 24 GB19.7 GB · fits 24 GB22.1 GB · fits 24 GB

Weights at this quant: 17.7 GB at Q4. Every extra 1,000 tokens of context adds about 0.0983 GB of KV cache at FP16. Try other settings in the VRAM calculator.

Speed by GPU, at Q4 and 8k context

GPUMemoryBandwidthFitsEst. tokens/sFeels like
RTX 3060 12 GB12 GB360 GB/snodoes not fit
RTX 4060 Ti 16 GB16 GB288 GB/snodoes not fit
RTX 4070 12 GB12 GB504 GB/snodoes not fit
RTX 3090 24 GB24 GB936 GB/syes342faster than you can read
RTX 4090 24 GB24 GB1,008 GB/syes369faster than you can read
RTX 5090 32 GB32 GB1,792 GB/syes655faster than you can read
RTX 6000 Ada 48 GB48 GB960 GB/syes351faster than you can read
L40S 48 GB48 GB864 GB/syes316faster than you can read
RTX PRO 6000 Blackwell 96 GB96 GB1,792 GB/syes655faster than you can read
A100 80 GB80 GB2,039 GB/syes746faster than you can read
H100 SXM 80 GB80 GB3,352 GB/syes1226faster than you can read

Single-stream decode ceiling from memory bandwidth at 70% efficiency. Prompt processing and batching not included. See the speed estimator for other quantisations and Apple hardware, or every card compared if you are choosing hardware rather than a model.

Notes

  • All experts stay in memory; only 3.3B are active per token, so it is fast for its size.
  • Architecture values from Qwen/Qwen3-30B-A3B config.json. Verify against the model card before buying hardware for this model.
  • Estimates assume a single conversation on a card that is otherwise free. A desktop on the same GPU takes 0.5 to 2 GB.

Questions

Can I run Qwen3 30B-A3B (MoE) on a 24 GB card?

Yes, at Q4_K_M and 8k context it needs about 19.7 GB, which fits a 24 GB card with room.

How much VRAM does Qwen3 30B-A3B (MoE) need at Q8?

About 34.9 GB at 8k context, or 37.3 GB at 32k. Q8 is near-lossless; use it when it fits.

How fast is Qwen3 30B-A3B (MoE) on an RTX 4090?

Roughly 369 tokens per second at Q4, single stream, which is faster than you can read. Real runtimes land within about 20% of this either way.

See how it compares in which models fit on 8 to 96 GB, or run it without buying the card: Nodegrove attaches a 24, 48 or 96 GB GPU to a workspace that stays saved.