24 GB VRAM 1,008 GB/s 450 W 2022 Mostly second-hand
RTX 4090 (24 GB) addresses 24 GB at 1,008 GB/s. Of the 16 open models tracked on this site, it runs 13 at Q4_K_M with 8k of context. The largest is Qwen3 32B, needing about 22.4 GB and generating an estimated 37 tokens per second — comfortable for chat.
The card most local-model advice is implicitly written for.
What an RTX 4090 runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 22.8 GB to spend once the 5% safety margin comes off its 24 GB, and it reads that memory at 1,008 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 380 | fits |
| Llama 3.1 8B | 8B | 6.4 GB | 152 | fits |
| Qwen3 8B | 8.2B | 6.7 GB | 148 | fits |
| Gemma 3 12B | 12.2B | 11.1 GB | 100 | fits |
| Qwen3 14B | 14.8B | 10.8 GB | 82 | fits |
| Phi-4 14B | 14.7B | 11 GB | 83 | fits |
| gpt-oss 20B | 21B · 3.6B active | 13.6 GB | 338 | fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 51 | fits |
| Gemma 3 27B | 27.4B | 21.2 GB | 44 | tight fit |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 369 | fits |
| Qwen3 32B | 32.8B | 22.4 GB | 37 | tight fit |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 37 | tight fit |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 37 | tight fit |
| Llama 3.3 70B | 70.6B | 45.8 GB | – | no |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | – | no |
| gpt-oss 120B | 117B · 5.1B active | 71.7 GB | – | no |
On this card that means anything at or under 22.8 GB counts as fitting, and anything above 20.4 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 24 GB can push it over.
The model to actually run on it
Qwen3 30B-A3B (MoE) is the best use of this card: 19.7 GB of the 24 GB available, an estimated 369 tokens per second, faster than you can read, and 38,912 tokens of context still available. It is a mixture-of-experts model, so all 30.5B parameters sit in memory but only 3.3B are read per token — which is why it is quick for its size.
Qwen3 32B is the largest model the card will hold, at 22.4 GB, but it is the wrong daily driver: filling 93% of memory with weights leaves room for only 9,216 tokens of conversation, and it generates at 37 tokens per second against 369. More parameters are not worth a context window that runs out mid-document.
If quality matters more than parameter count, Qwen3 14B fits at Q8_0 in about 18.2 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 24 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Qwen3 32B | 9,216 tokens | 18,432 tokens |
| Qwen2.5 Coder 32B | 9,216 tokens | 18,432 tokens |
| DeepSeek-R1 Distill Qwen 32B | 9,216 tokens | 18,432 tokens |
| Qwen3 30B-A3B (MoE) | 38,912 tokens | 78,848 tokens |
Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 32B on this card goes from 9,216 to 18,432 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.
Where this card stops
The first model out of reach is Llama 3.3 70B: about 45.8 GB at Q4_K_M and 8k context, against 24 GB of usable memory. Dropping to Q3_K_M would need about 37.7 GB, which still does not fit. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.
The honest take
The 4090 is the reference point: 24 GB, 1008 GB/s, and enough compute that nothing else in a desktop is the bottleneck. It runs 32B-class models at Q4 quickly and comfortably, handles image and video generation without complaint, and stops dead at 70B, which needs roughly twice its memory. Most "can I run this locally" questions are really asking whether it fits on a 4090.
What it is good at
- 1008 GB/s puts 32B models comfortably above reading speed.
- Compute headroom that a 3090 does not have, which shows up in image and video work and in prompt processing.
- The best-supported consumer card in every runtime, so problems have already been solved by someone else.
What it is not
- 24 GB is the same ceiling as the 3090 that costs much less. You are paying for speed, not capability.
- 450 W and a four-slot cooler on most models; plan the case and the power supply around it.
- No NVLink, so two cards cannot pool memory the way two 3090s can.
The thing people get wrong: A 4090 is about 7% faster than a 3090 at generating text, not the two-generation leap the rest of its specification suggests. Single-stream decoding reads weights out of memory, and 1008 GB/s against 936 GB/s is the whole story. The gap only opens up on prompt processing, image generation and batching.
Buy or rent
Worth it if you also generate images or video, or if prompt processing speed matters to you. If you only chat with text models, the 3090 does the same work for much less.
Second-hand pricing is the whole argument for this card, and it moves week to week, so printing a number here would be worse than printing none. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 450 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- RTX 3090 — 24 GB · 936 GB/s · best fit Qwen3 30B-A3B (MoE) at ~342 tok/s. The value benchmark for local models, and it has been for years.
- RTX 5090 — 32 GB · 1,792 GB/s · best fit Qwen3 32B at ~66 tok/s. The first consumer card whose memory bandwidth changes what is comfortable.
- RTX 6000 Ada — 48 GB · 960 GB/s · best fit Qwen3 32B at ~35 tok/s. Forty-eight gigabytes in a normal computer, at 300 watts, without the noise.
All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can an RTX 4090 run a 70B model?
Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 24 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.
What is the best model to run on an RTX 4090?
Qwen3 30B-A3B (MoE). At Q4_K_M and 8k context it needs about 19.7 GB of the 24 GB available and generates an estimated 369 tokens per second, which is faster than you can read. If quality matters more than size, Qwen3 14B fits at Q8 in about 18.2 GB.
How many tokens per second does an RTX 4090 generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 1,008 GB/s and 70% efficiency, this card produces an estimated 152 tokens per second on an 8B model at Q4, and about 37 on the largest model it holds, Qwen3 32B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.
Is 24 GB enough for running LLMs locally?
It runs 13 of the 16 models tracked here at Q4_K_M with 8k of context, up to 32.8B parameters. The honest test is not the model list but the context: Qwen3 32B on this card holds about 9,216 tokens before memory runs out.
RTX 4090 or RTX 3090 for local models?
Both address about the same memory, so they run the same models. On speed, this card is faster: 1,008 GB/s against 936 GB/s, and bandwidth is what sets chat speed.
RTX 4090 or RTX 3090 for running LLMs?
For text generation they are close: both hold 24 GB and the 4090 reads memory about 7% faster, which is what sets chat speed. The 4090 pulls ahead on prompt processing, image and video generation, and anything compute-bound. If you only chat with text models, the 3090 does the same work for much less; if you also generate images, the 4090 earns the difference.
Can two RTX 4090s run a 70B model?
Two cards give 48 GB of combined memory, which holds a 70B model at Q4 when the runtime splits it across both. What the 4090 cannot do is pool that memory over NVLink, which was removed on this generation, so the cards talk over PCIe and the link between them becomes the constraint for training.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
- Card specification: 24 GB, 1,008 GB/s, 450 W — manufacturer figures.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.