What LLMs can an RTX 3090 run?

Which open models run on an RTX 3090: what fits at Q4 and Q8, estimated tokens per second, how much context fits, and where the card stops.

Updated 9 September 2026 · estimates are labelled as estimates

24 GB VRAM 936 GB/s 350 W 2020 Mostly second-hand

RTX 3090 24 GB addresses 24 GB at 936 GB/s. Of the 16 open models tracked on this site, it runs 13 at Q4_K_M with 8k of context. The largest is Qwen3 32B, needing about 22.4 GB and generating an estimated 34 tokens per second — comfortable for chat.

The value benchmark for local models, and it has been for years.

What an RTX 3090 runs, model by model

Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 22.8 GB to spend once the 5% safety margin comes off its 24 GB, and it reads that memory at 936 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.

ModelSizeNeedsEst. tokens/sOn this card
Llama 3.2 3B 3.2B 3.4 GB 353 fits
Llama 3.1 8B 8B 6.4 GB 141 fits
Qwen3 8B 8.2B 6.7 GB 138 fits
Gemma 3 12B 12.2B 11.1 GB 93 fits
Qwen3 14B 14.8B 10.8 GB 76 fits
Phi-4 14B 14.7B 11 GB 77 fits
gpt-oss 20B 21B · 3.6B active 13.6 GB 314 fits
Mistral Small 3.1 24B 24B 16.3 GB 47 fits
Gemma 3 27B 27.4B 21.2 GB 41 tight fit
Qwen3 30B-A3B (MoE) 30.5B · 3.3B active 19.7 GB 342 fits
Qwen3 32B 32.8B 22.4 GB 34 tight fit
Qwen2.5 Coder 32B 32.8B 22.4 GB 34 tight fit
DeepSeek-R1 Distill Qwen 32B 32.8B 22.4 GB 34 tight fit
Llama 3.3 70B 70.6B 45.8 GB no
DeepSeek-R1 Distill Llama 70B 70.6B 45.8 GB no
gpt-oss 120B 117B · 5.1B active 71.7 GB no

On this card that means anything at or under 22.8 GB counts as fitting, and anything above 20.4 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 24 GB can push it over.

The model to actually run on it

Qwen3 30B-A3B (MoE) is the best use of this card: 19.7 GB of the 24 GB available, an estimated 342 tokens per second, faster than you can read, and 38,912 tokens of context still available. It is a mixture-of-experts model, so all 30.5B parameters sit in memory but only 3.3B are read per token — which is why it is quick for its size.

Qwen3 32B is the largest model the card will hold, at 22.4 GB, but it is the wrong daily driver: filling 93% of memory with weights leaves room for only 9,216 tokens of conversation, and it generates at 34 tokens per second against 342. More parameters are not worth a context window that runs out mid-document.

If quality matters more than parameter count, Qwen3 14B fits at Q8_0 in about 18.2 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.

How much context actually fits

Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 24 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.

ModelMax context, FP16 KVWith Q8 KV cache
Qwen3 32B 9,216 tokens 18,432 tokens
Qwen2.5 Coder 32B 9,216 tokens 18,432 tokens
DeepSeek-R1 Distill Qwen 32B 9,216 tokens 18,432 tokens
Qwen3 30B-A3B (MoE) 38,912 tokens 78,848 tokens

Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 32B on this card goes from 9,216 to 18,432 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.

Where this card stops

The first model out of reach is Llama 3.3 70B: about 45.8 GB at Q4_K_M and 8k context, against 24 GB of usable memory. Dropping to Q3_K_M would need about 37.7 GB, which still does not fit. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.

The honest take

Twenty-four gigabytes at 936 GB/s, on the second-hand market, is still the best ratio of capability to money in local inference. It runs everything a 4090 runs, at roughly 93% of the speed, because the two cards are much closer in memory bandwidth than in compute — and text generation is bound by bandwidth, not compute. The catch is that you are buying a five-year-old card, often one that has worked hard.

What it is good at

  • 24 GB runs 32B-class models at Q4, which is the size where open models stop feeling like toys.
  • Within about 7% of a 4090 for single-stream text generation, at a fraction of the price.
  • Supports NVLink in pairs, so two of them can share a 48 GB pool for training work.

What it is not

  • 350 W, hot, and physically large; two of them need a case and a power supply that were planned for it.
  • Used stock has often run mining or rendering workloads. Thermal pads on the rear memory are a known weak point.
  • Slower than an Ada card for image and video generation, where compute matters more.

The thing people get wrong: The 3090 is the last GeForce card with NVLink. The 4090 dropped it. If you want two consumer cards to behave as one 48 GB pool for fine-tuning, this is the only GeForce generation that does it.

Buy or rent

The strongest buy case on this page, and also the one with the most risk attached, because condition varies. If you are going to own hardware for local models, this is usually the card the arithmetic points at.

Second-hand pricing is the whole argument for this card, and it moves week to week, so printing a number here would be worse than printing none. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 350 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.

Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.

Compared with the alternatives

  • RTX 409024 GB · 1,008 GB/s · best fit Qwen3 30B-A3B (MoE) at ~369 tok/s. The card most local-model advice is implicitly written for.
  • RTX 509032 GB · 1,792 GB/s · best fit Qwen3 32B at ~66 tok/s. The first consumer card whose memory bandwidth changes what is comfortable.
  • RTX 6000 Ada48 GB · 960 GB/s · best fit Qwen3 32B at ~35 tok/s. Forty-eight gigabytes in a normal computer, at 300 watts, without the noise.

All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.

Questions

Can an RTX 3090 run a 70B model?

Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 24 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.

What is the best model to run on an RTX 3090?

Qwen3 30B-A3B (MoE). At Q4_K_M and 8k context it needs about 19.7 GB of the 24 GB available and generates an estimated 342 tokens per second, which is faster than you can read. If quality matters more than size, Qwen3 14B fits at Q8 in about 18.2 GB.

How many tokens per second does an RTX 3090 generate?

It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 936 GB/s and 70% efficiency, this card produces an estimated 141 tokens per second on an 8B model at Q4, and about 34 on the largest model it holds, Qwen3 32B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.

Is 24 GB enough for running LLMs locally?

It runs 13 of the 16 models tracked here at Q4_K_M with 8k of context, up to 32.8B parameters. The honest test is not the model list but the context: Qwen3 32B on this card holds about 9,216 tokens before memory runs out.

RTX 3090 or RTX 4090 for local models?

Both address about the same memory, so they run the same models. On speed, the RTX 4090 is faster: 1,008 GB/s against 936 GB/s, and bandwidth is what sets chat speed.

Is a used RTX 3090 safe to buy for AI work?

Usually, with care. Many were used for mining or rendering, and the known weak point is the memory on the back of the board, where the original thermal pads degrade — repadding is a common and inexpensive fix. Ask for a photo of the card running a load, check the fans, and prefer a seller who will take a return.

Can two RTX 3090s be used together for larger models?

Yes, and this is the 3090 argument that no newer GeForce card can make: it is the last with NVLink, so a pair can behave as one 48 GB pool for fine-tuning. For inference, splitting a model across two cards works over PCIe without NVLink too, but the bandwidth between them becomes the limit.

  • Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
  • Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
  • Card specification: 24 GB, 936 GB/s, 350 W — manufacturer figures.
  • Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.