12 GB VRAM 504 GB/s 200 W 2023 Sold new
RTX 4070 12 GB addresses 12 GB at 504 GB/s. Of the 16 open models tracked on this site, it runs 6 at Q4_K_M with 8k of context. The largest is Qwen3 14B, needing about 10.8 GB and generating an estimated 41 tokens per second — comfortable for chat.
Fast for its class, and capped by 12 GB.
What an RTX 4070 runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 11.4 GB to spend once the 5% safety margin comes off its 12 GB, and it reads that memory at 504 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 190 | fits |
| Llama 3.1 8B | 8B | 6.4 GB | 76 | fits |
| Qwen3 8B | 8.2B | 6.7 GB | 74 | fits |
| Gemma 3 12B | 12.2B | 11.1 GB | 50 | tight fit |
| Qwen3 14B | 14.8B | 10.8 GB | 41 | tight fit |
| Phi-4 14B | 14.7B | 11 GB | 41 | tight fit |
| gpt-oss 20B | 21B · 3.6B active | 13.6 GB | – | no |
| Mistral Small 3.1 24B | 24B | 16.3 GB | – | no |
| Gemma 3 27B | 27.4B | 21.2 GB | – | no |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | – | no |
| Qwen3 32B | 32.8B | 22.4 GB | – | no |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | – | no |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | – | no |
| Llama 3.3 70B | 70.6B | 45.8 GB | – | no |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | – | no |
| gpt-oss 120B | 117B · 5.1B active | 71.7 GB | – | no |
On this card that means anything at or under 11.4 GB counts as fitting, and anything above 10.2 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 12 GB can push it over.
The model to actually run on it
Qwen3 8B is the best use of this card: 6.7 GB of the 12 GB available, an estimated 74 tokens per second, faster than you can read, and 39,936 tokens of context still available.
Qwen3 14B is the largest model the card will hold, at 10.8 GB, but it is the wrong daily driver: filling 90% of memory with weights leaves room for only 11,264 tokens of conversation, and it generates at 41 tokens per second against 74. More parameters are not worth a context window that runs out mid-document.
It also fits at Q8_0, in about 10.7 GB, which is worth taking whenever the memory allows: Q8 is near-lossless where Q4 costs a little accuracy.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 12 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Qwen3 14B | 11,264 tokens | 23,552 tokens |
| Phi-4 14B | 9,216 tokens | 19,456 tokens |
| Gemma 3 12B | 8,192 tokens | 17,408 tokens |
| Qwen3 8B | 39,936 tokens | 79,872 tokens |
Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 14B on this card goes from 11,264 to 23,552 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.
Where this card stops
The first model out of reach is gpt-oss 20B: about 13.6 GB at Q4_K_M and 8k context, against 12 GB of usable memory. Dropping to Q3_K_M would need about 11.2 GB, which fits, though Q3 loses enough quality that a smaller model at Q4 is usually the better trade. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.
The honest take
The 4070 generates noticeably faster than either the 3060 or the 16 GB 4060 Ti, at 504 GB/s, and then runs out of memory at exactly the point where local models get interesting. For 7B and 8B models it is the best of the small cards: quick, quiet, efficient. For anything larger, the 12 GB ceiling arrives before the silicon is tired.
What it is good at
- 504 GB/s makes 8B models feel instant rather than merely usable.
- 200 W and a compact board; comfortable in a normal desktop.
- Strong at image generation, where compute matters more than for text.
What it is not
- 12 GB caps you at roughly 14B at Q4 with a short context, and 8B if you want long documents.
- The next real step up in capability is a 24 GB card, not another 12 GB one.
The thing people get wrong: For language models the 4070 and the 16 GB 4060 Ti trade places depending on what you value: the 4070 is faster on everything that fits, the 4060 Ti fits more. Neither wins outright, which is why the tier above exists.
Buy or rent
A reasonable buy if you already want the card for games or image work and language models are a bonus. A poor buy if models are the reason — the money is better spent on memory.
Retail prices drift, and a figure written today would be wrong within a month, so this page does not carry one. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 200 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- RTX 3090 — 24 GB · 936 GB/s · best fit Qwen3 30B-A3B (MoE) at ~342 tok/s. The value benchmark for local models, and it has been for years.
- RTX 4060 Ti — 16 GB · 288 GB/s · best fit gpt-oss 20B at ~97 tok/s. More memory than its neighbours, and less bandwidth than a card two years older.
- RTX 3060 — 12 GB · 360 GB/s · best fit Qwen3 8B at ~53 tok/s. The cheapest card that still makes local models worth doing.
All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can an RTX 4070 run a 70B model?
Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 12 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.
What is the best model to run on an RTX 4070?
Qwen3 8B. At Q4_K_M and 8k context it needs about 6.7 GB of the 12 GB available and generates an estimated 74 tokens per second, which is faster than you can read. It also fits at Q8, at about 10.7 GB, which is worth taking when it fits.
How many tokens per second does an RTX 4070 generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 504 GB/s and 70% efficiency, this card produces an estimated 76 tokens per second on an 8B model at Q4, and about 41 on the largest model it holds, Qwen3 14B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.
Is 12 GB enough for running LLMs locally?
It runs 6 of the 16 models tracked here at Q4_K_M with 8k of context, up to 14.8B parameters. The honest test is not the model list but the context: Qwen3 14B on this card holds about 11,264 tokens before memory runs out.
RTX 4070 or RTX 3090 for local models?
The RTX 3090 holds more: 24 GB against 12 GB, so it runs 13 of these models to this card's 6. On speed, the RTX 3090 is faster: 936 GB/s against 504 GB/s, and bandwidth is what sets chat speed.
RTX 4070 or RTX 4060 Ti 16 GB for LLMs?
The 4070 is faster on everything that fits in 12 GB, by a wide margin: 504 GB/s against 288 GB/s. The 4060 Ti holds models the 4070 cannot. If your models are 8B or smaller, take the 4070; if you specifically need a 14B model with long context, take the 4060 Ti and accept the pace.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
- Card specification: 12 GB, 504 GB/s, 200 W — manufacturer figures.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.