12 GB VRAM 360 GB/s 170 W 2021 Mostly second-hand
RTX 3060 12 GB addresses 12 GB at 360 GB/s. Of the 16 open models tracked on this site, it runs 6 at Q4_K_M with 8k of context. The largest is Qwen3 14B, needing about 10.8 GB and generating an estimated 29 tokens per second — comfortable for chat.
The cheapest card that still makes local models worth doing.
What an RTX 3060 runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 11.4 GB to spend once the 5% safety margin comes off its 12 GB, and it reads that memory at 360 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 136 | fits |
| Llama 3.1 8B | 8B | 6.4 GB | 54 | fits |
| Qwen3 8B | 8.2B | 6.7 GB | 53 | fits |
| Gemma 3 12B | 12.2B | 11.1 GB | 36 | tight fit |
| Qwen3 14B | 14.8B | 10.8 GB | 29 | tight fit |
| Phi-4 14B | 14.7B | 11 GB | 30 | tight fit |
| gpt-oss 20B | 21B · 3.6B active | 13.6 GB | – | no |
| Mistral Small 3.1 24B | 24B | 16.3 GB | – | no |
| Gemma 3 27B | 27.4B | 21.2 GB | – | no |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | – | no |
| Qwen3 32B | 32.8B | 22.4 GB | – | no |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | – | no |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | – | no |
| Llama 3.3 70B | 70.6B | 45.8 GB | – | no |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | – | no |
| gpt-oss 120B | 117B · 5.1B active | 71.7 GB | – | no |
On this card that means anything at or under 11.4 GB counts as fitting, and anything above 10.2 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 12 GB can push it over.
The model to actually run on it
Qwen3 8B is the best use of this card: 6.7 GB of the 12 GB available, an estimated 53 tokens per second, comfortable for chat, and 39,936 tokens of context still available.
Qwen3 14B is the largest model the card will hold, at 10.8 GB, but it is the wrong daily driver: filling 90% of memory with weights leaves room for only 11,264 tokens of conversation, and it generates at 29 tokens per second against 53. More parameters are not worth a context window that runs out mid-document.
It also fits at Q8_0, in about 10.7 GB, which is worth taking whenever the memory allows: Q8 is near-lossless where Q4 costs a little accuracy.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 12 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| Qwen3 14B | 11,264 tokens | 23,552 tokens |
| Phi-4 14B | 9,216 tokens | 19,456 tokens |
| Gemma 3 12B | 8,192 tokens | 17,408 tokens |
| Qwen3 8B | 39,936 tokens | 79,872 tokens |
Quantising the KV cache to Q8 roughly doubles what you can hold — Qwen3 14B on this card goes from 11,264 to 23,552 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.
Where this card stops
The first model out of reach is gpt-oss 20B: about 13.6 GB at Q4_K_M and 8k context, against 12 GB of usable memory. Dropping to Q3_K_M would need about 11.2 GB, which fits, though Q3 loses enough quality that a smaller model at Q4 is usually the better trade. Offloading the remainder to system RAM works and is ten to fifty times slower; it is a way to see a model run, not a way to use one.
The honest take
Twelve gigabytes on a budget card was an accident of Nvidia's memory-bus maths, and it turned the 3060 into the default first GPU for local models. It runs 7B and 8B models properly, handles 12B to 14B at low context, and stops there. Bandwidth is modest, so answers arrive at reading speed rather than instantly, which is fine for chat and irritating for long code output.
What it is good at
- 12 GB is enough for an 8B model at Q4 with a long context, or a 14B model at Q4 with a short one.
- A 170 W board runs on almost any existing power supply, so it is often a drop-in upgrade rather than a new build.
- Widely available second-hand, which is where the value is.
What it is not
- At 360 GB/s it is the slowest card here that anyone should still consider. Larger models feel it.
- No headroom for image or video generation at high resolution, where VRAM disappears quickly.
- Anything above 14B is out of reach at usable quality.
The thing people get wrong: Buy the 12 GB version deliberately. An 8 GB RTX 3060 exists, and for language models the memory is the entire point — the 8 GB card cannot hold a useful model plus its context.
Buy or rent
This is the one card on this list where buying is usually the right answer, because the outlay is small enough that break-even arrives quickly even at a few hours a month. Put your real second-hand price into the build-vs-rent calculator and see.
Second-hand pricing is the whole argument for this card, and it moves week to week, so printing a number here would be worse than printing none. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator — it already knows this card draws 170 W under load. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- RTX 4070 — 12 GB · 504 GB/s · best fit Qwen3 8B at ~74 tok/s. Fast for its class, and capped by 12 GB.
- RTX 4060 Ti — 16 GB · 288 GB/s · best fit gpt-oss 20B at ~97 tok/s. More memory than its neighbours, and less bandwidth than a card two years older.
- RTX 3090 — 24 GB · 936 GB/s · best fit Qwen3 30B-A3B (MoE) at ~342 tok/s. The value benchmark for local models, and it has been for years.
All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can an RTX 3060 run a 70B model?
Not at Q4_K_M in one card. Llama 3.3 70B needs about 45.8 GB at 8k context and this card can address 12 GB. Your options are a smaller model, a harsher quantisation with very little context, splitting across two cards, or renting a larger one for the hours you need it.
What is the best model to run on an RTX 3060?
Qwen3 8B. At Q4_K_M and 8k context it needs about 6.7 GB of the 12 GB available and generates an estimated 53 tokens per second, which is comfortable for chat. It also fits at Q8, at about 10.7 GB, which is worth taking when it fits.
How many tokens per second does an RTX 3060 generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 360 GB/s and 70% efficiency, this card produces an estimated 54 tokens per second on an 8B model at Q4, and about 29 on the largest model it holds, Qwen3 14B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.
Is 12 GB enough for running LLMs locally?
It runs 6 of the 16 models tracked here at Q4_K_M with 8k of context, up to 14.8B parameters. The honest test is not the model list but the context: Qwen3 14B on this card holds about 11,264 tokens before memory runs out.
RTX 3060 or RTX 4070 for local models?
Both address about the same memory, so they run the same models. On speed, the RTX 4070 is faster: 504 GB/s against 360 GB/s, and bandwidth is what sets chat speed.
Is the RTX 3060 8 GB good enough for AI?
No, and the difference is not marginal. An 8 GB card cannot hold a useful model and its context at the same time, so you end up offloading to system memory, which is ten to fifty times slower. The 12 GB version is the one worth buying; check the listing carefully, because both were sold under the same name.
Can an RTX 3060 run Stable Diffusion as well as language models?
Yes, and image generation is where its 12 GB is most comfortable, because image models are far smaller than language models. Video generation is where it runs out — that is where VRAM disappears fastest.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
- Card specification: 12 GB, 360 GB/s, 170 W — manufacturer figures.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.