80 GB VRAM 2,039 GB/s 2020 Rented, not bought
A100 80 GB addresses 80 GB at 2,039 GB/s. Of the 16 open models tracked on this site, it runs 16 at Q4_K_M with 8k of context. The largest is gpt-oss 120B, needing about 71.7 GB and generating an estimated 483 tokens per second — faster than you can read.
The old datacenter workhorse, still fast where it counts, and cheap to rent.
What an A100 runs, model by model
Every model at Q4_K_M with 8k of context, the setting most people actually use. This card has 76 GB to spend once the 5% safety margin comes off its 80 GB, and it reads that memory at 2,039 GB/s. Those two numbers decide the last two columns: the first says what loads, the second says how fast it answers. Both are estimates from stated formulas, and the VRAM calculator shows the working.
| Model | Size | Needs | Est. tokens/s | On this card |
|---|---|---|---|---|
| Llama 3.2 3B | 3.2B | 3.4 GB | 769 | fits |
| Llama 3.1 8B | 8B | 6.4 GB | 308 | fits |
| Qwen3 8B | 8.2B | 6.7 GB | 300 | fits |
| Gemma 3 12B | 12.2B | 11.1 GB | 202 | fits |
| Qwen3 14B | 14.8B | 10.8 GB | 166 | fits |
| Phi-4 14B | 14.7B | 11 GB | 167 | fits |
| gpt-oss 20B | 21B · 3.6B active | 13.6 GB | 684 | fits |
| Mistral Small 3.1 24B | 24B | 16.3 GB | 103 | fits |
| Gemma 3 27B | 27.4B | 21.2 GB | 90 | fits |
| Qwen3 30B-A3B (MoE) | 30.5B · 3.3B active | 19.7 GB | 746 | fits |
| Qwen3 32B | 32.8B | 22.4 GB | 75 | fits |
| Qwen2.5 Coder 32B | 32.8B | 22.4 GB | 75 | fits |
| DeepSeek-R1 Distill Qwen 32B | 32.8B | 22.4 GB | 75 | fits |
| Llama 3.3 70B | 70.6B | 45.8 GB | 35 | fits |
| DeepSeek-R1 Distill Llama 70B | 70.6B | 45.8 GB | 35 | fits |
| gpt-oss 120B | 117B · 5.1B active | 71.7 GB | 483 | tight fit |
On this card that means anything at or under 76 GB counts as fitting, and anything above 68 GB is marked tight: it loads, but there is little left for the conversation, and a desktop on the same 80 GB can push it over.
The model to actually run on it
Llama 3.3 70B is the best use of this card: 45.8 GB of the 80 GB available, an estimated 35 tokens per second, comfortable for chat, and 100,352 tokens of context still available.
gpt-oss 120B is the largest model the card will hold, at 71.7 GB, but it is the wrong daily driver: filling 90% of memory with weights leaves room for only 66,560 tokens of conversation, and it generates at 483 tokens per second against 35. More parameters are not worth a context window that runs out mid-document.
If quality matters more than parameter count, Qwen3 32B fits at Q8_0 in about 38.8 GB. Q8 is near-lossless; a smaller model at Q8 often beats a larger one squeezed into Q3.
How much context actually fits
Model size is the question people ask; context is the one that bites. The KV cache grows linearly with the conversation, so a model that loads comfortably can still run out of memory halfway through a long document. On 80 GB, the context you get is whatever the weights leave behind — which is why two cards that run the same model can feel completely different in use. These are the largest contexts this card holds, at Q4_K_M weights.
| Model | Max context, FP16 KV | With Q8 KV cache |
|---|---|---|
| gpt-oss 120B | 66,560 tokens | 131,072 tokens |
| Llama 3.3 70B | 100,352 tokens | 131,072 tokens |
| DeepSeek-R1 Distill Llama 70B | 100,352 tokens | 131,072 tokens |
| Qwen3 32B | 131,072 tokens | 131,072 tokens |
Quantising the KV cache to Q8 roughly doubles what you can hold — gpt-oss 120B on this card goes from 66,560 to 131,072 tokens — at a quality cost most people never notice. Capped at 128k tokens here; a model's own architecture may stop lower, and sliding-window models such as Gemma 3 use less KV than the formula assumes.
Where this card stops
Nothing in this list is out of reach. Every model tracked here fits at Q4_K_M with 8k of context, so the ceiling you will meet is context length and batch size, not parameter count.
The honest take
The A100 is two generations old and still has 2,039 GB/s of HBM2e, which is more than double a 4090. For running a large model in one conversation, it remains excellent, and because newer hardware has moved the market on, it is usually one of the cheapest ways to rent real capacity. It is not the card to pick for training at scale any more.
What it is good at
- 80 GB and 2,039 GB/s: 70B models at Q4 run quickly with long context.
- Mature software support; nothing about it surprises a runtime.
- Generally the cheapest 80 GB rental class available.
What it is not
- Ampere-generation compute; slower than an H100 on prompt processing and on training.
- No FP8 support, which newer runtimes increasingly assume.
- Rental only, in practice. This is not a card you put in a desktop.
The thing people get wrong: Its memory bandwidth still beats every consumer card except the 5090 and RTX PRO 6000, and it launched in 2020. Bandwidth ages more slowly than compute, which is why old datacenter cards stay useful for chat long after they stop being bought for training.
Buy or rent
Rent. The purchase question does not arise outside a rack, and if it does, the answer involves cooling and power that a home does not have.
There is no purchase price to reason about, only an hourly one, and it differs by provider and by week. Put the figure you are actually looking at, with your electricity rate and your real monthly hours, into the build vs rent calculator. It returns the month owning becomes cheaper, or tells you it never does. Hours per month decides it far more often than the price does.
Nodegrove is the rented side of that comparison: a workspace that stays saved, with a GPU attached only while you are using it. It is the right answer when your hours are low or your needs change; buying is the right answer when they are high and stable.
Compared with the alternatives
- H100 SXM — 80 GB · 3,352 GB/s · best fit Llama 3.3 70B at ~57 tok/s. The fastest memory here by a wide margin, and almost always more than one person needs.
- L40S — 48 GB · 864 GB/s · best fit Qwen3 32B at ~32 tok/s. The 48 GB card you meet as a rental line item, not as a purchase.
- RTX PRO 6000 Blackwell — 96 GB · 1,792 GB/s · best fit gpt-oss 120B at ~424 tok/s. Ninety-six gigabytes and 5090-class bandwidth on one card.
All 13 cards are compared side by side on the GPU index, and the same numbers from the model's point of view are on each model fit table. To check one specific pairing rather than read a table, use the can-I-run-it checker.
Questions
Can an A100 run a 70B model?
Yes. Llama 3.3 70B at Q4_K_M with 8k of context needs about 45.8 GB, inside the 80 GB this card can address, and it generates an estimated 35 tokens per second — comfortable for chat.
What is the best model to run on an A100?
Llama 3.3 70B. At Q4_K_M and 8k context it needs about 45.8 GB of the 80 GB available and generates an estimated 35 tokens per second, which is comfortable for chat. If quality matters more than size, Qwen3 32B fits at Q8 in about 38.8 GB.
How many tokens per second does an A100 generate?
It depends entirely on the size of the model, because generating a token means reading every active weight out of memory. At 2,039 GB/s and 70% efficiency, this card produces an estimated 308 tokens per second on an 8B model at Q4, and about 483 on the largest model it holds, gpt-oss 120B. Reading speed is roughly 5 to 8 words per second, so anything above 25 feels immediate.
Is 80 GB enough for running LLMs locally?
It runs 16 of the 16 models tracked here at Q4_K_M with 8k of context, up to 117B parameters. The honest test is not the model list but the context: gpt-oss 120B on this card holds about 66,560 tokens before memory runs out.
A100 or H100 SXM for local models?
Both address about the same memory, so they run the same models. On speed, the H100 SXM is faster: 3,352 GB/s against 2,039 GB/s, and bandwidth is what sets chat speed.
Is an A100 still worth renting in 2026?
For single-user inference, yes, and often it is the sensible choice. Its 2,039 GB/s of memory bandwidth still beats every consumer card except the 5090 and the RTX PRO 6000, and because newer hardware has moved the market on, it is usually the cheapest way to rent 80 GB. What it lacks is FP8 and the training throughput of an H100.
A100 40 GB or 80 GB?
The 80 GB version is the one that matters for language models: it holds a 70B model at Q4 with a long context, where the 40 GB card is limited to roughly the same tier as a 48 GB workstation card. It also has higher memory bandwidth.
- Memory: weights (parameters × bytes per parameter) + KV cache (2 × layers × KV heads × head dim × context × bytes) + 0.5 GB runtime + 4% of weights. Full derivation in the VRAM calculator.
- Speed: 70% × bandwidth ÷ active weight bytes, single stream, no batching, prompt processing excluded. See the tokens-per-second estimator.
- Card specification: 80 GB, 2,039 GB/s — manufacturer figures.
- Model architecture values come from each model's published config; the self-hosted LLM guide explains what each one changes.