Claim. An 8B model at Q4_K_M needs 6.8 GB at a short context.
Method. 8 × 0.5 × 1.2 + 2. Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. Same formula as llms.txt.
As of.
Estimate GPU memory for a local model. Updated 2 October 2026. Author: Billy G.R.
Agent stacks add a browser, a tool process, and a longer KV cache on top of the weights. Those extras are not inside the 1.2 factor. The VRAM guide for local agents walks 7B, 14B, 27B, 30B, 32B, and 70B, and labels an extra 1–4 GB scaffold allowance as editorial headroom rather than a vendor specification.
Claim. An 8B model at Q4_K_M needs 6.8 GB at a short context.
Method. 8 × 0.5 × 1.2 + 2. Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. Same formula as llms.txt.
As of.
Claim. A 32B model at Q4_K_M needs 21.2 GB, so a 24 GB GPU fits a short context and a 16 GB GPU does not.
Method. 32 × 0.5 × 1.2 + 2. The 1.2 factor is a short-context allowance, not a 32K KV-cache estimate.
As of.
Claim. A 70B model at Q4_K_M needs 44 GB and does not fit one 32 GB RTX 5090.
Method. 70 × 0.5 × 1.2 + 2. ASUS TUF RTX 5090 ASIN B0DS2X13PH is 32 GB and showed $7,398 on Amazon on this date.
As of.
Methodology — parameter math, quantization bytes, and the source list.
Quick presets
e.g. 7 for a 7B model, 70 for 70B
Estimated VRAM needed
— GB
—
Apple Mac Studio (M5 Max, 36GB)
36 GB unified
—
Apple Mac mini (M5 Pro, 24GB)
24 GB unified
—
Apple Mac mini (M4, 16GB)
16 GB unified
—
Apple Mac mini (M4, 24GB)
24 GB unified
—
Apple Mac mini (M4 Pro, 24GB)
24 GB unified
—
Apple Mac mini (M4 Pro, 48GB)
48 GB unified
—
Apple Mac Studio (M4 Max, 64GB)
64 GB unified
—
Apple Mac Studio (M4 Max, 128GB)
128 GB unified
—
NVIDIA GeForce RTX 4060 (8GB)
8 GB VRAM
—
NVIDIA GeForce RTX 4070 (12GB)
12 GB VRAM
—
NVIDIA GeForce RTX 4090 (24GB)
24 GB VRAM
—
NVIDIA GeForce RTX 3090 (24GB)
24 GB VRAM
—
NVIDIA GeForce RTX 5090 (32GB)
32 GB VRAM
—
Intel Arc B580 (12GB)
12 GB VRAM
—
AMD Radeon RX 7900 XTX (24GB)
24 GB VRAM
—
NVIDIA RTX 4080 (16GB)
16 GB VRAM
—
NVIDIA RTX 4060 Ti (16GB)
16 GB VRAM
—
NVIDIA RTX 5080 (16GB)
16 GB VRAM
—
NVIDIA RTX 5070 (12GB)
12 GB VRAM
—
NVIDIA RTX 5070 Ti (16GB)
16 GB VRAM
—
NVIDIA RTX 4070 Ti Super (16GB)
16 GB VRAM
—
NVIDIA DGX Spark
128 GB VRAM
—
PNY GeForce RTX 5060 Ti (16GB)
16 GB VRAM
—
GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)
128 GB unified
—
Beelink GTR9 Pro (128GB) — caution
128 GB unified
—
Beelink GTR9 Pro (128GB, OOS watchlist)
128 GB unified
—
GPU VRAM (or Apple Silicon unified memory) is the primary bottleneck when running large language models locally. The model weights must fit entirely in VRAM for efficient inference — otherwise the system falls back to slow RAM offloading.
The VRAM requirement scales linearly with parameter count and quantization bit-width:
| Model Size | Q4_K_M | Q5_K_M | Q8 | FP16 | Ollama tag |
|---|---|---|---|---|---|
| 7B | 6.2 GB | 7.3 GB | 10.4 GB | 18.8 GB | ollama run qwen3:8b is the nearby 8B tag (6.8 GB at Q4) |
| 8B | 6.8 GB | 8.0 GB | 11.6 GB | 21.2 GB | qwen3:8b, Llama 3.1 8B |
| 14B | 10.4 GB | 12.5 GB | 18.8 GB | 35.6 GB | ollama run qwen3:14b or phi4:14b |
| 32B | 21.2 GB | 26.0 GB | 40.4 GB | 78.8 GB | ollama run qwen3:32b — needs 24 GB at Q4 |
| 70B | 44.0 GB | 54.5 GB | 86.0 GB | 170 GB | llama3.3:70b — not one 32 GB GPU |
Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. The 1.2 factor covers activations and a short context (about 4K). It is not a 32K or 128K estimate. The context control above adds an approximate KV-cache allowance on top. A worked example: 8 × 0.5 × 1.2 + 2 = 6.8 GB for an 8B model at Q4_K_M. That is the same formula as llms.txt and the methodology page.
For Mixture-of-Experts models (Llama 4 Scout, Mixtral), use the total parameter count in the calculator — all expert weights must be loaded into VRAM even though only a fraction activate per token. For example, Llama 4 Scout is named "17B active" but has ~109B total parameters. At Q4_K_M that is 109 × 0.5 × 1.2 + 2 = 67.4 GB, because every expert weight is resident. See the Llama 4 guide for details.
Use VRAM (GB) = parameters (B) × bytes_per_param × 1.2 + 2 GB for the OS. Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. An 8B model at Q4_K_M is 8 × 0.5 × 1.2 + 2 = 6.8 GB at a short context. Longer context adds KV cache on top. The method is on the methodology page.
7 × 0.5 × 1.2 + 2 = 6.2 GB at Q4_K_M, 7 × 1.0 × 1.2 + 2 = 10.4 GB at Q8, and 7 × 2.0 × 1.2 + 2 = 18.8 GB at FP16. An 8 GB card fits Q4_K_M with a short context. Q8 needs 12 GB. Pull it with ollama pull qwen3:8b when you want that size.
32 × 0.5 × 1.2 + 2 = 21.2 GB. A 24 GB GPU such as an RTX 3090 or RTX 4090 fits it at a short context. Q8 is 32 × 1.2 + 2 = 40.4 GB and does not fit one 24 GB card. The matching Ollama tag is qwen3:32b.
70 × 0.5 × 1.2 + 2 = 44 GB at Q4_K_M. That does not fit a 32 GB RTX 5090. A 64 GB or larger unified-memory Mac, or two 24 GB GPUs, can hold the weights. ollama pull llama3.3:70b is the tag, not a promise that one consumer GPU can run it.
No. The 1.2 factor plus the 2 GB OS reserve covers activations and a short context, about 4K tokens. This calculator adds an extra KV-cache allowance for 16K, 32K, and 128K. Those extras are approximate and depend on layer count and KV heads. The methodology page shows the per-token KV formula.
13 × 0.5 × 1.2 + 2 = 9.8 GB at Q4_K_M. An 8 GB card is short. A 12 GB card such as an Intel Arc B580 or an RTX 4070 fits that Q4_K_M estimate. qwen3:14b is the nearby Ollama tag and needs 10.4 GB.
Pull a size tag, not a VRAM setting. If a 14B Q4_K_M fits, run ollama pull qwen3:14b and then ollama run qwen3:14b. In LM Studio, download the Q4_K_M GGUF of the same size. If ollama ps shows GPU under 100%, the weights spilled to system RAM.
Unified memory is shared with macOS. Treat usable memory as the catalog figure minus several GB for the OS, on top of the formula’s 2 GB reserve. A 24 GB Mac is not the same fit as a 24 GB discrete GPU. See the Apple Silicon guide for the wired-memory caveat.
Add the VRAM calculator to your blog or documentation site. Copy the snippet below — no signup required.
<iframe src="https://llmhardware.io/embed/vram-calculator" width="100%" height="420" style="border:none;border-radius:10px;" loading="lazy" title="LLM VRAM Calculator"></iframe> <p><a href="https://llmhardware.io/guides" target="_blank" rel="noopener">Calculator by LLMHardware.io</a></p>
Keep the credit line. Calculator by LLMHardware.io should stay linked to the guide hub.
How much VRAM do I need?
The same formula walked through for 7B, 14B, 32B, and 70B, plus which GPU tier fits.
Read guide →Best GPU for LLMs
Pick a card by the gigabytes the calculator just asked for, not by the gaming score.
Read guide →Best budget GPU for LLMs
12 GB and 16 GB cards, with the RTX 5070 called out as 12 GB.
Read guide →Best 24 GB VRAM GPU
Where qwen3:32b (21.2 GB at Q4_K_M) actually fits.
Read guide →Ollama cheat sheet
ollama pull qwen3:14b and the other size tags that match these estimates.
Read guide →