LLM VRAM Calculator

Estimate GPU memory for a local model. Updated 2 October 2026. Author: Billy G.R.

Agent stacks add a browser, a tool process, and a longer KV cache on top of the weights. Those extras are not inside the 1.2 factor. The VRAM guide for local agents walks 7B, 14B, 27B, 30B, 32B, and 70B, and labels an extra 1–4 GB scaffold allowance as editorial headroom rather than a vendor specification.

Citeable facts

Claim. An 8B model at Q4_K_M needs 6.8 GB at a short context.

Method. 8 × 0.5 × 1.2 + 2. Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. Same formula as llms.txt.

As of.

Claim. A 32B model at Q4_K_M needs 21.2 GB, so a 24 GB GPU fits a short context and a 16 GB GPU does not.

Method. 32 × 0.5 × 1.2 + 2. The 1.2 factor is a short-context allowance, not a 32K KV-cache estimate.

As of.

Claim. A 70B model at Q4_K_M needs 44 GB and does not fit one 32 GB RTX 5090.

Method. 70 × 0.5 × 1.2 + 2. ASUS TUF RTX 5090 ASIN B0DS2X13PH is 32 GB and showed $7,398 on Amazon on this date.

As of.

Methodology — parameter math, quantization bytes, and the source list.

Quick presets

e.g. 7 for a 7B model, 70 for 70B

Estimated VRAM needed

— GB

—

Compatible Hardware

 

Apple Mac Studio (M5 Max, 36GB)

36 GB unified

—

Apple Mac mini (M5 Pro, 24GB)

24 GB unified

—

Apple Mac mini (M4, 16GB)

16 GB unified

—

Apple Mac mini (M4, 24GB)

24 GB unified

—

Apple Mac mini (M4 Pro, 24GB)

24 GB unified

—

Apple Mac mini (M4 Pro, 48GB)

48 GB unified

—

Apple Mac Studio (M4 Max, 64GB)

64 GB unified

—

Apple Mac Studio (M4 Max, 128GB)

128 GB unified

—

NVIDIA GeForce RTX 4060 (8GB)

8 GB VRAM

—

NVIDIA GeForce RTX 4070 (12GB)

12 GB VRAM

—

NVIDIA GeForce RTX 4090 (24GB)

24 GB VRAM

—

NVIDIA GeForce RTX 3090 (24GB)

24 GB VRAM

—

NVIDIA GeForce RTX 5090 (32GB)

32 GB VRAM

—

Intel Arc B580 (12GB)

12 GB VRAM

—

AMD Radeon RX 7900 XTX (24GB)

24 GB VRAM

—

NVIDIA RTX 4080 (16GB)

16 GB VRAM

—

NVIDIA RTX 4060 Ti (16GB)

16 GB VRAM

—

NVIDIA RTX 5080 (16GB)

16 GB VRAM

—

NVIDIA RTX 5070 (12GB)

12 GB VRAM

—

NVIDIA RTX 5070 Ti (16GB)

16 GB VRAM

—

NVIDIA RTX 4070 Ti Super (16GB)

16 GB VRAM

—

NVIDIA DGX Spark

128 GB VRAM

—

PNY GeForce RTX 5060 Ti (16GB)

16 GB VRAM

—

GMKtec EVO-X2 (Ryzen AI Max+ 395, 128GB)

128 GB unified

—

Beelink GTR9 Pro (128GB) — caution

128 GB unified

—

Beelink GTR9 Pro (128GB, OOS watchlist)

128 GB unified

—

How Much VRAM Do You Need for LLMs?

GPU VRAM (or Apple Silicon unified memory) is the primary bottleneck when running large language models locally. The model weights must fit entirely in VRAM for efficient inference — otherwise the system falls back to slow RAM offloading.

The VRAM requirement scales linearly with parameter count and quantization bit-width:

VRAM Requirements by Model Size

Model SizeQ4_K_MQ5_K_MQ8FP16Ollama tag
7B 6.2 GB 7.3 GB 10.4 GB 18.8 GB ollama run qwen3:8b is the nearby 8B tag (6.8 GB at Q4)
8B 6.8 GB 8.0 GB 11.6 GB 21.2 GB qwen3:8b, Llama 3.1 8B
14B 10.4 GB 12.5 GB 18.8 GB 35.6 GB ollama run qwen3:14b or phi4:14b
32B 21.2 GB 26.0 GB 40.4 GB 78.8 GB ollama run qwen3:32b — needs 24 GB at Q4
70B 44.0 GB 54.5 GB 86.0 GB 170 GB llama3.3:70b — not one 32 GB GPU

The Formula

VRAM (GB) = parameters(B) × bytes_per_param × 1.2 + 2 GB OS

Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. The 1.2 factor covers activations and a short context (about 4K). It is not a 32K or 128K estimate. The context control above adds an approximate KV-cache allowance on top. A worked example: 8 × 0.5 × 1.2 + 2 = 6.8 GB for an 8B model at Q4_K_M. That is the same formula as llms.txt and the methodology page.

For Mixture-of-Experts models (Llama 4 Scout, Mixtral), use the total parameter count in the calculator — all expert weights must be loaded into VRAM even though only a fraction activate per token. For example, Llama 4 Scout is named "17B active" but has ~109B total parameters. At Q4_K_M that is 109 × 0.5 × 1.2 + 2 = 67.4 GB, because every expert weight is resident. See the Llama 4 guide for details.

Frequently Asked Questions

How do I calculate VRAM for an LLM?

Use VRAM (GB) = parameters (B) × bytes_per_param × 1.2 + 2 GB for the OS. Bytes per parameter: FP16 = 2, Q8 = 1, Q5_K_M = 0.625, Q4_K_M = 0.5. An 8B model at Q4_K_M is 8 × 0.5 × 1.2 + 2 = 6.8 GB at a short context. Longer context adds KV cache on top. The method is on the methodology page.

How much VRAM does a 7B model need?

7 × 0.5 × 1.2 + 2 = 6.2 GB at Q4_K_M, 7 × 1.0 × 1.2 + 2 = 10.4 GB at Q8, and 7 × 2.0 × 1.2 + 2 = 18.8 GB at FP16. An 8 GB card fits Q4_K_M with a short context. Q8 needs 12 GB. Pull it with ollama pull qwen3:8b when you want that size.

How much VRAM does a 32B model need at Q4_K_M?

32 × 0.5 × 1.2 + 2 = 21.2 GB. A 24 GB GPU such as an RTX 3090 or RTX 4090 fits it at a short context. Q8 is 32 × 1.2 + 2 = 40.4 GB and does not fit one 24 GB card. The matching Ollama tag is qwen3:32b.

How much VRAM does a 70B model need?

70 × 0.5 × 1.2 + 2 = 44 GB at Q4_K_M. That does not fit a 32 GB RTX 5090. A 64 GB or larger unified-memory Mac, or two 24 GB GPUs, can hold the weights. ollama pull llama3.3:70b is the tag, not a promise that one consumer GPU can run it.

Does the 1.2 factor include a long context window?

No. The 1.2 factor plus the 2 GB OS reserve covers activations and a short context, about 4K tokens. This calculator adds an extra KV-cache allowance for 16K, 32K, and 128K. Those extras are approximate and depend on layer count and KV heads. The methodology page shows the per-token KV formula.

Why does a 13B model miss an 8 GB GPU?

13 × 0.5 × 1.2 + 2 = 9.8 GB at Q4_K_M. An 8 GB card is short. A 12 GB card such as an Intel Arc B580 or an RTX 4070 fits that Q4_K_M estimate. qwen3:14b is the nearby Ollama tag and needs 10.4 GB.

Which Ollama or LM Studio file matches a calculator result?

Pull a size tag, not a VRAM setting. If a 14B Q4_K_M fits, run ollama pull qwen3:14b and then ollama run qwen3:14b. In LM Studio, download the Q4_K_M GGUF of the same size. If ollama ps shows GPU under 100%, the weights spilled to system RAM.

How should I read Apple unified memory in this calculator?

Unified memory is shared with macOS. Treat usable memory as the catalog figure minus several GB for the OS, on top of the formula’s 2 GB reserve. A 24 GB Mac is not the same fit as a 24 GB discrete GPU. See the Apple Silicon guide for the wired-memory caveat.

Embed This Calculator

Add the VRAM calculator to your blog or documentation site. Copy the snippet below — no signup required.

<iframe src="https://llmhardware.io/embed/vram-calculator" width="100%" height="420" style="border:none;border-radius:10px;" loading="lazy" title="LLM VRAM Calculator"></iframe>
<p><a href="https://llmhardware.io/guides" target="_blank" rel="noopener">Calculator by LLMHardware.io</a></p>

Keep the credit line. Calculator by LLMHardware.io should stay linked to the guide hub.

Related Guides