How Much VRAM Do I Need for Local LLMs?

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

Updated 2 October 2026 — formula aligned with the methodology page and the VRAM calculator

Citeable facts

Claim. An 8B model at Q4_K_M needs 6.8 GB.

Method. VRAM (GB) = params(B) × bytes_per_param × 1.2 + 2. Q4_K_M = 0.5, so 8 × 0.5 × 1.2 + 2 = 6.8.

As of.

Claim. A 70B model at Q4_K_M needs 44 GB and does not fit a 24 GB or 32 GB consumer GPU.

Method. 70 × 0.5 × 1.2 + 2 = 44.

As of.

Methodology — parameter math, quantization bytes, and the source list.

Quick rule: VRAM (GB) = parameters (billions) × bytes per parameter × 1.2 + 2. Q4_K_M is 0.5, Q5_K_M is 0.625, Q8 is 1.0, FP16 is 2.0. A 7B model at Q4_K_M is 6.2 GB. A 70B model at Q4_K_M is 44 GB. The same math is on the VRAM calculator.

VRAM by Model Size and Quantization

Model Size Q4_K_M Q8_0 FP16 Minimum GPU
7B 6.2 GB 10.4 GB 18.8 GB 8 GB GPU for Q4 only. Q8 needs 12 GB.
8B (qwen3:8b) 6.8 GB 11.6 GB 21.2 GB RTX 4060 8 GB at Q4_K_M
13B 9.8 GB 17.6 GB 33.2 GB 12 GB for Q4. Q8 misses 16 GB.
14B (qwen3:14b) 10.4 GB 18.8 GB 35.6 GB 12 GB or 16 GB at Q4. Not Q8.
32B (qwen3:32b) 21.2 GB 40.4 GB 78.8 GB 24 GB at Q4. 16 GB misses it.
70B (llama3.3:70b) 44 GB 86 GB 170 GB Two 24 GB cards or 64 GB unified memory
405B 245 GB 488 GB 974 GB Multi-GPU workstation

Each cell is params × bytes × 1.2 + 2, rounded to one decimal. The 1.2 factor already includes a short context. Longer windows are the add-on in the next section, and on the VRAM calculator.

Ollama and LM Studio

Pull the tag that matches the VRAM you actually have. In LM Studio, download the Q4_K_M GGUF of the same size.

After the model loads, run ollama ps. GPU% under 100 means part of the model is in system RAM.

The Formula

VRAM (GB) = params(B) × bytes_per_param × 1.2 + 2

bytes_per_param:
  Q4_K_M = 0.5
  Q5_K_M = 0.625
  Q8     = 1.0
  FP16    = 2.0

The 1.2 covers a short context. The +2 GB is the OS.

Example: 13B at Q4_K_M = 13 × 0.5 × 1.2 + 2 = 9.8 GB. The old weight-only sum (13 × 0.5 + 2 = 8.5 GB) left out the 1.2 factor.

Context on top of the formula

The published formula already includes a short context inside the 1.2 factor. The VRAM calculator adds the same extra for a longer window on every model size: 0 GB at 4K, 2 GB at 16K, 6 GB at 32K, and 22 GB at 128K. That add-on is approximate KV cache beyond the short window. It is not a second copy of the weights.

Context 7B 13B–14B 32B–70B
4K+0 GB+0 GB+0 GB
16K+2 GB+2 GB+2 GB
32K+6 GB+6 GB+6 GB
128K+22 GB+22 GB+22 GB

A 32B Q4_K_M at 4K is 21.2 GB. The same model at 32K is 21.2 + 6 = 27.2 GB, which misses a 24 GB card. Check the context control on the calculator before you buy for long documents.

Q4 vs Q8: Is Quality Affected?

Q4_K_M is the sweet spot for most users. It cuts VRAM roughly in half compared to Q8 while retaining 95-98% of the model's quality. The difference is imperceptible in normal use. Q8 is only worth it if you need maximum accuracy for tasks like math or coding benchmarks.

Format VRAM vs FP16 Quality loss Use when
Q2_K~88% lessNoticeableDesperate for VRAM
Q4_K_M~75% lessMinimalMost users (recommended)
Q5_K_M~69% lessNegligibleYou have headroom
Q8_0~50% lessNoneAccuracy-critical tasks
FP16BaselineNoneFine-tuning / training

What Can I Run on My GPU?

GPU / Memory Best model at Q4 Notes
8 GB VRAM7B Q4 (6.2) or qwen3:8b (6.8)Q8 of an 8B model is 11.6 GB and does not fit. 13B Q4 is 9.8 GB and does not fit.
12 GB VRAMqwen3:14b Q4 (10.4)RTX 5070 is 12 GB, $937.39 on 2 October 2026 (B0DS6S98ZF). 14B Q8 is 18.8 GB.
16 GB VRAM14B Q4 (10.4) and 20B Q4 (14.0)14B Q8 is 18.8 GB. 27B Q4 is 18.2 GB. 32B Q4 is 21.2 GB. None of those fit.
24 GB VRAMqwen3:32b Q4 (21.2) and 34B Q4 (22.4)70B Q4 is 44 GB and does not fit. See the 24 GB GPU guide.
32 GB VRAM32B Q5 (26.0) and 34B Q4 (22.4)RTX 5090 OC was $7,398 (B0DS2X13PH). 32B Q8 is 40.4 GB. 70B Q4 is 44 GB.
48 GB or 64 GB unified70B Q4 (44 GB)Two 24 GB cards, or a 64 GB unified-memory Mac. 405B Q4 is 245 GB.

Recommended GPUs by VRAM

RTX 4060 8 GB
MSI Ventus 2X 8G OC (B0C8BPW1SP). Listing extract $609.99 on 2 October 2026. Fits qwen3:8b at Q4_K_M (6.8 GB). Does not fit 13B Q4 (9.8 GB).
View on Amazon
RTX 4060 Ti 16 GB
MSI Ventus 2X 16G OC (B0CBK7BRL9). Listing extract $629.99 on 2 October 2026. Fits qwen3:14b Q4 (10.4 GB) and 20B Q4 (14.0 GB). 14B Q8 is 18.8 GB and does not fit.
View on Amazon
RTX 4090 24 GB
MSI Gaming X Trio 24G (B0BG94PS2F). Catalog price $1,599 from 4 May 2026, not rechecked this pass. Fits qwen3:32b Q4 (21.2 GB). 70B Q4 is 44 GB and does not fit.
View on Amazon

Frequently Asked Questions

How much VRAM do I need for a 7B LLM?

A 7B model at Q4_K_M is 7 × 0.5 × 1.2 + 2 = 6.2 GB. Q8 is 10.4 GB and FP16 is 18.8 GB. An 8 GB GPU fits Q4_K_M at a short context. Q8 needs 12 GB. The nearby Ollama tag is ollama run qwen3:8b (8B, 6.8 GB at Q4_K_M).

How much VRAM do I need for a 13B model?

A 13B model at Q4_K_M is 13 × 0.5 × 1.2 + 2 = 9.8 GB. Q8 is 17.6 GB and does not fit a 16 GB card. A 12 GB GPU fits Q4_K_M at a short context. qwen3:14b is the practical Ollama tag and needs 10.4 GB.

How much VRAM do I need for a 70B model?

A 70B model at Q4_K_M is 70 × 0.5 × 1.2 + 2 = 44 GB. That does not fit a 24 GB or 32 GB consumer GPU. Two 24 GB cards, or a 64 GB unified-memory Mac, can hold it. CPU offload runs, and tokens per second collapse.

Does VRAM include context window memory?

The published formula already includes a short context inside the 1.2 factor, plus 2 GB for the OS. A 32K or 128K window is extra KV cache on top of that, and it depends on layer count. Use the VRAM calculator context control for a rough add-on. It is not inside the weight-only table.

Why did RTX 50-series prices explode if VRAM did not?

Tokens per second, once a model fits, follow memory bandwidth. On 2 October 2026 an ASUS TUF RTX 5070 12 GB (ASIN B0DS6S98ZF) was $937.39, an MSI RTX 5070 Ti 16 GB (B0DV9GMDLR) was $1,490, an ASUS TUF RTX 5080 16 GB (B0DQSMMCSH) was $1,809.86, and an ASUS TUF RTX 5090 32 GB OC (B0DS2X13PH) was $7,398. A 27B Q4_K_M is 18.2 GB, so the 16 GB cards miss it. A 24 GB card still wins on fit.

What happens if my model does not fit in VRAM?

The model falls back to CPU offloading, which is 10-50x slower. A 13B Q4_K_M is 9.8 GB, so an RTX 4060 8 GB is short by about 2 GB and tokens per second collapse. Fitting the entire model in VRAM is the single biggest factor for inference speed.

Related Guides

Sources & methodology

VRAM and tokens-per-second figures on this page are synthesised from open community benchmarks. The sitewide formula and the full source list are on the methodology page. For this guide specifically I leaned on:

Spot a number that does not match the linked source? Contact us and I will update the guide.