Claim. An 8B model at Q4_K_M needs 6.8 GB.
Method. VRAM (GB) = params(B) × bytes_per_param × 1.2 + 2. Q4_K_M = 0.5, so 8 × 0.5 × 1.2 + 2 = 6.8.
As of.
Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Updated 2 October 2026 — formula aligned with the methodology page and the VRAM calculator
Claim. An 8B model at Q4_K_M needs 6.8 GB.
Method. VRAM (GB) = params(B) × bytes_per_param × 1.2 + 2. Q4_K_M = 0.5, so 8 × 0.5 × 1.2 + 2 = 6.8.
As of.
Claim. A 70B model at Q4_K_M needs 44 GB and does not fit a 24 GB or 32 GB consumer GPU.
Method. 70 × 0.5 × 1.2 + 2 = 44.
As of.
Methodology — parameter math, quantization bytes, and the source list.
| Model Size | Q4_K_M | Q8_0 | FP16 | Minimum GPU |
|---|---|---|---|---|
| 7B | 6.2 GB | 10.4 GB | 18.8 GB | 8 GB GPU for Q4 only. Q8 needs 12 GB. |
8B (qwen3:8b) | 6.8 GB | 11.6 GB | 21.2 GB | RTX 4060 8 GB at Q4_K_M |
| 13B | 9.8 GB | 17.6 GB | 33.2 GB | 12 GB for Q4. Q8 misses 16 GB. |
14B (qwen3:14b) | 10.4 GB | 18.8 GB | 35.6 GB | 12 GB or 16 GB at Q4. Not Q8. |
32B (qwen3:32b) | 21.2 GB | 40.4 GB | 78.8 GB | 24 GB at Q4. 16 GB misses it. |
70B (llama3.3:70b) | 44 GB | 86 GB | 170 GB | Two 24 GB cards or 64 GB unified memory |
| 405B | 245 GB | 488 GB | 974 GB | Multi-GPU workstation |
Each cell is params × bytes × 1.2 + 2, rounded to one decimal. The 1.2 factor already includes a short context. Longer windows are the add-on in the next section, and on the VRAM calculator.
Pull the tag that matches the VRAM you actually have. In LM Studio, download the Q4_K_M GGUF of the same size.
ollama pull qwen3:8b then ollama run qwen3:8b (6.8 GB). Q8 of the same model is 11.6 GB and spills.ollama run qwen3:14b (10.4 GB). Do not load the Q8 tag (18.8 GB) or qwen3:32b (21.2 GB).ollama run qwen3:32b (21.2 GB). There is no current qwen3:34b tag.After the model loads, run ollama ps. GPU% under 100 means part of the model is in system RAM.
Example: 13B at Q4_K_M = 13 × 0.5 × 1.2 + 2 = 9.8 GB. The old weight-only sum (13 × 0.5 + 2 = 8.5 GB) left out the 1.2 factor.
The published formula already includes a short context inside the 1.2 factor. The VRAM calculator adds the same extra for a longer window on every model size: 0 GB at 4K, 2 GB at 16K, 6 GB at 32K, and 22 GB at 128K. That add-on is approximate KV cache beyond the short window. It is not a second copy of the weights.
| Context | 7B | 13B–14B | 32B–70B |
|---|---|---|---|
| 4K | +0 GB | +0 GB | +0 GB |
| 16K | +2 GB | +2 GB | +2 GB |
| 32K | +6 GB | +6 GB | +6 GB |
| 128K | +22 GB | +22 GB | +22 GB |
A 32B Q4_K_M at 4K is 21.2 GB. The same model at 32K is 21.2 + 6 = 27.2 GB, which misses a 24 GB card. Check the context control on the calculator before you buy for long documents.
Q4_K_M is the sweet spot for most users. It cuts VRAM roughly in half compared to Q8 while retaining 95-98% of the model's quality. The difference is imperceptible in normal use. Q8 is only worth it if you need maximum accuracy for tasks like math or coding benchmarks.
| Format | VRAM vs FP16 | Quality loss | Use when |
|---|---|---|---|
| Q2_K | ~88% less | Noticeable | Desperate for VRAM |
| Q4_K_M | ~75% less | Minimal | Most users (recommended) |
| Q5_K_M | ~69% less | Negligible | You have headroom |
| Q8_0 | ~50% less | None | Accuracy-critical tasks |
| FP16 | Baseline | None | Fine-tuning / training |
| GPU / Memory | Best model at Q4 | Notes |
|---|---|---|
| 8 GB VRAM | 7B Q4 (6.2) or qwen3:8b (6.8) | Q8 of an 8B model is 11.6 GB and does not fit. 13B Q4 is 9.8 GB and does not fit. |
| 12 GB VRAM | qwen3:14b Q4 (10.4) | RTX 5070 is 12 GB, $937.39 on 2 October 2026 (B0DS6S98ZF). 14B Q8 is 18.8 GB. |
| 16 GB VRAM | 14B Q4 (10.4) and 20B Q4 (14.0) | 14B Q8 is 18.8 GB. 27B Q4 is 18.2 GB. 32B Q4 is 21.2 GB. None of those fit. |
| 24 GB VRAM | qwen3:32b Q4 (21.2) and 34B Q4 (22.4) | 70B Q4 is 44 GB and does not fit. See the 24 GB GPU guide. |
| 32 GB VRAM | 32B Q5 (26.0) and 34B Q4 (22.4) | RTX 5090 OC was $7,398 (B0DS2X13PH). 32B Q8 is 40.4 GB. 70B Q4 is 44 GB. |
| 48 GB or 64 GB unified | 70B Q4 (44 GB) | Two 24 GB cards, or a 64 GB unified-memory Mac. 405B Q4 is 245 GB. |
qwen3:8b at Q4_K_M (6.8 GB). Does not fit 13B Q4 (9.8 GB).qwen3:14b Q4 (10.4 GB) and 20B Q4 (14.0 GB). 14B Q8 is 18.8 GB and does not fit.qwen3:32b Q4 (21.2 GB). 70B Q4 is 44 GB and does not fit.A 7B model at Q4_K_M is 7 × 0.5 × 1.2 + 2 = 6.2 GB. Q8 is 10.4 GB and FP16 is 18.8 GB. An 8 GB GPU fits Q4_K_M at a short context. Q8 needs 12 GB. The nearby Ollama tag is ollama run qwen3:8b (8B, 6.8 GB at Q4_K_M).
A 13B model at Q4_K_M is 13 × 0.5 × 1.2 + 2 = 9.8 GB. Q8 is 17.6 GB and does not fit a 16 GB card. A 12 GB GPU fits Q4_K_M at a short context. qwen3:14b is the practical Ollama tag and needs 10.4 GB.
A 70B model at Q4_K_M is 70 × 0.5 × 1.2 + 2 = 44 GB. That does not fit a 24 GB or 32 GB consumer GPU. Two 24 GB cards, or a 64 GB unified-memory Mac, can hold it. CPU offload runs, and tokens per second collapse.
The published formula already includes a short context inside the 1.2 factor, plus 2 GB for the OS. A 32K or 128K window is extra KV cache on top of that, and it depends on layer count. Use the VRAM calculator context control for a rough add-on. It is not inside the weight-only table.
Tokens per second, once a model fits, follow memory bandwidth. On 2 October 2026 an ASUS TUF RTX 5070 12 GB (ASIN B0DS6S98ZF) was $937.39, an MSI RTX 5070 Ti 16 GB (B0DV9GMDLR) was $1,490, an ASUS TUF RTX 5080 16 GB (B0DQSMMCSH) was $1,809.86, and an ASUS TUF RTX 5090 32 GB OC (B0DS2X13PH) was $7,398. A 27B Q4_K_M is 18.2 GB, so the 16 GB cards miss it. A 24 GB card still wins on fit.
The model falls back to CPU offloading, which is 10-50x slower. A 13B Q4_K_M is 9.8 GB, so an RTX 4060 8 GB is short by about 2 GB and tokens per second collapse. Fitting the entire model in VRAM is the single biggest factor for inference speed.
VRAM and tokens-per-second figures on this page are synthesised from open community benchmarks. The sitewide formula and the full source list are on the methodology page. For this guide specifically I leaned on:
Spot a number that does not match the linked source? Contact us and I will update the guide.