Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Citeable facts
Claim. The site formula is VRAM (GB) ≈ parameters in billions × bytes per parameter × 1.2 + 2.
Method. Published on the methodology page and in llms.txt. Q4_K_M = 0.5, Q5_K_M = 0.625, Q8 = 1.0, FP16 = 2.0. Q4 therefore simplifies to about params × 0.6 + 2.
Method. Direct substitution. 30B matches Muse Glimmer’s parameter count. Meta’s own K-quant language model is about 17 GB and is a measured file, not this formula.
As of.
Claim. An extra 1–4 GB for a browser, tool process, or agent scaffold is an editorial planning pad. It is not part of the formula and it is not an NVIDIA or Apple specification.
Method. Labeled as headroom so a 20 GB estimate is not treated as a 20 GB card. The interactive calculator has its own context slider.
As of.
Methodology
— parameter math, quantization bytes, and the source list.
Every agent guide on this site points here when the question is “will it fit,” and points at the VRAM calculator when you want to change the number yourself. There is no single vendor behind an “agent GPU.” There is a weight file, a context cache, and whatever process you wrapped around the model. This page separates those three so the wrapper does not get smuggled into the formula.
Bytes: Q4_K_M = 0.5, Q5_K_M = 0.625, Q8 = 1.0, FP16 or BF16 = 2.0. The 1.2 is a short-context allowance (the methodology text says about 4K to 8K). The +2 is an operating-system reserve. At Q4 the expression collapses to about params × 0.6 + 2. That is the line the other guides use. It is not a promise of tokens per second. Bandwidth decides speed after the weights fit, and I am not inventing a speed.
Mixture-of-experts models put total parameters into that formula. Active parameters change how fast a token arrives. They do not excuse you from loading every expert. DeepSeek-V4-Flash-0731 at 284B total is not a 13B memory problem just because about 13B are active.
Worked weights, no scaffold yet
Model
Q4
Q5
Q8
FP16 formula
7B, K2-Horizon-7B
6.2
7.25
10.4
18.8
14B
10.4
12.5
18.8
35.6
27B, Qwen3.8-27B
18.2
22.25
34.4
66.8
30B, Muse Glimmer
20
24.5
38
74
32B, K2-Horizon-32B
21.2
26
40.4
78.8
70B dense
44
54.5
86
170
Q5 is params × 0.625 × 1.2 + 2, which is params × 0.75 + 2. Check my arithmetic if you republish a cell: 32 × 0.75 + 2 = 26, 70 × 0.75 + 2 = 54.5. Meta’s Glimmer post says full precision is over 55 GB, while this table’s FP16 cell is 74 GB because of the 1.2 and the +2. Quote Meta’s 55 GB for their weights. Use 74 GB only as the formula’s overhead-inclusive sketch. Their K-quant language model at about 17 GB is the number to plan a 24 GB card around, not the 20 GB cell alone, and not without KV cache and the vision encoder. The Glimmer guide keeps both numbers.
The pad for tools and a browser
After the formula, I add 1–4 GB in my head when the agent is more than chat.
+0 GB in the formula for “the model answers in Open WebUI.” The 1.2 and the +2 are already there.
+1–2 GB editorial when a tool process or LM Studio’s GUI is open beside a short chat.
+2–4 GB editorial when Playwright or a Docker sandbox is up. On a discrete GPU this is mostly host RAM. On a Mac it is the same unified pool, so it can push a “fits in 24 GB” 27B into swap.
Do not write “NVIDIA recommends +4 GB for agents.” Nobody said that. If your context is 32K, ignore this pad and use the calculator’s context control instead. A long KV cache dwarfs 4 GB. The general VRAM guide is the non-agent version of the same math.
What I would not buy from a formula cell
A 20 GB estimate does not mean a 20 GB card exists. It means a 24 GB card. A 44 GB estimate does not fit the 32 GB RTX 5090. A 6.2 GB estimate does fit an 8 GB card only if you keep the chat short. Agent guides that name Amazon listings — the mini, the Studio, the EVO-X2 — are applying this table, not a different formula. If a number on one of those pages disagrees with a cell here, this page wins, and tell me.
Why does Meta say Muse Glimmer is about 17 GB if the formula says 20?
Meta published a specific K-quant of the language model, about 17 GB, and said 4-bit weights come in under 20 GB. The site formula is a planning estimate that does not know their quantization recipe. Use 17 GB when you are quoting Meta, and 20 GB when you are comparing Glimmer to other 30B Q4 models. Their full working set, including KV cache, the perception encoder, and the drafter, is a 24–32 GB envelope.
Does the 1.2 factor include a browser agent?
No. The 1.2 factor on the methodology page is a short-context allowance for KV cache, activations, and the runtime, aimed at roughly 4K to 8K tokens. A Playwright process or a Docker sandbox is additional RAM. On a unified-memory Mac it comes out of the same pool as the model. The +1–4 GB note on this page is editorial, not a second term in the formula.
Where do I punch in my own numbers?
The VRAM calculator on this site. This guide is the worked examples for agent stacks. The calculator is the interactive tool. They share the same bytes-per-parameter table. The calculator also exposes a context overhead control for long chats.
This site uses cookies and shows personalised ads via Google AdSense. We and our partners store and access information on your device to serve relevant ads and improve your experience.
You can accept all cookies, decline (non-personalised ads only), or
manage preferences.
See our Privacy Policy.
Cookie preferences
Choose which cookies you allow. Strictly necessary cookies are always active.
Strictly necessary
Session state, security, and performance. Cannot be disabled.