VRAM Calculator for Local Agents (Chat + Tools + Browser + KV)

By Billy G.R. · 2 October 2026

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

Citeable facts

Claim. The site formula is VRAM (GB) ≈ parameters in billions × bytes per parameter × 1.2 + 2.

Method. Published on the methodology page and in llms.txt. Q4_K_M = 0.5, Q5_K_M = 0.625, Q8 = 1.0, FP16 = 2.0. Q4 therefore simplifies to about params × 0.6 + 2.

As of.

Claim. Worked Q4 results: 7B = 6.2 GB, 14B = 10.4 GB, 27B = 18.2 GB, 30B = 20 GB, 32B = 21.2 GB, 70B = 44 GB.

Method. Direct substitution. 30B matches Muse Glimmer’s parameter count. Meta’s own K-quant language model is about 17 GB and is a measured file, not this formula.

As of.

Claim. An extra 1–4 GB for a browser, tool process, or agent scaffold is an editorial planning pad. It is not part of the formula and it is not an NVIDIA or Apple specification.

Method. Labeled as headroom so a 20 GB estimate is not treated as a 20 GB card. The interactive calculator has its own context slider.

As of.

Methodology — parameter math, quantization bytes, and the source list.

Every agent guide on this site points here when the question is “will it fit,” and points at the VRAM calculator when you want to change the number yourself. There is no single vendor behind an “agent GPU.” There is a weight file, a context cache, and whatever process you wrapped around the model. This page separates those three so the wrapper does not get smuggled into the formula.

The formula, once

From the methodology page:

VRAM (GB) ≈ params(B) × bytes_per_param × 1.2 + 2

Bytes: Q4_K_M = 0.5, Q5_K_M = 0.625, Q8 = 1.0, FP16 or BF16 = 2.0. The 1.2 is a short-context allowance (the methodology text says about 4K to 8K). The +2 is an operating-system reserve. At Q4 the expression collapses to about params × 0.6 + 2. That is the line the other guides use. It is not a promise of tokens per second. Bandwidth decides speed after the weights fit, and I am not inventing a speed.

Mixture-of-experts models put total parameters into that formula. Active parameters change how fast a token arrives. They do not excuse you from loading every expert. DeepSeek-V4-Flash-0731 at 284B total is not a 13B memory problem just because about 13B are active.

Worked weights, no scaffold yet

Model Q4 Q5 Q8 FP16 formula
7B, K2-Horizon-7B 6.2 7.25 10.4 18.8
14B 10.4 12.5 18.8 35.6
27B, Qwen3.8-27B 18.2 22.25 34.4 66.8
30B, Muse Glimmer 20 24.5 38 74
32B, K2-Horizon-32B 21.2 26 40.4 78.8
70B dense 44 54.5 86 170

Q5 is params × 0.625 × 1.2 + 2, which is params × 0.75 + 2. Check my arithmetic if you republish a cell: 32 × 0.75 + 2 = 26, 70 × 0.75 + 2 = 54.5. Meta’s Glimmer post says full precision is over 55 GB, while this table’s FP16 cell is 74 GB because of the 1.2 and the +2. Quote Meta’s 55 GB for their weights. Use 74 GB only as the formula’s overhead-inclusive sketch. Their K-quant language model at about 17 GB is the number to plan a 24 GB card around, not the 20 GB cell alone, and not without KV cache and the vision encoder. The Glimmer guide keeps both numbers.

The pad for tools and a browser

After the formula, I add 1–4 GB in my head when the agent is more than chat.

Do not write “NVIDIA recommends +4 GB for agents.” Nobody said that. If your context is 32K, ignore this pad and use the calculator’s context control instead. A long KV cache dwarfs 4 GB. The general VRAM guide is the non-agent version of the same math.

What I would not buy from a formula cell

A 20 GB estimate does not mean a 20 GB card exists. It means a 24 GB card. A 44 GB estimate does not fit the 32 GB RTX 5090. A 6.2 GB estimate does fit an 8 GB card only if you keep the chat short. Agent guides that name Amazon listings — the mini, the Studio, the EVO-X2 — are applying this table, not a different formula. If a number on one of those pages disagrees with a cell here, this page wins, and tell me.

Related guides

Sources

Questions

Why does Meta say Muse Glimmer is about 17 GB if the formula says 20?

Meta published a specific K-quant of the language model, about 17 GB, and said 4-bit weights come in under 20 GB. The site formula is a planning estimate that does not know their quantization recipe. Use 17 GB when you are quoting Meta, and 20 GB when you are comparing Glimmer to other 30B Q4 models. Their full working set, including KV cache, the perception encoder, and the drafter, is a 24–32 GB envelope.

Does the 1.2 factor include a browser agent?

No. The 1.2 factor on the methodology page is a short-context allowance for KV cache, activations, and the runtime, aimed at roughly 4K to 8K tokens. A Playwright process or a Docker sandbox is additional RAM. On a unified-memory Mac it comes out of the same pool as the model. The +1–4 GB note on this page is editorial, not a second term in the formula.

Where do I punch in my own numbers?

The VRAM calculator on this site. This guide is the worked examples for agent stacks. The calculator is the interactive tool. They share the same bytes-per-parameter table. The calculator also exposes a context overhead control for long chats.

A correction or a hardware question goes to the contact page.