Meta Muse Glimmer 30B Local Hardware Guide (VRAM, Quants, Agent Stacks)

By Billy G.R. · 2 October 2026

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

Citeable facts

Claim. Muse Glimmer is a 30B Apache-2.0 model for local agents, not the consumer Muse app.

Method. Meta research post, 10 August 2026, and the Hugging Face card meta-models/Muse-Glimmer-30B.

As of.

Claim. Site-formula Q4 for 30B is 20 GB. Meta’s own K-quant language model is about 17 GB and is aimed at a 24–32 GB envelope once KV, perception, and the drafter are included.

Method. 30 × 0.5 × 1.2 + 2 = 20. Meta’s post: 4-bit language-model weights under 20 GB, K-Quant-17GB called out by name, full precision over 55 GB.

As of.

Claim. Q8 for the same 30B is about 38 GB. That is a 48–64 GB unified-memory tier, not a 24 GB GPU.

Method. 30 × 1.0 × 1.2 + 2 = 38.

As of.

Methodology — parameter math, quantization bytes, and the source list.

Muse Glimmer is Meta’s open-weight agent model: 30 billion parameters, Apache-2.0, built for tool use, failure recovery, screenshots, and local coding. Weights are on Hugging Face. This is not consumer Muse. That product stays on a cloud VM and is powered by Muse Spark. Buying a GPU does not install the Muse app, and installing the Muse app does not put Glimmer on your disk.

What Meta actually promised about memory

The 10 August 2026 research post says full precision for a 30B model is over 55 GB, which is already past any single consumer GPU. They compress the language-model weights to about 4-bit, under 20 GB, and they name a K-Quant around 17 GB. The working set they describe is that language model plus KV cache, a perception encoder for images, and a small DFlash drafter. Their envelope for all of that is 24 GB or 32 GB. They say they measured that 17 GB quant with the drafter on an M4 Max, an M5 Max, and an RTX 5090. The speeds are theirs. I did not re-run them, and the copy I read did not include a numeric table I am willing to retype, so I am not printing tokens per second. The local-models guide quotes the model card’s speed table if you want those figures attributed.

Meta also said llama.cpp, MLX, and ExecuTorch integrations were landing, and that Ollama, LM Studio, and Unsloth were partner paths “in the coming days” from 10 August. By 2 October that may or may not be a one-click pull in your app. Confirm the tag exists before you plan a weekend around it. vLLM and SGLang are the multi-user serve path they named, which is a LAN box, not a laptop.

Quant table: their number and ours

Two numbers can both be honest. Meta measured a specific K-quant. The site formula is a planning estimate that does not know their exact quantization recipe: params × bytes × 1.2 + 2. Use the vendor file size when you have it, and the formula when you are comparing Glimmer to Qwen3.8-27B or K2-32B. Check a context length on the VRAM calculator.

Mode Figure Box that matches
Q4, site formula 30 × 0.6 + 2 = 20 GB 24 GB GPU, or 32 GB unified, before a long chat
Meta K-quant LM About 17 GB language model. Under 20 GB at 4-bit. 24–32 GB once KV, vision, and the drafter join
Q5 30 × 0.625 × 1.2 + 2 = 24.5 GB 32 GB, tight. 36 GB unified is calmer.
Q8 30 × 1.2 + 2 = 38 GB 48–64 GB unified
BF16 / FP16 Formula 30 × 2.4 + 2 = 74 GB. Meta: over 55 GB at full precision. Not a 24 GB or 32 GB card. Quote Meta’s 55 GB for their weights; use 74 GB only as the formula’s overhead-inclusive sketch.

A 16 GB gaming card is below the target Meta published. I would not buy one “for Glimmer.” A 12 GB card is a 7B machine: K2-Horizon-7B at Q4 is 6.2 GB. Different job.

Agent stack around the weights

Glimmer is the brain. It still needs a scaffold. Meta says it works with OpenClaw and other orchestration patterns, and they show an OpenCode demo. I am not deep-linking OpenClaw; the project URL was not re-confirmed in this pass. What I will point at:

Screenshot tasks use the perception encoder, so they sit in the 24–32 GB envelope Meta described, not in the 17 GB language-model file alone. Text-only tool calls can be leaner. Do not plan the vision agent on the language-model number.

What I would buy

The comfortable single-user machine in the Amazon snapshot is the Mac Studio M5 Max 36 GB, ASIN B0HGKSQMX6, $2,449 on 1 October 2026. A 24 GB discrete card also matches Meta’s lower envelope; see an RTX 4090 rather than a 16 GB Blackwell card. If you want Q8 (38 GB) or to serve Glimmer beside another model, 36 GB is the floor and 64 GB configured at Apple, or the EVO-X2 128 GB (B0F53MLYQ6), is the headroom. The 64 GB Studio was not a clean Amazon buy in that pass.

Mac Studio M5 Max 36 GB on Amazon

ASIN B0HGKSQMX6. Amazon Associates link. The live price is on that page, not locked in here.

RTX 4090 24 GB on Amazon

ASIN B0BG94PS2F. Amazon Associates link. The live price is on that page, not locked in here.

Related guides

Sources

A correction or a hardware question goes to the contact page.