Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Citeable facts
Claim. Consumer Meta Muse runs on a cloud Muse Secure VM. It is not a model you load on a GPU.
Method. Meta’s 8 September 2026 announcement (updated 30 September) says Muse lives on a dedicated cloud VM, powered by Muse Spark, and is messaged from the Muse app or WhatsApp.
As of.
Claim. A 7B dense model at Q4_K_M needs about 6.2 GB. A 27B dense model needs about 18.2 GB.
Method. Site formula: params × 0.5 × 1.2 + 2. K2-Horizon-7B is the 7B example. Qwen3.8-27B is the 27B example.
As of.
Claim. Muse Glimmer is a different product: a 30B Apache-2.0 model meant to run locally.
Method. Meta’s 10 August 2026 research post. Consumer Muse and Muse Glimmer are not interchangeable names.
As of.
Methodology
— parameter math, quantization bytes, and the source list.
Meta Muse is Meta’s consumer personal AI agent. The announcement describes a dedicated Muse Secure VM in the cloud, a separate Sentinel that gates what reaches the internet, and a chat that feels like messaging a contact in the Muse app or WhatsApp. The model behind that product is Muse Spark. Muse is free for most of what people need, with subscriptions if you want more. None of that is a GGUF on your GPU.
What the cloud agent actually does
Meta’s post says Muse can keep working after you close the app, open a browser, fill forms, and check with you before it sends mail or pays. Credentials go into secure storage so the agent uses them without seeing the password. You pick which apps it connects to, you get an audit trail, and you can tell it to forget something. The rollout described there is the US, on iOS, Android, and muse.ai, with AI glasses later. A download page lives at ai.meta.com/muse/download/. I am not claiming a local Mac runtime from that page.
Later in the same post Meta says a Muse Confidential VM, encrypted with a key only you hold, is planned. That is still their computer. If the requirement is “the weights and the files stay on hardware I own,” Muse is the wrong tool, and the rest of this page is the replacement.
The local stack I would actually stand up
A home always-on agent is three jobs, not one app. Something has to hold the weights. Something has to be the chat door you can leave running. Something has to touch files, a shell, or a browser, and it should ask first.
Front door.Open WebUI on the same box, pointed at that API. Tools and a knowledge folder live here. This is the piece that stays up when you are not at the keyboard.
Hands. OpenHands (Agent Canvas or the CLI — the old Docker Local GUI is deprecated in their docs) for code. browser-use when the task is a website. Open Interpreter when you want a terminal agent with an approval mode. Meta’s Glimmer post names OpenClaw as a compatible scaffold. I am not linking an OpenClaw repository until that URL is confirmed.
This stack will not book a flight through Meta’s checkout or inherit your WhatsApp identity. It will read a folder you mounted and call tools you enabled. That is the trade. The Muse, dots, and Grok Bot comparison puts the three cloud products on one page if you are still choosing which bill to keep.
VRAM and unified memory, by the site formula
Cloud Muse needs no VRAM. The local brain does. The formula on the methodology page is VRAM ≈ params × bytes × 1.2 + 2. At Q4_K_M, bytes are 0.5, so that is about params × 0.6 + 2. I am not publishing tokens per second. Fit comes first. Run a specific context length through the VRAM calculator, and read the agent-workload walkthrough before you add a browser on top.
Memory
Agent brain
Q4 site formula
6–12 GB
K2-Horizon-7B, always on
7 × 0.6 + 2 = 6.2 GB
16–24 GB
Qwen3.8-27B. Muse Glimmer 30B if you want the open agent model.
27B = 18.2 GB. 30B Q4 = 20 GB. Meta’s K-quant language model is about 17 GB, before KV and the perception encoder.
36–64 GB
32B comfortable. 70B Q4 only from 48 GB up.
32 × 0.6 + 2 = 21.2 GB. 70 × 0.6 + 2 = 44 GB.
96–128 GB
Large MoE experiments, not a smarter 7B.
DeepSeek-V4-Flash-0731 is 284B total / about 13B active. Memory follows the total.
A 14B model lands at 10.4 GB, which is why a 12 GB card is the tight end of “always-on 14B” and a bad place for a 27B. Leave a few gigabytes for the OS and the tool process. That headroom is editorial, not a line in NVIDIA’s spec sheet.
Hardware that matches those tiers
For a quiet always-on box, the Mac mini M5 Pro with 24 GB (ASIN B0HGGHNQY6) was $1,669.99 on Amazon on 1 October 2026. That is the 7B-comfortable / 27B-tight machine. The Mac Studio M5 Max with 36 GB (ASIN B0HGKSQMX6) was $2,449 the same day, which is where Qwen3.8-27B and a 32B Q4 stop feeling cramped. A large MoE wants the GMKtec EVO-X2 128 GB (ASIN B0F53MLYQ6, $3,649.99 that day) or an Apple Ultra you configure at Apple. I am not inventing a 64 GB Amazon ASIN. Those Studio configs were an Apple order, not a listing I could buy off the shelf.
ASIN B0HGGHNQY6. Amazon Associates link. The live price is on that page, not locked in here.
On a discrete GPU, 24 GB is the step that matters for a 27–30B Q4. An RTX 4090 does that job. An RTX 5090 is 32 GB; the ASUS TUF listing B0DS2X13PH showed $7,398 on 2 October 2026. Extra dollars do not make a 7B faster once it fits. A 70B Q4 at 44 GB still does not fit that 32 GB card.
This site uses cookies and shows personalised ads via Google AdSense. We and our partners store and access information on your device to serve relevant ads and improve your experience.
You can accept all cookies, decline (non-personalised ads only), or
manage preferences.
See our Privacy Policy.
Cookie preferences
Choose which cookies you allow. Strictly necessary cookies are always active.
Strictly necessary
Session state, security, and performance. Cannot be disabled.