Qwen3.8-27B, Muse Glimmer-30B, and K2-Horizon on 16, 24, and 32 GB

By Billy G.R. · 2 October 2026 · models from 29 May through 1 October 2026

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

The models worth a local install in that window are not the announcement posts. I checked the Hugging Face cards, then mapped them onto memory tiers. I am not quoting tokens per second I did not measure. Where a vendor published a speed, I say whose number it is.

What fits

MemoryReasonable to loadLeave for a bigger box
8 GB K2-Horizon-7B Q4 (about 4–5 GB). K2-3.7B if you want smaller. Qwen3.8-27B, Muse Glimmer-30B, K2-Horizon-MoVA
16 GB 7B–14B Q4/Q8. MoVA does not fit. Qwen3.8-27B Q4 is the weights alone. Leave this tier for smaller models
24 GB Qwen3.8-27B Q4 with modest context. Muse K-Quant-17GB (their 24 GB target). MoVA Q4, tight. Muse full BF16. DeepSeek-V4.
32 GB Muse dynamic quant (their 32 GB target). Qwen3.8-27B with more KV cache. A 5090 does this, at October 2026 scarcity prices. Still not a huge MoE
64–128 GB unified Muse BF16 (their 64 GB target). DeepSeek-V4-Flash at an aggressive quant on an EVO-X2 128 GB. 64 GB+ Mac Studio configures at Apple — the Amazon 64 GB M4 Max was unavailable on 2026-10-01. Do not expect discrete-GPU tokens/sec

Run your own context length through the VRAM calculator. The ranges above are weights, not weights plus a 32K chat.

Qwen3.8-27B

Qwen/Qwen3.8-27B is a 27B dense model from Alibaba. The LICENSE file on the repo is Apache-2.0 (copyright 2026 Alibaba Cloud). The card describes a vision-language model with a 262,144-token native context. Q4 weights land around 16–19 GB. That is the whole 16 GB card before KV cache, so I do not call 16 GB a fit. 24 GB is the discrete tier where a Q4 download plus a normal chat actually stays on the GPU. 32 GB is headroom, not a different model.

The older Qwen3 hardware guide still covers the 8B/14B/32B generation. This 27B is a later model. Do not assume an Ollama tag from the Qwen3 era points at Qwen3.8.

Muse Glimmer-30B

meta-models/Muse-Glimmer-30B is a ~29.6B dense model with a perception encoder. The model card says Apache-2.0 and an August 2026 release. Their own fit table is the one to trust: a K-Quant around 17 GB aimed at 24 GB, a dynamic quant aimed at 32 GB, and full precision aimed at 64 GB. They also say the 4-bit language-model weights come in under 20 GB, with KV cache and the vision encoder on top. A 16 GB gaming card is below the target they published.

Meta's card reports generation speed for the 17 GB quant plus their DFlash drafter: 74.9 tok/s baseline and 233.4 tok/s with speculation on an RTX 5090, 23.7 / 37.8 on an M4 Max, measured by them with batch size 1 and greedy decoding. Those are their numbers, not a llama-bench run I did. Speculation changes the picture; do not compare that 233 figure to a plain Q4 llama.cpp run of a different model.

K2-Horizon

IFM/K2-Horizon-7B is the dense 7B-class model (the Hub sidebar lists about 9B parameters, which usually includes embeddings). The card calls the release fully open — weights, data, and code — and dates the artifact index 1 October 2026. Read the LICENSE file before you ship a product on it. Q4 is about 4–5 GB, so an 8 GB GPU is enough for the weights. It has a 512K context on the card; you will not get 512K on 8 GB, because the KV cache will not fit. K2-3.7B is the smaller sibling if 8 GB feels tight once you turn context up.

K2-Horizon-MoVA-36B-A4B is the 36B / ~4B-active MoE from the same family (3 September 2026 in the research pass, Apache-2.0 listed there — confirm the file). Q4 of the full expert set is about 20–24 GB. Speed can resemble a small model. Memory cannot. A 16 GB card does not hold it. 24 GB is the minimum discrete tier, and it will be tight with context.

K2-32B, if you grab that checkpoint, behaves like other dense ~32B models: Q4 around 20 GB of weights, so 24 GB again, 32 GB if you want Q5/Q6 or a long chat.

The 24 GB card, and what it costs now

For this set, 24 GB is the useful discrete step: Qwen3.8-27B Q4, Muse's 17 GB quant, and MoVA Q4. A used RTX 3090 or a 4090 does that job. An RTX 5090 adds 32 GB and a lot of bandwidth, and on 2 October 2026 the ASUS TUF 5090 (ASIN B0DS2X13PH) was $7,399 on Amazon with two left. Extra VRAM does not make a 7B model faster once it already fits. It lets you load the next size up. The 5090 vs 4090 guide has the bandwidth comparison; the price explosion is why I would not "upgrade for tok/s" on an 8B.

See an RTX 4090 24 GB on Amazon

Live price only. I am not repeating the May catalog number.

On a Mac, the same 24–36 GB band was in stock on Amazon on 1 October 2026: Mac mini M5 Pro 24 GB and M4 Pro 24 GB for Qwen3.8-27B Q4 (tight) and K2-7B, and Mac Studio M5 Max 36 GB for Qwen3.8-27B plus K2-32B / 36B-A4B. DeepSeek-V4-Flash at an aggressive quant is the EVO-X2 (128 GB at ~256 GB/s) or an Apple Ultra, not those minis and not a 64 GB DDR5 box at ~90 GB/s. The Beelink GTR9 Pro matches the EVO-X2’s bandwidth and adds a NIC caveat; the ranking is in the Strix Halo guide. Mac prices and the out-of-stock 48 GB / 64 GB listings are in the Apple Silicon guide.

Also on the list, with caveats

Announced, not a recommendation

Nex-N2.5-Pro, DeepSeek-V4.1-Pro, and MiMo-V2.6 showed up as announcements or incomplete drops. I will not tell you which GPU "runs" them. When a repo has files, a license, and a size, it can go in the calculator. Until then it is a press post.

Questions

Does a 16 GB GPU run Qwen3.8-27B?

Barely, and often not once KV cache is included. A 27B dense model at Q4 is about 16–19 GB of weights. A 16 GB card has no room left for context. Use 24 GB for a practical Q4 setup. Muse Glimmer's own card targets a 17 GB quant at 24 GB, not 16 GB.

Why is K2-Horizon-MoVA slower to load than it is to run?

It is a 36B-total Mixture-of-Experts with about 4B active parameters. Active parameters set tokens per second. Total parameters set VRAM, because every expert has to be resident. Q4 of the full model is about 20–24 GB, so plan a 24 GB card even though each token only touches a 4B slice.

Can I run DeepSeek-V4 on one gaming GPU?

No. Open weights are not the same thing as "fits in 24 GB." DeepSeek-V4-class MoEs are multi-GPU or large-unified-memory workloads. DeepSeek-V4.1-Pro was announced without a verified local weight drop in this pass, so I am not treating it as runnable. Use a distilled or Flash build only after you have confirmed the file size on Hugging Face.

Which new models should I ignore until weights exist?

Nex-N2.5-Pro, DeepSeek-V4.1-Pro, and MiMo-V2.6 were announcements or incomplete releases in the 29 May–1 October 2026 window. Do not buy hardware for them until a Hugging Face repo you can download is up, with a license and a real file size.