Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Citeable facts
Claim. Open WebUI is a self-hosted interface for Ollama and OpenAI-compatible APIs, including a fully offline install.
Method. docs.openwebui.com, read 2 October 2026. Docker, pip, uv, and a desktop app are all documented. The :main image bundles embeddings; :slim does not.
As of.
Claim. Mac mini M5 Pro 24 GB maps to a 7B comfortably and a 27B Q4 tightly (18.2 GB). Studio M5 Max 36 GB maps to 27–32B with headroom. EVO-X2 128 GB is the large-MoE tier.
Method. Q4 formula params × 0.5 × 1.2 + 2. ASINs B0HGGHNQY6, B0HGKSQMX6, B0F53MLYQ6. Prices checked 1 October 2026 and will move.
As of.
Claim. Open WebUI Computer is a separate, faster-moving agent harness from the same project. It is not required to run the chat UI.
Method. Same docs page. Computer uses approvals and plan mode on the machine you install it on. The stable multi-user product described there is Open WebUI itself.
As of.
Methodology
— parameter math, quantization bytes, and the source list.
Open WebUI plus Ollama is the stack I would tell a person to install if they want a ChatGPT-shaped window on models they downloaded. Open WebUI’s docs call it a self-hosted platform that can run offline, with Ollama and any OpenAI-compatible API behind it. Ollama is the model runtime: pull a GGUF-style tag, serve it on localhost. Neither product is a cloud agent. Meta Muse, dots, and Grok Bot are the cloud products this box is an alternative to. The comparison is on the hub page.
What you are installing, in the docs’ own split
The docs give four ways in: Docker, pip, uv, and a desktop app from their GitHub. For a machine that should come back after a reboot, Docker with a restart policy is the boring choice. They distinguish a :main image, which carries embedding, speech, and reranking models so retrieval works immediately, from a :slim image when your chat model is already served elsewhere. Those bundled models are small helpers. They are not Qwen3.8.
The same page describes two faster-moving siblings. Open Terminal gives an agent a terminal and a file browser inside the chat. Open WebUI Computer is a separate harness for the real machine — files, git, a terminal — with approvals and a plan mode, installed with a uv tool, not by clicking chat. The docs say Computer moves faster and that pieces graduate into Open WebUI when they settle. If you want an always-on family chat box, install Open WebUI and Ollama and stop. If you want the agent to edit the disk, read the Computer page and leave approvals on. I am not walking through unattended root.
Their “connect an agent” list also names OpenClaw. I am not linking a repository for it. Meta’s Glimmer post uses the same name for a scaffold, and I still do not have a confirmed project URL I am willing to publish.
Three boxes, three model tiers
Always-on means the weights stay loaded. The formula is params × bytes × 1.2 + 2. At Q4_K_M that is params × 0.6 + 2. Prices below are the 1 October 2026 Amazon snapshot. Check the live page. The calculator is how you test a context length I did not.
Machine
Leave loaded
Do not expect
Mac mini M5 Pro 24 GB. B0HGGHNQY6. $1,669.99 on 1 Oct 2026.
K2-Horizon-7B at 6.2 GB all day. Qwen3.8-27B at 18.2 GB if the chat stays short.
A 32B (21.2 GB) plus macOS plus a browser. A 70B.
Mac Studio M5 Max 36 GB. B0HGKSQMX6. $2,449 the same day.
27B with KV headroom. K2-32B. Muse Glimmer’s ~17 GB K-quant plus cache.
An aggressive quant of a large MoE (DeepSeek-V4-Flash is 284B total). Or a 70B Q4 at 44 GB with room.
Discrete-GPU tokens per second. A second 70B resident beside the first.
The Strix Halo guide has the bandwidth and the Beelink caveat. The Apple Silicon guide has the configurations that were out of stock. I am not adding a 64 GB Amazon ASIN. Order that from Apple if you need it. Worked agent overhead, separate from this formula, is in the agent VRAM guide.
Which listing, after you know the tier
Start with the mini if the household agent is a 7B that answers from a notes folder. Move to the Studio when you want the 27B to stay loaded while someone else uses the Mac. Move to the EVO-X2 when the model itself does not fit in 36 GB. Ollama on all three. Open WebUI on all three. The software does not care. The weights do.
No. You connect Ollama, vLLM, or another provider. The docs say a built-in path stays empty until you configure one. The :main Docker image does bundle embedding and speech models for retrieval and voice. That is not the chat model.
Which of the three Amazon boxes matches which model?
Mac mini M5 Pro 24 GB (B0HGGHNQY6): K2-Horizon-7B at Q4 (about 6.2 GB), and Qwen3.8-27B at Q4 (about 18.2 GB) only if you keep context short. Mac Studio M5 Max 36 GB (B0HGKSQMX6): that 27B with room, or K2-32B at about 21.2 GB. GMKtec EVO-X2 128 GB (B0F53MLYQ6): a large MoE such as an aggressive DeepSeek-V4-Flash quant. It is not CUDA.
Is Open WebUI Computer the same app?
No. The docs describe Computer as a faster agent runtime for the actual machine, installed separately, with approvals and plan mode. Open WebUI is the stable multi-user platform. You can point Open WebUI at a Computer workspace, but you do not need Computer to chat with Ollama.
Where do I ask a question about a fit?
Use the contact page on this site. Do not send hardware questions to a personal mailbox. Prices on Amazon move; the ASINs above are the listings, not a quote.
This site uses cookies and shows personalised ads via Google AdSense. We and our partners store and access information on your device to serve relevant ads and improve your experience.
You can accept all cookies, decline (non-personalised ads only), or
manage preferences.
See our Privacy Policy.
Cookie preferences
Choose which cookies you allow. Strictly necessary cookies are always active.
Strictly necessary
Session state, security, and performance. Cannot be disabled.