Guides
A short list of guides, each one written from real time on the hardware. If you are new here, start with the beginner's guide or check what can I run on my GPU.
Start here
- Beginner's guide to local AI
New to local LLMs? Start here. What "local" means, what hardware you need, and how to pick a first model.
- How to run LLMs locally
Three ways to get a model running on your machine: Ollama, LM Studio, llama.cpp. Walks through install + first chat.
- What can I run on my GPU?
Use this when you already have a GPU and want to know which models fit.
- How much VRAM do I need?
The simple formula: parameters x bytes per param x overhead. Worked through Q4, Q5, Q8, and FP16.
GPU buying guides
- Best GPU for LLMs (overall)
My picks by budget and use case. Real tokens-per-second numbers where I have them.
- Best 24GB VRAM GPU for LLMs
RTX 3090, 4090, 7900 XTX, A5000. Which makes sense in 2026.
- Best budget GPU for LLMs
Under $500 picks. RTX 3060 12GB still leads on $/GB VRAM.
- RTX 4090 LLM guide
What it actually runs, what it does not, and how thermals behave in a real room.
- RTX 3090 LLM guide
The used 3090 is the best $/GB you can buy. Caveats included.
- RTX 4070 vs 4080 vs 4090
Which tier is the right break-point for you. Not always the 4090.
- RTX 5090 vs 4090 for LLMs
Bandwidth still favors the 5090. October 2026 Amazon asks are a different question from the May price.
- Strix Halo mini PCs
Default buy is the GMKtec EVO-X2 128GB at ~256 GB/s. The Beelink GTR9 Pro is the same silicon with a ~$700 premium and an Intel E610 NIC caveat. SER9/SER10 at ~90 GB/s are hobby 7–14B boxes.
Apple Silicon and laptops
- Apple Silicon for LLMs
October 2026 Amazon buy-now: Mac Studio M5 Max 36GB, Mac mini M5 Pro and M4 Pro 24GB, plus the Strix Halo EVO-X2. 64GB+ configures at Apple.
- Mac vs PC for LLMs
When a Mac wins, when a PC wins. Different workloads land in different places.
Software setup
- Ollama cheat sheet
Every command I actually use, with the flags I always forget.
- LM Studio vs Ollama
Pick the right tool for the job. They are not interchangeable.
- How to run LLMs on Windows
Native Windows + WSL2 routes. Driver gotchas included.
- How to run LLMs on Mac
Ollama, LM Studio, and llama.cpp on Apple Silicon. What works best for which model size.
Concepts
- Quantization explained
Why Q4 fits where FP16 cannot, and where quality actually drops.
- GGUF vs GPTQ
The two formats you keep seeing. Which to use when.
- LLM system requirements
CPU, RAM, storage, and PSU notes most guides skip.
Model-specific
- Llama 3.1 hardware requirements
What 8B, 70B, and 405B need. With Q4/Q8/FP16 breakdowns.
- Qwen3 hardware requirements
Original Qwen3 sizes, from 0.6B through the big MoE. VRAM, not a leaderboard.
- Qwen3.8, Muse, and K2 hardware
What 16, 24, and 32 GB actually hold from the May–October 2026 local models. Announcements without weights stay off the list.
- LFM2.5-8B-A1B hardware
Liquid's 8.3B MoE with 1.5B active. Q4 fits an 8 GB GPU; the license is lfm1.0.
- DeepSeek hardware requirements
DeepSeek R1 and V3, including realistic options for the 671B beast.
- Phi-4 hardware requirements
Microsoft Phi-4 14B is the easiest mid-size model to run well.
- Gemma 3 hardware requirements
Google Gemma 3, including the 27B that fits in 24GB VRAM at Q4.
- Mistral hardware requirements
Mistral 7B, Mixtral 8x7B, and Mistral Small / Large.
Cloud agents and the local stack
- Muse vs dots vs Grok Bot
Three cloud always-on agents, then one local blueprint and the full memory legend.
- Meta Muse vs a local agent
Muse stays on a Secure VM. The local path is Ollama, Open WebUI, and a model that fits.
- Muse Glimmer 30B hardware
Meta’s open 30B agent model. Their 17 GB K-quant next to the site formula.
- Grok Bot vs local teammates
Two words, cloud computer, no local Grok weights. OpenBot or OpenHands if the box is yours.
- ChatGPT dots vs local agents
Lowercase dots, GPT-6 Astra, paid plans. Operator is not the current product.
- OpenBot vs OpenBots.ai
CopilotKit’s self-hosted agents, kept apart from the healthcare RPA company.
- Claude Cowork vs a local desktop agent
Folder handoff in the cloud. Locally: OpenHands, browser-use, and approval on.
- Cursor Cloud Agents vs local coding
Formerly Background Agents. Aider, OpenHands, and Continue on hardware you own.
- OpenHands by model size
Agent Canvas and the CLI. 7B is a demo. 27–32B is the coding tier.
- Open WebUI + Ollama home box
Mac mini 24 GB, Studio 36 GB, and the 128 GB EVO-X2, mapped to model tiers.
- VRAM for local agents
Worked 7B through 70B examples, plus a scaffold pad that is not a vendor spec.
- Perplexity Comet vs a local browser
Comet keeps their index. browser-use keeps the click-path on your machine.
- Workspace Agents vs a LAN stack
The team product, not dots. Size one shared model for the group.
- Copilot Studio vs local business agents
Microsoft Graph stays with Microsoft. Local RAG is the documents you actually stored.
- Manus vs a local sandbox
The homepage was too thin to quote. OpenHands is the sandbox I can describe.
- Devin Desktop vs a local fleet
Cloud VMs plus a Desktop. One shared coding model beats a GPU per role.
- Aider and local models
Terminal, git, Ollama. 7B for a small edit, 24 GB for a 27–32B.
- browser-use VRAM tiers
Text extraction on 12 GB. Screenshots want 24 GB and a multimodal model.
- CrewAI on one GPU
Roles share a server. Three copies of a 7B is the expensive way to get a weak crew.
- LangGraph on local hardware
The graph is cheap. The model inside the node is the bill.
- LM Studio as an agent runtime
Proprietary desktop, local server, Bionic on the current homepage. Not open source.
- Open Interpreter, approval on
A terminal agent with a sandbox. No unattended root recommendation.
- Continue after the Cursor acquisition
Still a local IDE client. Not the platform I would build a team on.
- Gemini Skills vs a local preset
Skills replace Gems, with Google’s dates. A modelfile is the local copy of a saved prompt.
Use cases
- Best local LLMs to run
The shortlist that actually holds up across coding, writing, and reasoning tasks.
- Best LLM for coding
Which open model gets closest to GPT-4-class code completion.
- Best LLM for writing locally
For drafting, editing, and tone control. Not the same as coding picks.
- Best LLM for data analysis
Models that handle tabular reasoning and structured outputs well.