Aider + Local Models — Which GPU or Unified Memory for Pair Programming

By Billy G.R. · 2 October 2026

Key facts

Checked
2026-10-02
Author
Billy G.R.

Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.

Citeable facts

Claim. Aider is a terminal pair programmer with git. Its docs describe Ollama and other local backends.

Method. aider.chat and aider.chat/docs/llms/ollama.html. Confirm the provider page still matches your Aider version before you copy a flag.

As of.

Claim. Small edits can sit on a 7B Q4 (about 6.2 GB). Repo-scale edits want a 27–32B Q4 (about 18–21 GB).

Method. Site formula. This is a fit guide, not a SWE-bench score.

As of.

Claim. Aider does not include a GPU. The model server does.

Method. Aider is the client. Ollama or LM Studio holds the weights.

As of.

Methodology — parameter math, quantization bytes, and the source list.

Aider is an open-source pair programmer that lives in the terminal and talks to git. You point it at a model. The model can be a hosted API or a local server. The Ollama notes are at aider.chat/docs/llms/ollama.html. LM Studio’s local server is the other common endpoint, because it speaks an OpenAI-compatible API. Aider is not an IDE fork and it is not a cloud VM. If you wanted Cursor’s remote computer, that is a different product. This page is what has to fit in memory when the model is yours.

How I would pick the model, before the GPU

Aider’s quality is the model’s quality. The git workflow will not rescue a 7B that mis-edits three files. Use K2-Horizon-7B when the change is “rename this function and fix the call sites in one file.” Use Qwen3.8-27B or K2-Horizon-32B when the change crosses a module. Muse Glimmer is an option if you also want an agent model Meta trained for tool use; it is 30B, not a free upgrade in memory. The coding model guide is the qualitative companion. I am not pasting a leaderboard.

Keep the map of files you add to the chat small. Aider’s repo map is a context decision. Context is KV cache. KV cache is why a model that “fits” at a short prompt falls off the card halfway through a refactor. The calculator slider is the right tool for that, not a guess.

Fit table

You have Run this in Aider Q4 formula
8 GB K2-Horizon-7B, short chats 6.2 GB. Leave the rest for the cache.
12–16 GB 14B Q4 on 12 GB is 10.4 GB and tight. 16 GB is still not a 27B card. 27B Q4 is 18.2 GB. It does not fit.
24 GB Qwen3.8-27B. K2-32B if you keep context modest. 18.2 GB and 21.2 GB.
36 GB unified 32B with a longer map, or Glimmer plus a bit of cache The headroom is the point of the extra 12 GB.

Q8 roughly doubles the weight term: a 32B at Q8 is 32 × 1.2 + 2 = 40.4 GB. That is a 48 GB-and-up decision, not a “quality mode” toggle on a 24 GB card. Quantization explains why Q4 is the default.

What to put under it

Any verified 16–24 GB card in the catalog works for the 14B-to-27B step, and only 24 GB works for the 27–32B step. The RTX 4090 listing is B0BG94PS2F. On Mac, the mini M5 Pro 24 GB (B0HGGHNQY6, $1,669.99 on 1 October 2026) is the tight 27B machine, and the Studio M5 Max 36 GB (B0HGKSQMX6, $2,449) is the one I would use if Aider is a daily driver. Aider itself is a pip or uv install. The money is the memory.

RTX 4090 24 GB on Amazon

ASIN B0BG94PS2F. Amazon Associates link. The live price is on that page, not locked in here.

Mac mini M5 Pro 24 GB on Amazon

ASIN B0HGGHNQY6. Amazon Associates link. The live price is on that page, not locked in here.

Related guides

Sources

A correction or a hardware question goes to the contact page.