Claim. The ASUS TUF RTX 5090 32 GB OC (B0DS2X13PH) was $7,398 on 2 October 2026, and 32 GB does not fit a 70B model at Q4_K_M.
Method. Amazon product page for B0DS2X13PH. 70 × 0.5 × 1.2 + 2 = 44 GB.
As of.
Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Updated 2 October 2026 · RTX 5090 OC $7,398 · 70B Q4 is 44 GB
Claim. The ASUS TUF RTX 5090 32 GB OC (B0DS2X13PH) was $7,398 on 2 October 2026, and 32 GB does not fit a 70B model at Q4_K_M.
Method. Amazon product page for B0DS2X13PH. 70 × 0.5 × 1.2 + 2 = 44 GB.
As of.
Claim. A 32B model at Q4_K_M needs 21.2 GB. A 34B model at Q8 needs 42.8 GB and does not fit 32 GB.
Method. 32 × 0.5 × 1.2 + 2 = 21.2. 34 × 1.0 × 1.2 + 2 = 42.8. Ollama tag: qwen3:32b.
As of.
Methodology — parameter math, quantization bytes, and the source list.
The right GPU for local LLM inference depends entirely on which models you want to run. An entry-level RTX 4060 is enough for 7B models at Q4 quantization. Running 70B without a multi-GPU setup requires Apple Silicon with 64 GB+ of unified memory. This guide covers every major budget tier and the exact models each card can handle.
Quick picks by model size
ollama run qwen3:32b.Advertisement
| GPU | VRAM | Price | Max model | Quantization | Best for | Buy |
|---|---|---|---|---|---|---|
| RTX 4060 8GB | 8 GB | $609.99 | qwen3:8b (6.8 GB) | Q4_K_M | Listing extract 2026-10-02. Not 14B. | Buy |
| RTX 4060 Ti 16GB | 16 GB | $629.99 | 20B Q4 (14.0 GB) | Q4_K_M | Listing extract 2026-10-02. Not 14B Q8. | Buy |
| RTX 4070 12GB | 12 GB | Not rechecked | qwen3:14b (10.4 GB) | Q4_K_M | May 2026 catalog price, not rechecked | Buy |
| RTX 5070 12GB | 12 GB | $937.39 | qwen3:14b (10.4 GB) | Q4_K_M | ASUS TUF 12 GB on 2026-10-02. Not 16 GB. | Buy |
| RTX 5070 Ti 16GB | 16 GB | $1,490 | 20B Q4 (14.0 GB) | Q4_K_M | MSI 16 GB on 2026-10-02. Not 14B Q8 (18.8 GB). | Buy |
| RTX 4070 Ti Super 16GB | 16 GB | Not rechecked | 20B Q4 (14.0 GB) | Q4_K_M | 16 GB fits 20B Q4, not 13B Q8 (17.6 GB) | Buy |
| RX 7900 XTX 24GB | 24 GB | See price | 34B | Q4_K_M | Budget 24 GB VRAM (AMD) | Buy |
| RTX 3090 24GB | 24 GB | See price | 34B | Q4_K_M | Best used-market value | Buy |
| RTX 4080 16GB | 16 GB | Not rechecked | 20B Q4 (14.0 GB) | Q4_K_M | Not rechecked 2026-10-02. Not 14B Q8. | Buy |
| RTX 4090 24GB | 24 GB | $1,599 (May, not rechecked) | qwen3:32b (21.2 GB) | Q4_K_M | Catalog 2026-05-04. Fits 32B Q4, not 70B Q4. | Buy |
| RTX 5090 32GB | 32 GB | $7,398 | 32B Q5 (26.0 GB) | Q5_K_M | ASUS TUF OC on 2026-10-02. Not 70B Q4 (44 GB). | Buy |
| Mac Studio M4 Max 64GB | 64 GB | Unavailable | 48–64GB tier | Q4 / higher precision 27B | Watchlist — OOS on Amazon 2026-10-01 | Buy |
| Mac Studio M4 Max 128GB | 128 GB | Not re-checked | large MoE quant | aggressive Q4 | Configure 64GB+ at Apple | Buy |
| NVIDIA DGX Spark | 128 GB | See price | 200B | Q4_K_M | Best desktop NVIDIA 70B–200B inference | Buy |
"Max model" is the largest parameter count that fits in VRAM at the listed quantization with at least 1 GB headroom. Tokens/sec ranges (in the tier breakdown below) are cross-referenced with the XiongjieDai community llama-bench runs and the Home GPU LLM Leaderboard. Use the VRAM Calculator for exact figures. Compare any two options side-by-side on the Compare page.
RTX 4060 8GB
Entry-level local inference, 7B chat & coding models
VRAM
8 GB
Max model
7B at Q4_K_M
Speed
20–30 t/s
Pros
Cons
The RTX 4060 is the go-to entry point for local LLMs in 2026. It runs any 7B model at Q4_K_M quantization with a comfortable 1–3 GB of headroom, delivering 20–30 tokens/second — fast enough for real-time chat. You cannot run 13B or larger models without slow CPU offloading, but a well-tuned 7B model handles most everyday tasks.
RTX 4060 Ti 16GB
Best VRAM-per-dollar — 16 GB CUDA. Listing extract 2 October 2026.
VRAM
16 GB
Max model
20B Q4 (14.0 GB) / 14B Q4 (10.4 GB)
Speed
20–35 t/s
Pros
Cons
The RTX 4060 Ti 16 GB (MSI Ventus 2X, B0CBK7BRL9) had a listing extract of $629.99 on 2 October 2026. 16 GB fits qwen3:14b at Q4_K_M (10.4 GB) and a 20B model at Q4_K_M (14.0 GB). It does not fit 13B at Q8 (17.6 GB) or 14B at Q8 (18.8 GB). In LM Studio, download the Q4_K_M GGUF. Run ollama run qwen3:14b, then ollama ps.
RTX 4070 12GB
Fast 7B at Q8, 13B at Q4, best CUDA support
VRAM
12 GB
Max model
13B at Q4_K_M
Speed
30–50 t/s
Pros
Cons
RTX 4070 Ti Super 16GB
Best CUDA 16GB value — 2.3x faster than 4060 Ti at same VRAM
VRAM
16 GB
Max model
20B Q4 (14.0 GB) / 14B Q4 (10.4 GB)
Speed
50–100 t/s at 7B Q4
Pros
Cons
RX 7900 XTX 24GB
34B models on a budget, large VRAM at low cost
VRAM
24 GB
Max model
34B at Q4_K_M
Speed
15–25 t/s
Pros
Cons
The RTX 4070 (12 GB) fits 13B at Q4_K_M (9.8 GB) and qwen3:14b (10.4 GB). It does not fit 14B at Q8 (18.8 GB). The RTX 4070 Ti Super (16 GB) fits 20B at Q4 (14.0 GB), not 13B at Q8 (17.6 GB). The RX 7900 XTX (24 GB) fits qwen3:32b at Q4 (21.2 GB). Its catalog price of $799 is from May 2026 and was not rechecked. AMD ROCm needs more setup than CUDA.
RTX 3090 24GB (used)
Budget 24 GB VRAM build — same capacity as RTX 4090
VRAM
24 GB
Max model
34B at Q4_K_M
Speed
15–22 t/s
Pros
Cons
The RTX 3090 is the used-market standout for 2026. You get 24 GB of CUDA VRAM — the same as the RTX 4090 — on eBay or local resale markets. Inference speed is about 25–30% lower than the RTX 4090 due to the older memory architecture, but you can run the same models. If you want 24 GB VRAM without paying flagship prices, this is the best-value move.
RTX 4080 16GB
Fast 14B–20B Q4. Catalog price not rechecked 2 October 2026.
VRAM
16 GB
Max model
20B Q4 (14.0 GB)
Speed
35–55 t/s at 13B
Pros
Cons
The RTX 4080 is the speed pick in the 16 GB tier. You get the same 16 GB of VRAM as the RTX 4060 Ti but with 716 GB/s bandwidth — roughly 2.5x more throughput. That translates to noticeably faster tokens per second on 13B–20B models. If inference speed matters more than VRAM capacity, the 4080 is worth the premium over the 4060 Ti. For max VRAM on a budget, go with the 4060 Ti instead.
RTX 4090 24GB
Best single-GPU fit for 32B Q4. Price not rechecked 2 October 2026.
VRAM
24 GB
Max model
34B Q4 (22.4 GB) / 32B Q4 (21.2 GB)
Speed
25–40 t/s at 13B
Pros
Cons
The RTX 4090 remains the best single consumer GPU for local LLMs in 2026. Its 24 GB GDDR6X runs 34B models at Q4_K_M and 13B models at Q8 with headroom to spare. Inference throughput is the fastest available in a consumer card — roughly twice as fast as the RTX 3090 on the same model. If you have the budget and want one card that handles everything up to 34B, this is it.
RTX 5090 32GB
32 GB GDDR7. Does not fit 70B Q4 (44 GB) or 34B Q8 (42.8 GB).
VRAM
32 GB
Max model
32B Q5 (26.0 GB) / 34B Q4 (22.4 GB)
Speed
35–55 t/s at 13B
Pros
Cons
The RTX 5090 is 32 GB of GDDR7. On 2 October 2026 the ASUS TUF OC (B0DS2X13PH) was $7,398. That fits a 34B model at Q4_K_M (22.4 GB) and a 32B model at Q5_K_M (26.0 GB). It does not fit 34B at Q8 (42.8 GB), 32B at Q8 (40.4 GB), or 70B at Q4_K_M (44 GB). Use ollama run qwen3:32b. In LM Studio, the matching file is Qwen3 32B Q4_K_M or Q5_K_M, not Q8.
Mac Studio M4 Max (64GB)
70B models, silent operation, macOS ecosystem
VRAM
64 GB unified
Max model
70B at Q4_K_M, 34B at Q8
Speed
8–15 t/s at 70B
Pros
Cons
Mac Studio M4 Max (128GB)
Near-server-grade local inference, macOS ecosystem, 70B at Q8
VRAM
128 GB unified
Max model
70B at Q8, 34B at FP16
Speed
6–12 t/s at 70B Q8
Pros
Cons
NVIDIA DGX Spark
70B–200B NVIDIA inference, CUDA/TensorRT-LLM stack
VRAM
128 GB unified (LPDDR5X)
Max model
70B at Q8, 200B at Q4
Speed
~8 t/s at 70B Q4, ~4 t/s at 70B Q8
Pros
Cons
For 70B+ models, the two practical desktop options are the Mac Studio M4 Max and the NVIDIA DGX Spark. The Mac Studio M4 Max 64 GB runs 70B at Q4_K_M. The NVIDIA DGX Spark runs 70B at Q8 and 200B at Q4 — the most capable desktop AI workstation for NVIDIA users, at a lower price than the Mac Studio 128 GB. For CUDA/TensorRT-LLM workflows, the DGX Spark is the clear choice. For macOS-native tooling, the Mac Studio wins.
For most budgets, the RTX 4090 24GB is the best GPU for local LLMs in 2026 — it runs 34B models at Q4 and 13B at FP16. On a tighter budget, the RTX 4070 12GB handles 7B–13B models well. Apple Silicon M4 Pro/Max wins when you need 48GB+ at low power.
Every "max model" figure in this guide assumes Q4_K_M quantization unless noted. Quantization determines how many bits each model weight uses — and therefore how much VRAM the model occupies.
| Format | Bits/weight | Bytes/param | Quality loss | Typical use |
|---|---|---|---|---|
| Q4_K_M | 4-bit | ~0.5 | ~1–3% | Default for consumer GPUs — best VRAM efficiency |
| Q5_K_M | 5-bit | ~0.625 | ~0.5–1% | Slight quality improvement over Q4 with moderate VRAM cost |
| Q8_0 | 8-bit | ~1.0 | <0.1% | Near-lossless — use when VRAM allows |
| FP16 | 16-bit | 2.0 | Reference | Full precision — fine-tuning, research |
Use the VRAM Calculator to compute exact memory requirements for any model size and quantization combination.
The RTX 4090 (24 GB) fits 32B and 34B at Q4_K_M (21.2 GB and 22.4 GB). The RTX 5090 (32 GB) fits 32B at Q5_K_M (26.0 GB) and 34B at Q4_K_M. It does not fit 34B at Q8 (42.8 GB). A 70B model at Q4_K_M is 44 GB, so it needs two 24 GB cards or 64 GB of unified memory. On 2 October 2026 the ASUS TUF RTX 5090 OC (B0DS2X13PH) was $7,398.
Yes. The RTX 4060 (8 GB) runs 7B models like Qwen3 8B, Llama 3.1 8B, and Gemma 3 4B at Q4_K_M quantization with smooth 20–30 tokens/second throughput. For most everyday chat and coding tasks, a well-tuned 7B model at Q4 is genuinely useful.
Yes — the RTX 3090 offers 24 GB VRAM on the used market, matching the RTX 4090 in memory capacity for a fraction of the cost. Inference speed is roughly 20% lower. For a budget-conscious builder who wants to run 13B–34B models at Q4_K_M, it is the best used-market value in 2026.
For 7B–34B models, a GPU PC with an RTX 4090 is faster and cheaper. For 70B models, the Mac Studio M4 Max is the only practical single-device option. Macs also have zero driver overhead, near-silent operation, and excellent llama.cpp Metal support.
The RTX 4070 (12 GB) is faster for 7B–13B models and has better software support via CUDA. The RX 7900 XTX (24 GB) doubles the VRAM, enabling 34B models at Q4_K_M — but AMD ROCm requires more setup. If you want maximum VRAM per dollar and are comfortable with AMD, the 7900 XTX wins. If you want the easiest setup, the RTX 4070 wins.
A 70B model at Q4_K_M is 70 × 0.5 × 1.2 + 2 = 44 GB. That does not fit a 24 GB or 32 GB consumer GPU, including the RTX 5090. Two 24 GB cards (48 GB) or 64 GB of unified memory can hold it. CPU offload runs, and tokens per second collapse.
Yes. The RTX 5090 (32 GB GDDR7, ASIN B0DS2X13PH) works with Ollama, llama.cpp, and LM Studio. On 2 October 2026 the ASUS TUF OC listing was $7,398. 32 GB fits a 34B model at Q4_K_M (22.4 GB) and a 32B model at Q5_K_M (26.0 GB). It does not fit 34B at Q8 (42.8 GB) or 70B at Q4_K_M (44 GB). Pull qwen3:32b for the 24–32 GB tier.
Find hardware for a specific model or check exact VRAM requirements.
VRAM Calculator
Same formula: 32B Q4 is 21.2 GB and 70B Q4 is 44 GB.
How Much VRAM Do I Need?
Q4, Q5, Q8, and FP16 for 7B through 70B.
Best Budget GPUs
The RTX 5070 is 12 GB at $937.39, not 16 GB.
Best 24 GB VRAM GPU
Cards that fit qwen3:32b at 21.2 GB.
Ollama Cheat Sheet
ollama run qwen3:14b and how to read ollama ps.
LLM Quantization Guide
Q4_K_M is 0.5 bytes per parameter. Q8 is 1.0.
VRAM and tokens-per-second figures on this page are synthesised from open community benchmarks. The sitewide formula and the full source list are on the methodology page. For this guide specifically I leaned on:
Spot a number that does not match the linked source? Contact us and I will update the guide.