Claim. A 32B model at Q4_K_M needs 21.2 GB. It fits 24 GB and does not fit 16 GB.
Method. 32 × 0.5 × 1.2 + 2 = 21.2. Ollama tag: ollama run qwen3:32b.
As of.
I used AI to sweep the 24 GB GPU landscape into a first draft, then walked the numbers row by row against the Home GPU LLM Leaderboard and the XiongjieDai runs.
Checked 2026-10-02. Author: Billy G.R. Retail prices move; the hardware catalog stores the Amazon snapshot, not a promise of stock.
Updated 2 October 2026 · qwen3:32b is 21.2 GB · 70B Q4 is 44 GB
Claim. A 32B model at Q4_K_M needs 21.2 GB. It fits 24 GB and does not fit 16 GB.
Method. 32 × 0.5 × 1.2 + 2 = 21.2. Ollama tag: ollama run qwen3:32b.
As of.
Claim. A 27B model at Q4_K_M needs 18.2 GB, so a 16 GB Blackwell card is not a substitute for 24 GB.
Method. 27 × 0.5 × 1.2 + 2 = 18.2.
As of.
Methodology — parameter math, quantization bytes, and the source list.
24 GB fits qwen3:32b at Q4_K_M (21.2 GB) and a 34B model at Q4_K_M (22.4 GB). It also fits 13B at Q8 (17.6 GB) and 14B at Q8 (18.8 GB). A 70B model at Q4_K_M is 44 GB and does not fit. This guide compares the 24 GB consumer cards. Prices for the RTX 4090 (catalog $1,599 on 4 May 2026), RTX 3090 (catalog $499), and RX 7900 XTX (catalog $799) were not rechecked on 2 October 2026.
Quick recommendations
12 GB fits qwen3:14b at Q4_K_M (10.4 GB). 24 GB fits qwen3:32b at Q4_K_M (21.2 GB) and a 34B model at Q4_K_M (22.4 GB). The 22.4 GB figure already includes a short context, so the leftover on a 24 GB card is about 1.6 GB. A 32K window adds about 6 GB on the calculator and misses 24 GB.
13B at Q8 is 17.6 GB and fits. 13B at FP16 is 33.2 GB and does not. 14B at Q8 is 18.8 GB and fits 24 GB, not 16 GB. Run ollama run qwen3:32b on these cards. In LM Studio, use the Q4_K_M GGUF.
What 24 GB unlocks vs 12 GB
RTX 4090 24 GB
Fastest 24 GBBest overall — fastest CUDA, new build, fine-tuning
VRAM
24 GB
BW
1008 GB/s
13B Q4
~55 tok/s
Price
Check price on Amazon
Pros
Cons
RTX 3090 Ti 24 GB
Best ValueBest value — same bandwidth as RTX 4090 at 60% less cost
VRAM
24 GB
BW
1008 GB/s
13B Q4
~55 tok/s
Price
Check price on Amazon
Pros
Cons
RX 7900 XTX 24 GB
Best AMDBest AMD option — excellent ROCm on Linux, competitive bandwidth
VRAM
24 GB
BW
960 GB/s
13B Q4
~50 tok/s
Price
Check price on Amazon
Pros
Cons
RTX 3090 24 GB
Budget PickBest budget 24 GB — 7% slower than RTX 4090 at less than a third of the price
VRAM
24 GB
BW
936 GB/s
13B Q4
~48 tok/s
Price
Check price on Amazon
Pros
Cons
| GPU | VRAM | Bandwidth | 13B Q4 Speed | 34B Q4 Speed | Price | Best For |
|---|---|---|---|---|---|---|
| RTX 4090 24 GB | 24 GB | 1008 GB/s | ~55 tok/s | ~25 tok/s | Check price on Amazon | CUDA, new builds |
| RTX 3090 Ti 24 GB | 24 GB | 1008 GB/s | ~55 tok/s | ~25 tok/s | Check price on Amazon | Best value |
| RX 7900 XTX 24 GB | 24 GB | 960 GB/s | ~50 tok/s | ~22 tok/s | Check price on Amazon | AMD, Linux |
| RTX 3090 24 GB | 24 GB | 936 GB/s | ~48 tok/s | ~22 tok/s | Check price on Amazon | Budget CUDA |
Speed figures are approximate on Linux with Ollama. Use the VRAM Calculator for exact memory requirements.
At Q4 quantization, a model uses roughly 0.65 GB per billion parameters. A 34B model needs about 20 GB, fitting in 24 GB with headroom for KV cache. Here is the full picture across popular model sizes:
| Model | VRAM Required | Fits 24 GB? | Notes |
|---|---|---|---|
| Llama 3.3 70B Q4 (llama3.3:70b) | 44 GB | No | 70 × 0.5 × 1.2 + 2. Does not fit 24 GB. |
| Qwen3 32B Q8 | 40.4 GB | No | 32 × 1.0 × 1.2 + 2. Use Q4 instead. |
| 34B Q4_K_M | 22.4 GB | Yes | 34 × 0.5 × 1.2 + 2. Tight on 24 GB at a short context. |
| qwen3:32b Q4_K_M | 21.2 GB | Yes | 32 × 0.5 × 1.2 + 2. The current Ollama tag. Not qwen3:34b. |
| 13B Q8 | 17.6 GB | Yes | 13 × 1.0 × 1.2 + 2. Fits 24 GB. Does not fit 16 GB. |
| 13B FP16 | 33.2 GB | No | 13 × 2.0 × 1.2 + 2. Use Q8 on a 24 GB card. |
| qwen3:14b Q8 | 18.8 GB | Yes | 14 × 1.0 × 1.2 + 2. Fits 24 GB. Does not fit 16 GB. |
| 7B FP16 | 18.8 GB | Yes | 7 × 2.0 × 1.2 + 2. |
| 27B Q4_K_M | 18.2 GB | Yes | 27 × 0.5 × 1.2 + 2. A 16 GB card misses this. |
LLM inference speed is bottlenecked by memory bandwidth, not compute. Models run as decode loops where each token requires a full pass through the model weights stored in VRAM. The faster the GPU can stream those weights, the faster inference runs. This is why bandwidth, not TFLOPS, determines LLM performance.
RTX 4090
1008 GB/s
priciest of the three
RTX 3090 Ti (used)
1008 GB/s
best value
RTX 3090 (used)
936 GB/s
cheapest of the three
The RTX 3090 Ti has the same 1008 GB/s bandwidth as the RTX 4090 because both use GDDR6X at the same effective memory clock. The architectural difference between Ampere and Ada Lovelace matters for compute-bound workloads like fine-tuning, but for pure inference the bandwidth numbers tell the whole story. Bought used, the RTX 3090 Ti delivers the same inference performance as the RTX 4090 for far less money.
RTX 4090 — CUDA power users and new builds
Buy the RTX 4090 if you want the best long-term CUDA support, plan to fine-tune models (Flash Attention 2, bfloat16 native), or are building a new machine and want a warranty. For pure inference it is equal to the RTX 3090 Ti, so the premium is for the ecosystem and peace of mind. If inference per dollar is your metric, look elsewhere.
RTX 3090 Ti — best value pick
The most compelling option in the 24 GB tier. Identical 1008 GB/s bandwidth to the RTX 4090 means identical LLM inference speed at 60% of the price. The trade-off is buying used with no warranty. Inspect listings carefully, check for memory errors with CUDA tools before committing to heavy use. For inference-focused workloads, this is the answer.
RX 7900 XTX — AMD Linux users
If you are on Linux and comfortable with the AMD ecosystem, the RX 7900 XTX delivers 960 GB/s with excellent ROCm and Ollama support. Performance is slightly behind the RTX 4090 (50 vs 55 tok/s on 13B Q4) but the price is lower and the Linux experience is solid. Avoid if you are on Windows — ROCm Windows support is not mature enough for smooth LLM inference workflows.
RTX 3090 — tightest budget
On the used market the RTX 3090 is the cheapest path to 24 GB of CUDA VRAM. With 936 GB/s bandwidth it is only 7% slower than the RTX 4090 for inference — a difference most users will never notice in practice. The gap versus the RTX 3090 Ti is also small. If you find a good deal, the RTX 3090 is excellent value. As pricing converges with the RTX 3090 Ti, the Ti is almost always worth the small premium.
Ollama on Linux (NVIDIA)
Install Ollama with curl -fsSL https://ollama.com/install.sh | sh. It auto-detects NVIDIA GPUs via CUDA. Run ollama run qwen3:32b. That tag is about 21.2 GB at Q4_K_M, so a 24 GB card holds it at a short context. There is no current qwen3:34b tag. In LM Studio, download the Qwen3 32B Q4_K_M GGUF. Then run ollama ps and check that GPU% is 100.
Ollama on Linux (AMD RX 7900 XTX)
ROCm is included in Ollama's Linux package. After installing Ollama, verify AMD detection with ollama info — you should see the RX 7900 XTX listed. If not, install ROCm manually from AMD's repository, then reinstall Ollama.
Power and cooling for RTX 3090 / 3090 Ti
Both the RTX 3090 and RTX 3090 Ti draw up to 350-450 W under full load. Ensure your PSU has at least 850 W headroom, and that your case has adequate airflow for sustained inference workloads. The cards run hot under LLM load — a 3-slot cooler or aftermarket cooler helps significantly.
Testing a used GPU before committing
Before intensive use, run nvidia-smi -q to check for ECC errors and temperature history. For a deeper VRAM test, use cuda-memcheck or run a full model load and verify token output quality. Allocate a few days of burn-in before trusting a used card for production inference.
In 2026 the best 24 GB GPU depends on your budget and priorities. For maximum performance and CUDA compatibility, the RTX 4090 24 GB delivers 1008 GB/s bandwidth and ~55 tok/s on 13B Q4. For the best value, the RTX 3090 Ti 24 GB matches the RTX 4090 bandwidth at 1008 GB/s for 60% less. For AMD Linux users, the RX 7900 XTX 24 GB offers excellent ROCm support at 960 GB/s.
24 GB fits a 32B model at Q4_K_M (21.2 GB) and a 34B model at Q4_K_M (22.4 GB). A 16 GB Blackwell card does not: 27B Q4_K_M is 18.2 GB. 13B at Q8 is 17.6 GB and fits 24 GB. 70B at Q4_K_M is 44 GB and does not. The Ollama tag is ollama run qwen3:32b.
Yes. The RTX 3090 24 GB bought used delivers 936 GB/s bandwidth and approximately 48 tok/s on 13B Q4 — only 7% slower than the RTX 4090 at a fraction of the price. For LLM inference the limiting factor is memory bandwidth, not compute architecture, so the generation gap matters less than it would for gaming. If budget is the primary concern, the RTX 3090 is hard to beat.
A 70B model at Q4_K_M is 70 × 0.5 × 1.2 + 2 = 44 GB. On a 24 GB GPU the layers that do not fit are processed on the CPU, and tokens per second collapse. Two 24 GB cards (48 GB) can hold the 44 GB estimate. A single 24 GB card cannot.
The RTX 3090 Ti is the better value. Both cards have 1008 GB/s memory bandwidth, which is the primary driver of LLM inference speed. The RTX 3090 Ti achieves the same ~55 tok/s on 13B Q4 as the RTX 4090 while costing far less when bought used. The RTX 4090 is worth the premium only if you need newer CUDA features, plan to fine-tune models, or want to buy new with a warranty.
Check exact VRAM requirements or compare any two GPUs side by side.
VRAM and tokens-per-second figures on this page are synthesised from open community benchmarks. The sitewide formula and the full source list are on the methodology page. For this guide I leaned on:
Spot a number that does not match the linked source? Contact us and I will update the guide.