The used RTX 3090 isn't just still alive in 2026 — it's the most cost-effective way to run 70B-class models at home. Here's what a quad-3090 build actually delivers, and why it might be smarter than a single RTX 5090.
Let's set the scene. It's August 2026. Nvidia just dropped $96 billion in quarterly revenue — yes, billion with a B — and forecast 70% growth next year. Jensen Huang's net worth just leapfrogged both Mark Zuckerberg and Larry Ellison in a single afternoon. The AI boom isn't slowing down. But here's the thing: while the cloud is printing money for Nvidia, the local AI community is quietly building something way more interesting in basements and home labs across the world.
I'm talking about the quad RTX 3090 build. Four cards. Ninety-six gigabytes of VRAM. Less money than a single scalped RTX 5090. And it runs Llama 3.1 70B like butter.
Three things just collided to make this the perfect moment for a quad-3090 build:
RTX 5090 prices are still insane. The cheapest 5090 on Amazon US right now is a GIGABYTE WINDFORCE at $4,488 USD. Most AIB cards sit between $5,000-$5,500. In Canada? The ASUS ROG Astral starts at $7,048 CAD. For 32GB of VRAM.
Used RTX 3090s have cratered to $600-800. That's $25-33 per gigabyte of VRAM — versus the 5090's $140+/GB. The math isn't subtle.
Multi-GPU inference stacks have matured massively. Ollama does multi-GPU out of the box. llama.cpp has --tensor-split. EXL2 handles tensor parallelism cleanly. The "it's janky" excuse doesn't fly anymore.
Here's what a serious quad-3090 inference rig looks like in late August 2026:
| Component | Choice | USD Price | CAD Price |
|---|---|---|---|
| GPUs | 4x Used RTX 3090 24GB | $2,800 | $3,800 |
| CPU | AMD Threadripper 3960X (used) | $500 | $680 |
| Motherboard | TRX40 (used) | $400 | $540 |
| RAM | 128GB DDR4 ECC (8x16GB) | $300 | $410 |
| PSU | 2x Corsair RM1000x (dual PSU) | $380 | $520 |
| Storage | 2TB NVMe Gen4 | $120 | $160 |
| Chassis | Mining frame or open bench | $80 | $110 |
| Cooling | Riser cables + fans | $120 | $160 |
| TOTAL | ~$4,700 USD | ~$6,380 CAD |
That's $4,700 USD — roughly the price of a single mid-tier RTX 5090 AIB card. And you get 96GB of total VRAM versus 32GB.
Simple: PCIe lanes. A consumer platform (AM5 or LGA1851) gives you 24-28 lanes. Four GPUs at x8 each need 32 lanes minimum. Threadripper gives you 64-88 lanes. It's not about CPU performance (LLM inference barely touches the CPU) — it's about feeding four GPUs without bottlenecking.
You can also go the EPYC route. A used EPYC 7302 + Supermicro H11SSL-i board runs about $450-550 total and gives you 128 PCIe 4.0 lanes. That's the home-lab sweet spot.
I've compiled real numbers from Hardware Corner, Compute Market, Local AI Master, Puget Systems, and the r/LocalLLaMA community. No synthetic fluff — these are llama.cpp and Ollama numbers measured by actual humans.
| GPU Configuration | Total VRAM | Model | Quant | Context | Token Gen (t/s) |
|---|---|---|---|---|---|
| 1x RTX 5090 | 32 GB | Qwen3 14B | Q4_K | 16k | 102.7 |
| 1x RTX 5090 | 32 GB | Qwen3 32B | Q4_K | 32k | 43.8 |
| 1x RTX 5090 | 32 GB | Qwen3.5 35B | MXFP4 | 256k | 97.3 |
| 1x RTX 4090 | 24 GB | Qwen3 14B | Q4_K | 16k | ~79 |
| 1x RTX 3090 | 24 GB | Qwen3 14B | Q4_K | 16k | ~52 |
| 2x RTX 3090 | 48 GB | Llama 3.1 70B | Q4 | 4k | 18-22 |
| 4x RTX 3090 | 96 GB | Llama 3.1 70B | Q4 | 32k | 35-40 |
| 4x RTX 3090 | 96 GB | Llama 3.1 405B | Q2 | 4k | 8-12 |
| 1x RX 7900 XTX | 24 GB | Llama 3.1 8B | Q4_K | 4k | ~96 |
| 2x RX 7900 XTX | 48 GB | Llama 3.1 70B | Q4 | 4k | ~12-15 |
| 1x Arc Pro B70 | 32 GB | Gemma 4 9B | Q4_K | 4k | ~54 |
| 4x Arc Pro B70 | 128 GB | Llama 3.1 8B | - | batch | ~12,000 (batch) |
Look at that dual 3090 row. $1,200-1,600 in GPUs and you're running Llama 3.1 70B at 18-22 tokens per second — perfectly usable chat speed. The quad setup at 35-40 t/s on 70B? That's API-level responsiveness, locally, with zero recurring cost.
And yes — you can technically run Llama 3.1 405B on quad 3090s at Q2 quantization. It's 8-12 t/s, which is slow, but it works. In 2024 that would have cost $40,000+ in enterprise hardware. Today? About $4,700.
Great question. Here's the honest comparison:
| Quad RTX 3090 | Single RTX 5090 | |
|---|---|---|
| Total VRAM | 96 GB | 32 GB |
| Max model (Q4) | Llama 70B comfortably, 405B at Q2 | Qwen 32B max |
| 14B speed | ~50 t/s (single card) | ~103 t/s |
| 70B speed | ~35-40 t/s | Can't run it fully |
| Power draw | ~1,400W total | ~575W |
| Noise | LOUD (4 blower cards) | Manageable |
| Setup complexity | High (PCIe bifurcation, risers, dual PSU) | Low (plug and play) |
| Cost USD | ~$4,700 total build | ~$5,000 GPU alone |
| Driver stability | Mature (Ampere, CUDA 12.x) | Mature (Blackwell) |
| Resale value | Already depreciated, stable | Will depreciate significantly |
The 5090 wins on single-card simplicity and raw speed on models that fit in 32GB. It's the right choice if you're doing agentic workflows with Qwen3 32B or running MXFP4-quantized models (where the 5090's NVFP4 hardware acceleration shines at 97 t/s on 35B models).
But if your goal is running large models — 70B and above — fully in VRAM, the quad 3090 is in a completely different league. The 5090 literally cannot run a 70B model at any usable quantization without offloading to system RAM, which drops you to 1-2 t/s. That's not inference — that's meditation.
AMD RX 7900 XTX (24GB): At ~$1,200 USD used, a dual XTX setup gives you 48GB for ~$2,400. ROCm 7.2 (June 2026) is genuinely good now — unified Windows+Linux installer, official RDNA 3 support, and FlashAttention-2 ported. Performance is solid (~96 t/s on 8B models). The catch? Tensor parallelism across AMD cards is still rougher than CUDA. You'll get 70B running across 2 XTXs, but expect 12-15 t/s and more debugging. If you're a Linux native and love tinkering, it's viable. If you want it to "just work," CUDA still wins.
Intel Arc Pro B70 (32GB): The $949 wildcard. Alex Ziskind did a great video on this — 32GB VRAM at under $1,000 is objectively compelling. Puget Systems benchmarked four B70s at nearly 12,000 tok/s on Llama 3.1 8B in batch. But for interactive single-user inference, the SYCL stack drags: ~54 t/s on a 9B dense model, and MoE models actually run slower than dense ones due to dispatch inefficiencies. The hardware is ready. The software isn't — yet. Watch this space, but don't build your daily driver around it in August 2026.
If you're building this, here's what the software side looks like:
# Pull your model
ollama pull llama3.1:70b
# Ollama auto-detects all GPUs and splits across them
# No config needed for basic multi-GPU
# For llama.cpp with fine control:
./llama-cli \
-m llama-3.1-70b-q4_K_M.gguf \
-ngl 99 \
--tensor-split 24,24,24,24 \
-c 32768 \
--flash-attn \
-p "Explain quantum computing in simple terms"
The --tensor-split 24,24,24,24 tells llama.cpp to distribute layers evenly across all four 24GB cards. With EXL2 you get even better multi-GPU scaling through tensor parallelism.
One critical tip: use NVLink bridges between pairs of 3090s if your board supports it. It doesn't double bandwidth (3090 NVLink is ~112 GB/s bidirectional), but it reduces the PCIe bottleneck for tensor parallelism. Without NVLink, you're still fine — llama.cpp uses a layer-split strategy that minimizes cross-GPU communication.
Build the quad 3090 rig if:
Buy a single RTX 5090 if:
Me? I'd build the quad 3090 rig. There's something deeply satisfying about piecing together "obsolete" gaming cards into a machine that runs models the cloud charges hundreds per month to access. The 5090 is faster on paper — but the quad 3090 runs models the 5090 literally cannot. That's not a benchmark win. That's a capability unlock.
And in August 2026, with Nvidia's stock price reminding us daily who's really profiting from the AI boom, there's something poetic about building your own inference server from cards Jensen Huang would rather you'd forgotten about.
What's your local LLM setup? Running a single card, dual GPUs, or a full rack? Drop your tokens-per-second in the comments — I want to see those numbers.