NVIDIA killed NVLink on consumer GPUs after the RTX 3090. Here's why that makes a pair of used RTX 3090s the smartest local AI rig you can build right now — and why the RTX 5090's street price is breaking everyone's heart.
Let's get one thing straight: this is the best time in history to run a large language model on your own hardware, and simultaneously the worst time to try buying a new GPU to do it.
The good news? Open-weight models like Llama 3.3 70B, Qwen 2.5 72B, and DeepSeek R1 distillations are now genuinely competitive with ChatGPT. Ollama has 174,000+ GitHub stars. llama.cpp runs on everything from a Raspberry Pi to an 8-GPU server. And a single consumer card can push 130+ tokens per second on a 7B model — faster than you can read.
The bad news? If you want to run a proper 70B model at Q4 quantization (roughly 40GB of VRAM), no single consumer GPU on the market can do it comfortably. The RTX 5090, with its 32GB of GDDR7, comes close — but at $4,400+ USD on the street (that's over 2x the mythical $1,999 MSRP), it's a tough pill to swallow for something that still has to offload layers to system RAM.
Enter the dual RTX 3090 with NVLink — the budget king of local AI that nobody's talking about enough.
Here's the reality check for anyone dreaming of running a 70B model on a single card:
| Model Size | FP16 VRAM | Q4 VRAM | Single GPU? |
|---|---|---|---|
| Llama 3 8B | 16GB | ~5GB | Yes — any 8GB+ GPU |
| Qwen 2.5 32B | 64GB | ~18GB | Yes — 24GB GPU |
| Llama 3.3 70B | 140GB | ~40GB | No — needs 2+ GPUs |
| DeepSeek R1 (dense) | 140GB | ~40GB | No — needs 2+ GPUs |
| Llama 3.1 405B | 810GB | ~230GB | Enterprise only |
The math is brutal: a 70B model at Q4_K_M quantization needs roughly 40GB of VRAM. The RTX 5090 has 32GB. You can squeeze a 70B into Q3 (~28GB), but quality drops noticeably. At Q4, you're offloading to system RAM — and that means 1-2 tokens per second instead of 20+. That's the difference between "interactive chat" and "go make coffee while it thinks."
Why it works: The NVLink bridge creates a unified 48GB memory pool. Both GPUs see the full VRAM as one address space — no software hacks, no layer splitting tricks. This is the last consumer GPU combination that gives you true hardware-level VRAM pooling. NVIDIA removed NVLink from RTX 40-series and RTX 50-series entirely.
The catch: both cards must be identical RTX 3090s (NOT 3090 Ti — the Ti dropped NVLink support). You need a motherboard with two properly spaced PCIe x16 slots and at minimum a 1000W PSU. Blower-style cards handle thermals better when stacked.
Why it works: The RTX 4090's raw compute (16,384 CUDA cores, 4th-gen tensor cores) means each card processes its layers faster. Despite lacking NVLink, the PCIe bottleneck matters less for inference than you'd think — generation speed is 30-50% faster per token than dual 3090s. But you're paying roughly double for that 50% speed bump.
Why it hurts: At MSRP ($1,999), the RTX 5090 would be compelling. At $4,400+ street price, it's in "just buy two 3090s and have money left over" territory. The 32GB GDDR7 at 1,792 GB/s is genuinely magnificent for models that fit — a 32B at Q4 absolutely flies. But for 70B? You're either degrading quality with Q3 or eating the offload penalty. For the price of one 5090, you can build an entire dual-3090 rig.
Here's how these setups perform on the models people actually run:
| Hardware | Llama 7B Q4 | Qwen 32B Q4 | Llama 70B Q4 | Llama 70B Q3 |
|---|---|---|---|---|
| RTX 5090 (32GB) | 130-150 t/s | 85-105 t/s | 14-22 t/s* | 45+ t/s |
| Dual RTX 3090 NVLink (48GB) | 100-140 t/s | 70-90 t/s | 14-16 t/s | 25-30 t/s |
| Dual RTX 4090 PCIe (48GB) | 140-180 t/s | 95-120 t/s | 20-24 t/s | 35-42 t/s |
| DGX Spark (128GB) | 105-125 t/s | 75-95 t/s | 35-45 t/s | 50-65 t/s |
| Mac Studio M5 Max (128GB) | 95-110 t/s | 65-85 t/s | 25-32 t/s | 38-48 t/s |
*With partial CPU offload. Speed tanks to 1-2 t/s with heavy offloading.
Key insight: For 70B inference specifically, the DGX Spark at 35-45 t/s is the outright winner. But at $3,000+ it's a purpose-built appliance. The dual 3090 delivers half the speed at a third of the cost — and it's actual PC hardware you can repurpose for gaming, rendering, or anything else.
The software side has matured beautifully in 2026. Here's your cheat sheet:
Ollama — Easiest entry point. Auto-detects GPUs, pulls models with one command. For multi-GPU, it handles distribution automatically. Just ollama run llama3.3:70b and it figures out layer placement across your cards.
llama.cpp — Maximum control. The --tensor-split flag lets you manually assign layer ratios. For a mixed RTX 5090 + RTX 4090 setup:
./llama-server -m llama-3-70b-q4.gguf --n-gpu-layers 99 --tensor-split 57,43
This assigns 57% of layers to the 5090 and 43% to the 4090, matching their VRAM ratio.
LM Studio — GUI-friendly with multi-GPU support. Great for Windows users who don't want to touch a terminal.
vLLM — Production-grade. Tensor parallelism for NVLink-connected cards, proper batching for multi-user serving.
| Component | US (USD) | Canada (CAD) |
|---|---|---|
| RTX 5090 (Zotac Solid OC 32GB) | $4,379.99 | $5,786.92 |
| RTX 5090 (ASUS ROG Astral OC) | $4,829.99 | $6,719.99 |
| RTX 4090 (various, low stock) | $2,650-3,500 | $5,299-7,134 |
| RTX 3090 (renewed, 24GB) | $1,430-1,580 | $2,300-2,780 |
| RTX 5080 (MSI Inspire 3X 16GB) | ~$1,199 | $2,099.99 |
| NVLink Bridge | ~$40-80 | ~$55-110 |
The RTX 5090 street price situation is absurd. At $4,400+ USD, you could buy three used RTX 3090s and still have change left over. Even the RTX 4090 market is wonky — discontinued, mostly third-party sellers, prices all over the map. The used RTX 3090 is genuinely the only GPU that makes mathematical sense for multi-GPU local AI right now.
The $2,000 USD / $2,800 CAD Ultimate Local LLM Box:
| Component | Pick | US Price | CA Price |
|---|---|---|---|
| GPUs | 2× RTX 3090 24GB (used/renewed) | ~$1,500 | ~$2,500 |
| NVLink Bridge | NVIDIA RTX NVLink 3-Slot | ~$60 | ~$85 |
| CPU | AMD Ryzen 7 7800X3D or Intel i7-14700K | ~$350 | ~$480 |
| Motherboard | X670E / Z790 with dual x16 slots | ~$250 | ~$350 |
| RAM | 64GB DDR5-6000 | ~$180 | ~$250 |
| PSU | 1200W 80+ Platinum | ~$200 | ~$280 |
| Storage | 2TB NVMe Gen4 | ~$120 | ~$165 |
| Case | Full-tower with good airflow | ~$150 | ~$210 |
| Total | ~$2,810 | ~$4,320 |
This rig runs Llama 3.3 70B at Q4 completely in VRAM at 14-16 tok/s. It runs 32B models at Q8 with room to spare. It games at 4K. It renders. It's a real computer, not an appliance.
Could you spend more? Sure. Dual RTX 4090s will get you to 20-24 tok/s on 70B for about $2,000 more. A DGX Spark hits 35-45 tok/s but costs $3,000+ and only does AI. The Mac Studio M5 Max with 128GB runs 70B at 25-32 tok/s and sips power, but it starts at $4,000+.
But for pure, unfiltered price-to-performance — the kind of value that makes you feel like you're getting away with something — nothing touches a pair of used RTX 3090s with the NVLink bridge NVIDIA wishes you'd forget about.
Published August 7, 2026 — Friday Ultimate Local LLM Box Edition. Prices valid at time of writing. The RTX 5090 MSRP is a work of fiction. Any resemblance to actual street prices is purely coincidental.