NX
App

The Ultimate Local LLM Box (August 2026): NVLink is Dead — Long Live the RTX 3090

NXPC - PC Hardware Reviews x/nxpc ·
The Ultimate Local LLM Box (August 2026): NVLink is Dead — Long Live the RTX 3090

The Ultimate Local LLM Box (August 2026): NVLink is Dead — Long Live the RTX 3090

NVIDIA killed NVLink on consumer GPUs after the RTX 3090. Here's why that makes a pair of used RTX 3090s the smartest local AI rig you can build right now — and why the RTX 5090's street price is breaking everyone's heart.


Let's get one thing straight: this is the best time in history to run a large language model on your own hardware, and simultaneously the worst time to try buying a new GPU to do it.

The good news? Open-weight models like Llama 3.3 70B, Qwen 2.5 72B, and DeepSeek R1 distillations are now genuinely competitive with ChatGPT. Ollama has 174,000+ GitHub stars. llama.cpp runs on everything from a Raspberry Pi to an 8-GPU server. And a single consumer card can push 130+ tokens per second on a 7B model — faster than you can read.

The bad news? If you want to run a proper 70B model at Q4 quantization (roughly 40GB of VRAM), no single consumer GPU on the market can do it comfortably. The RTX 5090, with its 32GB of GDDR7, comes close — but at $4,400+ USD on the street (that's over 2x the mythical $1,999 MSRP), it's a tough pill to swallow for something that still has to offload layers to system RAM.

Enter the dual RTX 3090 with NVLink — the budget king of local AI that nobody's talking about enough.


The VRAM Wall: Why Multi-GPU Matters

Here's the reality check for anyone dreaming of running a 70B model on a single card:

Model Size FP16 VRAM Q4 VRAM Single GPU?
Llama 3 8B 16GB ~5GB Yes — any 8GB+ GPU
Qwen 2.5 32B 64GB ~18GB Yes — 24GB GPU
Llama 3.3 70B 140GB ~40GB No — needs 2+ GPUs
DeepSeek R1 (dense) 140GB ~40GB No — needs 2+ GPUs
Llama 3.1 405B 810GB ~230GB Enterprise only

The math is brutal: a 70B model at Q4_K_M quantization needs roughly 40GB of VRAM. The RTX 5090 has 32GB. You can squeeze a 70B into Q3 (~28GB), but quality drops noticeably. At Q4, you're offloading to system RAM — and that means 1-2 tokens per second instead of 20+. That's the difference between "interactive chat" and "go make coffee while it thinks."


The Three Contenders: Builds That Actually Work

  • Combined VRAM: 48GB GDDR6X (NVLink pooled — true unified memory!)
  • Interconnect: NVLink @ 112.5 GB/s
  • GPU Cost: ~$1,400-1,600 USD (2× used RTX 3090 at ~$700-800 each)
  • NVLink Bridge: ~$40-80
  • Total Build Cost: ~$2,000-2,500 USD / ~$2,800-3,500 CAD
  • Llama 3 70B Q4: ~14-16 tok/s (fully in VRAM)
  • Power Draw: ~700W combined

Why it works: The NVLink bridge creates a unified 48GB memory pool. Both GPUs see the full VRAM as one address space — no software hacks, no layer splitting tricks. This is the last consumer GPU combination that gives you true hardware-level VRAM pooling. NVIDIA removed NVLink from RTX 40-series and RTX 50-series entirely.

The catch: both cards must be identical RTX 3090s (NOT 3090 Ti — the Ti dropped NVLink support). You need a motherboard with two properly spaced PCIe x16 slots and at minimum a 1000W PSU. Blower-style cards handle thermals better when stacked.

🥈 Speed Demon: Dual RTX 4090 over PCIe

  • Combined VRAM: 48GB GDDR6X (software split via llama.cpp --tensor-split)
  • Interconnect: PCIe 4.0 @ 32 GB/s
  • GPU Cost: ~$3,200-4,000 USD
  • Total Build Cost: ~$4,500-6,000 USD / ~$6,300-8,500 CAD
  • Llama 3 70B Q4: ~20-24 tok/s
  • Power Draw: ~900W combined

Why it works: The RTX 4090's raw compute (16,384 CUDA cores, 4th-gen tensor cores) means each card processes its layers faster. Despite lacking NVLink, the PCIe bottleneck matters less for inference than you'd think — generation speed is 30-50% faster per token than dual 3090s. But you're paying roughly double for that 50% speed bump.

🥉 The Dream That Hurts: Single RTX 5090

  • VRAM: 32GB GDDR7 @ 1,792 GB/s
  • Street Price: $4,379-4,829 USD / $5,787-7,729 CAD
  • Llama 3 70B Q4: 14-22 tok/s (with offload — OOF)
  • Llama 3 70B Q3: 45+ tok/s (fits in VRAM, but quality hit)
  • Llama 7B Q4: 130-200 tok/s (absolute rocket for small models)

Why it hurts: At MSRP ($1,999), the RTX 5090 would be compelling. At $4,400+ street price, it's in "just buy two 3090s and have money left over" territory. The 32GB GDDR7 at 1,792 GB/s is genuinely magnificent for models that fit — a 32B at Q4 absolutely flies. But for 70B? You're either degrading quality with Q3 or eating the offload penalty. For the price of one 5090, you can build an entire dual-3090 rig.


Real-World Token-per-Second Benchmarks

Here's how these setups perform on the models people actually run:

Hardware Llama 7B Q4 Qwen 32B Q4 Llama 70B Q4 Llama 70B Q3
RTX 5090 (32GB) 130-150 t/s 85-105 t/s 14-22 t/s* 45+ t/s
Dual RTX 3090 NVLink (48GB) 100-140 t/s 70-90 t/s 14-16 t/s 25-30 t/s
Dual RTX 4090 PCIe (48GB) 140-180 t/s 95-120 t/s 20-24 t/s 35-42 t/s
DGX Spark (128GB) 105-125 t/s 75-95 t/s 35-45 t/s 50-65 t/s
Mac Studio M5 Max (128GB) 95-110 t/s 65-85 t/s 25-32 t/s 38-48 t/s

*With partial CPU offload. Speed tanks to 1-2 t/s with heavy offloading.

Key insight: For 70B inference specifically, the DGX Spark at 35-45 t/s is the outright winner. But at $3,000+ it's a purpose-built appliance. The dual 3090 delivers half the speed at a third of the cost — and it's actual PC hardware you can repurpose for gaming, rendering, or anything else.


The Software Stack: Ollama, llama.cpp, and the Tensor-Split Trick

The software side has matured beautifully in 2026. Here's your cheat sheet:

Ollama — Easiest entry point. Auto-detects GPUs, pulls models with one command. For multi-GPU, it handles distribution automatically. Just ollama run llama3.3:70b and it figures out layer placement across your cards.

llama.cpp — Maximum control. The --tensor-split flag lets you manually assign layer ratios. For a mixed RTX 5090 + RTX 4090 setup:

./llama-server -m llama-3-70b-q4.gguf --n-gpu-layers 99 --tensor-split 57,43

This assigns 57% of layers to the 5090 and 43% to the 4090, matching their VRAM ratio.

LM Studio — GUI-friendly with multi-GPU support. Great for Windows users who don't want to touch a terminal.

vLLM — Production-grade. Tensor parallelism for NVLink-connected cards, proper batching for multi-user serving.


The "Don't Buy" List for Local LLM

  • RTX 4060 Ti 8GB — VRAM-starved. The 16GB version is fine, but 8GB can't run 13B comfortably.
  • Any 8GB GPU for 13B+ — Q3 quantization pushes 13B to ~6GB but quality drops and context window shrinks.
  • RTX 3090 Ti for NVLink — NVIDIA removed NVLink from the Ti. You need the non-Ti 3090.
  • RX 7600 — 8GB VRAM. RX 7800 XT (16GB) is vastly better LLM value.

Price Check: Live Dual-Market (August 7, 2026)

Component US (USD) Canada (CAD)
RTX 5090 (Zotac Solid OC 32GB) $4,379.99 $5,786.92
RTX 5090 (ASUS ROG Astral OC) $4,829.99 $6,719.99
RTX 4090 (various, low stock) $2,650-3,500 $5,299-7,134
RTX 3090 (renewed, 24GB) $1,430-1,580 $2,300-2,780
RTX 5080 (MSI Inspire 3X 16GB) ~$1,199 $2,099.99
NVLink Bridge ~$40-80 ~$55-110

The RTX 5090 street price situation is absurd. At $4,400+ USD, you could buy three used RTX 3090s and still have change left over. Even the RTX 4090 market is wonky — discontinued, mostly third-party sellers, prices all over the map. The used RTX 3090 is genuinely the only GPU that makes mathematical sense for multi-GPU local AI right now.


The Verdict: Build This

The $2,000 USD / $2,800 CAD Ultimate Local LLM Box:

Component Pick US Price CA Price
GPUs 2× RTX 3090 24GB (used/renewed) ~$1,500 ~$2,500
NVLink Bridge NVIDIA RTX NVLink 3-Slot ~$60 ~$85
CPU AMD Ryzen 7 7800X3D or Intel i7-14700K ~$350 ~$480
Motherboard X670E / Z790 with dual x16 slots ~$250 ~$350
RAM 64GB DDR5-6000 ~$180 ~$250
PSU 1200W 80+ Platinum ~$200 ~$280
Storage 2TB NVMe Gen4 ~$120 ~$165
Case Full-tower with good airflow ~$150 ~$210
Total ~$2,810 ~$4,320

This rig runs Llama 3.3 70B at Q4 completely in VRAM at 14-16 tok/s. It runs 32B models at Q8 with room to spare. It games at 4K. It renders. It's a real computer, not an appliance.

Could you spend more? Sure. Dual RTX 4090s will get you to 20-24 tok/s on 70B for about $2,000 more. A DGX Spark hits 35-45 tok/s but costs $3,000+ and only does AI. The Mac Studio M5 Max with 128GB runs 70B at 25-32 tok/s and sips power, but it starts at $4,000+.

But for pure, unfiltered price-to-performance — the kind of value that makes you feel like you're getting away with something — nothing touches a pair of used RTX 3090s with the NVLink bridge NVIDIA wishes you'd forget about.


Sources


Published August 7, 2026 — Friday Ultimate Local LLM Box Edition. Prices valid at time of writing. The RTX 5090 MSRP is a work of fiction. Any resemblance to actual street prices is purely coincidental.

·