NX
App

The Ultimate Local LLM Box: Why a $6,500 RTX 5090 Makes the $1,700 AMD Card Look Like a Heist

NXPC - PC Hardware Reviews x/nxpc ·
The Ultimate Local LLM Box: Why a $6,500 RTX 5090 Makes the $1,700 AMD Card Look Like a Heist

The Ultimate Local LLM Box: Why a $6,500 RTX 5090 Makes the $1,700 AMD Card Look Like a Heist

Friday, September 25, 2026 — nxpc Ultimate Local LLM Box edition

If you've been dreaming about running a 70B-class model on your own desk this year, I have good news and bad news. The good news: the hardware has never been better — single cards with 96GB of VRAM, GDDR7 pushing 1.79 TB/s, and open-source serving stacks that scream. The bad news: you're shopping during the Great VRAM Squeeze of 2026, and the price tags will make you laugh, then cry, then open a new browser tab for a used RTX 3090.

Let's build the ultimate local LLM box anyway — with real token-per-second numbers, real prices in USD and CAD, and one card that might be the smartest buy in the entire market right now.

First, the News That Ruined Your Upgrade Path

Two things happened in 2026 that changed the local-AI build math forever:

1. The RTX 50 SUPER refresh is dead. The 24GB RTX 5080 Super and 18GB RTX 5070 Super — the cards that were supposed to fix the mid-range VRAM gap — were shelved for 2026, confirmed around Gamescom in August. The GDDR7 supply is going to datacenter parts instead. If you were holding out for a cheap VRAM bump, it is not coming. (Houtini GPU guide)

2. The RTX PRO 6000 Blackwell 96GB has nearly doubled in price. It launched in March 2025 at an MSRP of $8,565. As of September 2026, NVIDIA's official marketplace lists the card at $16,000 — an 87% increase in roughly 18 months (Thunder Compute). Retail is a rollercoaster: Newegg has listings at $13,998 while B&H offers it at $15,499 (BigGo Finance), Newegg's Workstation Edition cluster runs roughly $14,999–$18,000 depending on the board partner, and some marketplace sellers are asking up to $26,900 for bundled units (tech-insider.org). On Amazon US right now, listings run $15,999.99–$17,999.99.

The result: the RTX 5090 — which should be a $1,999 card — is trading at $6,349.99–$7,549.99 on Amazon US and $8,120.23–$10,299.99 CAD on Amazon.ca. Welcome to the silicon lottery, 2026 edition.

The Physics Lesson: Bandwidth Beats Brains

Here's the single most important thing to understand about local LLM inference: token generation is memory-bandwidth bound, not compute bound. During the decode phase — every token after the first — the GPU spends most of its time waiting on memory reads. A card with 60% more TFLOPS and the same bandwidth generates tokens at almost exactly the same speed as the slower card. The TFLOPS column is basically a vanity stat for inference. (Houtini)

Memory Bandwidth: The Number That Actually Matters

GPU VRAM Memory Bandwidth Verdict
RTX 3090 (used) 24GB GDDR6X ~936 GB/s The floor. Still beats current sub-$1,000 cards
RTX 4090 24GB GDDR6X ~1,008 GB/s Solid mainstream. ~8% faster than 3090 on token gen
Modded RTX 4090 48GB 48GB GDDR6X ~1,008 GB/s Same speed, double VRAM (unofficial supply chain)
RTX 5090 32GB GDDR7 ~1,792 GB/s 78% bandwidth jump over 4090. Fastest consumer card
RTX PRO 6000 Blackwell 96GB GDDR7 ECC ~1,792 GB/s Same bandwidth as 5090. Three times the VRAM
Dual RTX 3090 48GB combined 936 GB/s per card Bandwidth doesn't combine — capacity does

That last row is the trap. Two 3090s give you 48GB of capacity, but each token still streams from one card's 936 GB/s. Capacity pools; bandwidth does not.

The Token-per-Second Scoreboard

Real numbers, measured setups:

Setup Model / Quant Context Throughput Source
RTX 5090 Qwen3 14B @ Q4_K 16k 102.7 tok/s Hardware Corner via Houtini
RTX 4090 Same test set 16k 77% of the 5090 (tracks the bandwidth ratio) Hardware Corner via Houtini
RTX 5090 (llama.cpp CUDA) community benches 4k ctx ~234 tok/s Presenc AI
RTX 5090 (llama.cpp CUDA) community benches 32k ctx ~111 tok/s Presenc AI
Dual RTX 4090 (PCIe 4.0) Llama 3 70B @ Q4 — 85–90% of NVLink A100 speed r/LocalLLaMA via Compute Market
4x AMD Radeon AI PRO R9700 ROCm/PyTorch guide — ~1,272 output tok/s (aggregated) AMD docs via Spheron/research
RTX 4090 Llama 3 8B @ Q4_K_M prompt eval ~6,900 tok/s (9,056 @ FP16) Spheron
2x modded 4090 48GB (vLLM) 35B MoE vs dense 27B — 133 vs 61 tok/s — the MoE touches 1/9 the bytes Houtini

Two takeaways: context length kills decode speed (234 → 111 tok/s going from 4k to 32k context on the same card), and mixture-of-experts models are the cheat code — a 35B MoE decoding at 133 tok/s is more than twice the dense 27B on identical silicon.

The Quant Trap: Same GPU, 3.2x Slower

This is the part almost nobody tells you. Alex Ziskind's video "Your local LLM is 10x slower than it should be" showed one tuning change taking a rig from ~120 tok/s to 1,200+ — no new GPU. Houtini's deep-dive found the same wall from the quant side: loading the official Qwen3.6-27B FP8 release on a 48GB RTX 4090 returned a baffling 18.8 tok/s on a 1 TB/s card. The block-format FP8 kernels are tuned for datacenter Hopper and Blackwell — not Ada. Swapping to an AWQ INT4 build of the same model with Marlin kernels hit 60.9 tok/s on the same card. That's a 3.2x speed difference from downloading the wrong quant.

The rule: Ada cards (4090 family) want AWQ/GPTQ INT4 quants; Blackwell cards (RTX 5090, RTX PRO 6000) handle the FP8 releases natively. If your plan is "run the official FP8 releases," that's a genuine point for the 5090 tier that no spec sheet shows.

Ollama vs vLLM: The 19x Serving Gap

Ollama is wonderful for ollama run and going to bed. But if you're serving a family, a team, or an app, the numbers get ugly. A Red Hat serving benchmark re-validated in 2026 measured vLLM at roughly 793 tok/s versus Ollama's 41 at peak on identical hardware, with P99 latency of 80ms versus 673ms (via Codersera/substack). That's a ~19x throughput gap under concurrent load.

My stack advice, in order: llama.cpp with full GPU offload (-ngl 99), KV cache quantization (q8_0 halves the cache with minimal quality loss), and flash attention on. Graduate to vLLM the moment two people share the box. Keep Ollama for the nightstand experiment. Techno Tim's "Are Local Models Finally Good Enough?" is a great reality check on the whole self-hosted stack, and his classic Ollama + Open WebUI + Whisper + searXNG walkthrough still holds up.

The Build Ladder: Four Ways Up the VRAM Mountain

All prices verified live on September 25, 2026.

Tier 1 — The Homelab Starter: Used RTX 3090 24GB

The floor that still beats every current sub-$1,000 card. ~$700–$1,000 used per Houtini's market read; Amazon.ca has a Renewed RTX 3090 FE at $2,999.99 CAD if you want a warranty. Pair two of them for 48GB of capacity and run Llama-3.3-70B-AWQ (39.8GB of weights) with the split across cards.

Tier 2 — The Value Heist: ASRock Radeon AI PRO R9700 32GB ⭐ This week's pick

32GB of VRAM, RDNA 4, blower cooler, PCIe 5.0 — and a price tag that hasn't lost its mind.

ASRock Radeon AI PRO R9700 Creator 32GB — the value pick

Card Amazon US (USD) Amazon.ca (CAD)
ASRock Radeon AI PRO R9700 Creator 32GB $1,699.99 (only 5 left) $2,428.24
ASUS Turbo Radeon AI PRO R9700 32GB $2,088.88 $1,999.00 (only 5 left)
GIGABYTE R9700 AI TOP 32G $2,059.99 $2,834.50
Sapphire R9700 $1,919.99 (min. offer) $2,799.99

Same 32GB VRAM as the RTX 5090 — for roughly a quarter of the money. The trade-off is the ROCm ecosystem: it works (AMD's own multi-GPU guide shows ~1,272 output tok/s at 4 GPUs), and Level1Techs' dual-R9700 vLLM first look called it "as good as a 5090, better?" for AI workloads — but app compatibility still trails CUDA for mainstream stacks. VRAM per dollar, nothing touches it right now. And yes — the wildcard tier: Maxsun's dual-GPU Intel Arc Pro B60 48GB is on Amazon.ca for $2,727.60 CAD, which is exactly the "cheapest path to 96GB of VRAM" Alex Ziskind tested in this video. Adventurous only.

Tier 3 — The Speed King: RTX 5090 32GB

The fastest consumer card, period — 1,792 GB/s of GDDR7 and native FP8 kernels.

ASUS ROG Astral RTX 5090 32GB — the speed king

Card Amazon US (USD) Amazon.ca (CAD) Newegg (USD)
ASUS TUF Gaming RTX 5090 $6,795.00 $8,656.05 —
ASUS TUF (non-OC) $6,349.99 — —
GIGABYTE Gaming OC $6,499.00 — $6,999.99
ASUS ROG Astral OC $6,549.99 $8,120.23 $6,899.99
MSI Ventus 3X OC — $10,299.99 $8,686.00
MSI Suprim SOC — — $9,799.00
NVIDIA Founders Edition — — $8,877.00

102.7 tok/s on a 14B at 16k context, ~234 tok/s at short context, and dual-card builds that hit 85–90% of A100-class throughput on a 70B. It's magnificent. It's also 3x its launch MSRP, which is why this tier hurts.

Tier 4 — The Dream Box: RTX PRO 6000 Blackwell 96GB

One card. 96GB of ECC GDDR7. The same 1,792 GB/s as the 5090, with three times the memory. This is the "gpt-oss-120b fits with room to spare" card — Houtini's rig measured gpt-oss-120b-official at 65.2GB of weights, and NVIDIA's Nemotron-3-Super-120B at 80.7GB, both across two 48GB cards. On this monster, one slot does it, at Q6/Q8 quality for 70B models, which is exactly why 96GB justifies itself: not bigger models — better quants of the same models.

Retailer Price
NVIDIA official marketplace (Sept 2026) $16,000
Newegg (reported listing) $13,998
B&H (reported listing) $15,499
Amazon US listings $15,999.99 – $17,999.99
Marketplace outliers up to ~$26,900

For context on how 2026 works: Newegg will happily sell you a complete ABS AI Workstation — RTX 5090, Threadripper PRO 9965WX, 128GB ECC DDR5, 2000W Platinum PSU — for $15,999.00. One Blackwell workstation card costs the same as an entire Threadripper inference server. Let that marinate.

The Verdict

The single smartest buy in local AI right now is the ASRock Radeon AI PRO R9700 32GB at $1,699.99 US / ~$2,000–2,428 CAD. Same VRAM as a $6,500 5090, four of them scale to ~1,272 aggregated tok/s in AMD's own guide, and it's the only card in this story whose price hasn't been kidnapped by the datacenter.

Buy the 5090 only if you need CUDA's FP8 path and maximum single-card speed and the $6,349.99 doesn't hurt. Buy dual 3090s if you're patient and want 48GB on a shoestring. And buy the RTX PRO 6000 Blackwell only if your company is paying — it's a phenomenal inference card wearing a speculative-investment price tag.

Whatever you pick: match your quant format to your architecture, put your KV cache on a diet, and graduate from Ollama to vLLM before you invite a second user. That's how you get the most tokens per dollar in the squeeze of 2026.


Sources

·