Friday, October 9, 2026 — nxpc Ultimate Local LLM Box edition
I was going to open this build guide with a clean little parts list. Then I checked Amazon pricing.
The ASUS ROG Astral RTX 5090 — NVIDIA's 32GB GDDR7 halo card — is listed at $6,629.99 USD on Amazon US right now. Across the border on Amazon.ca, the same card starts at $9,299.99 CAD, with an AORUS Master ICE variant at $13,499 CAD and a PNY Triple Fan listed at a frankly deranged $19,829.07 CAD. The 2026 VRAM shortage hasn't just nudged prices; it has lobbed them into low orbit.
Meanwhile, sitting quietly in the same Amazon search results, is the ASRock Intel Arc Pro B60 Creator: 24GB of VRAM for $649.99 USD — with 200+ bought in the past month. Do the math on $/GB and it isn't a contest: the B60 is roughly $27/GB, the 5090 at these street prices is north of $207/GB.
So today's question isn't "which single GPU is fastest." It's: what actually builds the ultimate local LLM box in October 2026? I dug through ServeTheHome's inference data, Level1Techs forum benchmarks, Alex Ziskind's and Techno Tim's latest builds, and pulled live dual-market pricing (USD + CAD, because Steve insists and he's right). Buckle up.
Here's the rule that governs everything: every token you generate requires reading the entire model out of memory once. That means your token-generation ceiling is roughly memory bandwidth ÷ model size. Quantization and frameworks are multipliers on top — but the bandwidth number sets the wall.

The bandwidth hierarchy in 2026:
| GPU | VRAM | Memory Bandwidth | Notes |
|---|---|---|---|
| RTX 5090 | 32GB GDDR7 | 1,792 GB/s | The bandwidth king (BIZON, InsiderLLM) |
| RX 7900 XTX | 24GB GDDR6 | ~960 GB/s | ~62% of 5090's effective baseline |
| RTX 3090 | 24GB GDDR6X | 936 GB/s | Still the homelab workhorse |
| Arc Pro B60 | 24GB GDDR6 | 456 GB/s | Half the speed, quarter the price |
InsiderLLM's 2026 comparison puts it perfectly: the 5090's 1,792 GB/s "explains every result," and the 3090's 936 GB/s is why its baseline lands around 62% of the 5090's. And before anyone says "just buy a 4090" — the r/LocalLLaMA crowd benchmarked it: a 4090 is only ~10% faster than a 3090 for single-user inference, because memory bandwidth is nearly identical.
Pulling together measured results from across the 2026 benchmark landscape:
| GPU (VRAM) | Model / Quant | Token Generation | Source |
|---|---|---|---|
| RTX 5090 (32GB) | ~30B-class Q4, single card | Fastest single-GPU tier; 32B Q4 fits on one card | BIZON, OpenClawDC |
| RTX 3090 (24GB) | Llama 8B Q4 | ~87 tok/s | GPU Hunter |
| RTX 3090 (24GB) | 14B @ 16k context | 52.1 tok/s avg | Hardware Corner |
| Dual RTX 4090 (48GB) | Llama 3.3 70B Q4 | ~100 tok/s, 5–10% multi-GPU penalty | PromptQuorum |
| RX 7900 XTX (24GB) | Llama 8B Q4 | ~66 tok/s | GPU Hunter |
| RX 7900 XTX (24GB) | Llama 3.1 8B | ~96 tok/s (~75% of a 4090) | LocalAI Master |
| RX 7900 XTX (24GB) | Qwen3-Coder 32B Q4_K_M (ROCm 6.4) | 92 tok/s | BestLLMfor |
| RX 9070 XT (16GB) | Llama 8B Q4 | ~56 tok/s | GPU Hunter |
| RX 9070 XT (16GB) | gpt-oss:20b | ~92 tok/s | LocalAI Master |
| 4× Arc Pro B60 (96GB) | Qwen3-Coder-30B | 82 tok/s | ninabot.ch measured, CC BY 4.0 |
| 4× Arc Pro B60 (96GB) | 3 models, 130B params resident | 119 tok/s aggregate @ 211W | ninabot.ch measured |
Two big takeaways:
One: the B60 punches far above its bandwidth. That ninabot.ch dataset is the sleeper hit of the year — three MoE models resident across 96GB of pooled B60 VRAM, 82 tok/s on Qwen3-Coder-30B, all while sipping 211W. The card runs ~25–30% below a 3090 in raw throughput (GIGAGPU attributes that directly to the 456 vs 936 GB/s gap), but nothing else touches its capacity-per-dollar. And the official vLLM blog confirms B60 support with linear throughput scaling from 16 to 64 concurrent requests — this is a serving card, not a toy.
Two: multi-GPU pools VRAM, not speed. The Level1Techs and r/LocalLLaMA wisdom holds — stacking cards mostly buys capacity so bigger models fit, not linear token speed. Dual 3090s over NVLink get ~112.5 GB/s of interconnect for true pooling (Compute Market), and vLLM's tensor parallelism or Ollama's auto-split handles the layer sharding. Expect a 5–10% penalty versus a hypothetical single card with the same total bandwidth.

Newegg's RTX 3090 listings start at $899.99 USD (options from $899.99–$1,395.00), the used market sits around $820 according to Alibaba's August 2026 price guide, and Amazon Renewed has cards like the MSI Ventus 3X at $1,719.99 if you want returns-backed peace of mind. Two of the cheap ones = 48GB of VRAM for under ~$1,800, enough for Llama 3.3 70B at Q4 with ~87 tok/s single-card performance on 8B models. ServeTheHome's classic 3090 compute review still holds: inference barely taxes these cores — the 24GB of memory is the whole point. One caveat from the 2026 build guides: pair it with a 1000W-class PSU, because two 350W boards under load is real heat.

Two ASRock B60 Creator cards at $649.99 each = 48GB of brand-new, blower-cooled, ECC GDDR6 for $1,299.98 USD. You trade raw speed (456 GB/s per card) for capacity, silent blowers, sane 200W power draw, and warranty coverage the used-3090 route can't match. Run llama.cpp's SYCL backend or IPEX-LLM, or go vLLM if you're serving multiple users. The Level1Techs forum has an active B60 benchmark thread with owners pushing these cards through LLM workloads and SR-IOV experiments. For a home AI server that runs 24/7 under your desk, this is my pick — and yes, four of them (96GB) for ~$2,600 still undercuts one 5090 by three grand.
If you want the fastest single-GPU inference on the consumer market, the 5090 remains untouchable: 32GB GDDR7 at 1,792 GB/s handles 32B models at Q4 on one card, and two of them do 70B at Q8 quality. Alex Ziskind's August build went even further — dual RTX Pro 6000s and a terrified power supply. But at $6,629.99 USD / $9,299.99+ CAD street, you're paying a ~4x premium over the B60 route for roughly 4x the bandwidth and less than half the total VRAM. If you have the budget, it's glorious. If you have sense, it hurts.
Here's my favorite part, and it costs $0. Alex Ziskind's January video — "Your local LLM is 10x slower than it should be" — documents taking a setup from ~120 tok/s to 1200+ tok/s without touching hardware. The 2026 consensus tuning checklist, distilled:
num_ctx — oversized context silently eats VRAM and speed. Set what you need.ollama ps — never assume the model is actually on the GPU.Bonus: Ollama 0.30's updated GGUF/llama.cpp backends reportedly deliver up to 20% faster performance over prior versions. If you haven't updated since summer, that's a free upgrade waiting in your terminal. And the AMD crowd has real reason to smile in 2026: ROCm 7.2 officially supports the RX 9070 XT (gfx1201), and Ollama's llama.cpp HIP backend means RX 7900 XTX and 9070 XT owners just run ollama run and it works — genuinely fine for inference, per IDFS AI's deep dive.
Signature nxpc dual-market table — USD and CAD, pulled live this morning:
| Card | VRAM | Amazon US (USD) | Amazon CA (CAD) | Notes |
|---|---|---|---|---|
| ASRock Arc Pro B60 Creator | 24GB | $649.99 | $1,052.16 (from $939.43) | 200+/mo sold; also $649.99 on Newegg |
| RX 7900 XTX (ASRock Phantom) | 24GB | $1,129.99 | $3,306.39 (ASRock PG) | Best AMD capacity card |
| RX 9070 XT (ASUS Prime) | 16GB | $829.00 (from $672.19) | $1,404.18 (ASRock TC) | ROCm 7.2 official |
| RTX 5090 (ASUS ROG Astral OC) | 32GB | $6,629.99 | from $9,299.99 | MSI Vanguard $8,997 USD |
| RTX 5090 (AORUS Master) | 32GB | — | $11,255–$13,499 | Master ICE at $13,499 CAD |
| RTX 3090 (used, Newegg) | 24GB | from $899.99 | — | Used market ~$820 |
| RTX 3090 (Amazon Renewed, MSI Ventus) | 24GB | $1,719.99 | — | 50+ bought/mo |
Note the B60 CAD quirk: importers list third-party variants (GUNNIR, MAXSUN, Sparkle) at $1,207–$1,550 CAD, but the ASRock card's marketplace offers start at $939.43 CAD. Canadians, watch that offer list like a hawk — it moves.
Who should buy what:
The 2026 meta is beautiful and strange: the "ultimate" local LLM box isn't one flagship card anymore — it's capacity engineering. And right now, capacity is cheapest on a card with an Intel logo.