RTX 4090 vs RTX 5090: used value vs the new flagship

Editorial comparison updated 2026-07-26 · live eBay medians refresh daily · Methodology

Verdict: The 5090 is simply the better card: 32GB instead of 24GB, roughly 75% more memory bandwidth, and FP4 support. But street pricing keeps it well above MSRP, so it typically runs 50–75% more than a used 4090. The extra 8GB mainly buys headroom: 32B models at higher quants and much larger context windows. It still falls short of 70B territory. If you need what 32GB uniquely enables, nothing else consumer-grade does it; otherwise the used 4090, or a 3090 for pure value, is the rational buy.
RTX 4090RTX 5090
Fair used value$2,325 ($1,975–$2,750)$3,725 ($3,175–$4,400)
eBay median (filtered)$3,199 38% over fair$6,000 61% over fair
Memory24GB VRAM32GB VRAM
$ / GB (at fair value)$97$116
Power450W TDP575W TDP
Launch MSRP$1,599$1,999

Fair values are editorial estimates; eBay medians are filtered Buy-It-Now samples, not checkout prices. See how we compute these.

Affiliate disclosure: Affiliate (paid) links: we may earn a commission from eBay Partner Network and Amazon Associates at no extra cost to you. As an Amazon Associate I earn from qualifying purchases. Full disclosure.

What each can run locally

Model class (weights @ Q4)RTX 4090 · 24GBRTX 5090 · 32GB
7–8B @ Q4 (Llama 3.1 8B, Qwen3 8B)✓ fits✓ fits
14B @ Q4 (DeepSeek-R1 14B distill)✓ fits✓ fits
30B MoE @ Q4 (Qwen3-coder 30B)✓ fits✓ fits
32B dense @ Q4 (Qwen3 32B, R1 32B)✓ fits✓ fits
70B dense @ Q4 (Llama 3.3 70B)— too big— too big
100B+ MoE (gpt-oss-120b class)— too big— too big

Approximate weights-only footprints — long context and KV cache need additional memory headroom.

What 32GB actually unlocks

The jump from 24GB to 32GB lets you run 32B-class models at Q5/Q6 instead of Q4, hold dramatically more KV cache for long-context work, and batch larger image or video generation jobs. What it does not do is reach 70B dense models — those still need 40GB+, which means dual GPUs, workstation cards, or a big-memory Mac.

Speed: bandwidth is the story

For LLM inference the headline spec is memory bandwidth: GDDR7 pushes the 5090 to roughly 1.79 TB/s against the 4090’s 1.0 TB/s, and token generation scales close to linearly with it. Blackwell’s FP4 support adds another gear in engines that exploit it. This is a real generational gap, not a refresh.

Price reality

Check the live medians above: the 5090’s used market has stayed at or above launch MSRP since release, driven by AI demand. On dollars-per-GB it is consistently the worst value on our board — you are paying for the ceiling, not the capacity.

Budget for the platform too: 575W of board power wants a 1000W+ PSU, and the same 12VHPWR connector diligence applies as with the 4090 — inspect the plug before buying used.

The value ladder

The used NVIDIA ladder for local AI is straightforward: the 3090 is the cheapest 24GB, the 4090 is speed at 24GB, and the 5090 is the only consumer 32GB. Buy the rung that matches your models — see our RTX 3090 vs 4090 comparison if the 5090 premium is out of reach.

FAQ

Is the RTX 5090 worth it over the 4090 for local AI?

Only if you specifically benefit from 32GB or the bandwidth: 32B models at higher quants, very long context, or heavy generation workloads. For everything that fits in 24GB, the used 4090 delivers most of the experience at roughly two-thirds of the price.

Can the RTX 5090 run 70B models?

Not alone. 70B at Q4 needs 40GB+, so a single 5090 still requires CPU offload at painful speeds. Two 3090s, an RTX A6000, or a 64GB+ Mac Studio are the realistic 70B paths.

Will used 4090 prices drop now that the 5090 exists?

They have been sticky — AI demand keeps every 24GB card in demand. Watch the live median on this page against our fair value; our data refreshes daily.

Also consider

ModelMemoryFair valueeBay median
RTX 309024GB$1,200$1,799
RTX A600048GB$3,750$5,183

More comparisons