RTX 4090 vs RTX 5090: used value vs the new flagship
| RTX 4090 | RTX 5090 | |
|---|---|---|
| Fair used value | $2,325 | $3,725 |
| eBay median (filtered) | $3,199 38% over fair | $6,000 61% over fair |
| Memory | 24GB VRAM | 32GB VRAM |
| $ / GB (at fair value) | $97 | $116 |
| Power | 450W TDP | 575W TDP |
| Launch MSRP | $1,599 | $1,999 |
Affiliate disclosure: Affiliate (paid) links: we may earn a commission from eBay Partner Network and Amazon Associates at no extra cost to you. As an Amazon Associate I earn from qualifying purchases. Full disclosure.
What each can run locally
| Model class (weights @ Q4) | RTX 4090 · 24GB | RTX 5090 · 32GB |
|---|---|---|
| 7–8B @ Q4 (Llama 3.1 8B, Qwen3 8B) | ✓ fits | ✓ fits |
| 14B @ Q4 (DeepSeek-R1 14B distill) | ✓ fits | ✓ fits |
| 30B MoE @ Q4 (Qwen3-coder 30B) | ✓ fits | ✓ fits |
| 32B dense @ Q4 (Qwen3 32B, R1 32B) | ✓ fits | ✓ fits |
| 70B dense @ Q4 (Llama 3.3 70B) | — too big | — too big |
| 100B+ MoE (gpt-oss-120b class) | — too big | — too big |
What 32GB actually unlocks
The jump from 24GB to 32GB lets you run 32B-class models at Q5/Q6 instead of Q4, hold dramatically more KV cache for long-context work, and batch larger image or video generation jobs. What it does not do is reach 70B dense models — those still need 40GB+, which means dual GPUs, workstation cards, or a big-memory Mac.
Speed: bandwidth is the story
For LLM inference the headline spec is memory bandwidth: GDDR7 pushes the 5090 to roughly 1.79 TB/s against the 4090’s 1.0 TB/s, and token generation scales close to linearly with it. Blackwell’s FP4 support adds another gear in engines that exploit it. This is a real generational gap, not a refresh.
Price reality
Check the live medians above: the 5090’s used market has stayed at or above launch MSRP since release, driven by AI demand. On dollars-per-GB it is consistently the worst value on our board — you are paying for the ceiling, not the capacity.
Budget for the platform too: 575W of board power wants a 1000W+ PSU, and the same 12VHPWR connector diligence applies as with the 4090 — inspect the plug before buying used.
The value ladder
The used NVIDIA ladder for local AI is straightforward: the 3090 is the cheapest 24GB, the 4090 is speed at 24GB, and the 5090 is the only consumer 32GB. Buy the rung that matches your models — see our RTX 3090 vs 4090 comparison if the 5090 premium is out of reach.
FAQ
Is the RTX 5090 worth it over the 4090 for local AI?
Only if you specifically benefit from 32GB or the bandwidth: 32B models at higher quants, very long context, or heavy generation workloads. For everything that fits in 24GB, the used 4090 delivers most of the experience at roughly two-thirds of the price.
Can the RTX 5090 run 70B models?
Not alone. 70B at Q4 needs 40GB+, so a single 5090 still requires CPU offload at painful speeds. Two 3090s, an RTX A6000, or a 64GB+ Mac Studio are the realistic 70B paths.
Will used 4090 prices drop now that the 5090 exists?
They have been sticky — AI demand keeps every 24GB card in demand. Watch the live median on this page against our fair value; our data refreshes daily.
Also consider
More comparisons
- RTX 3090 vs RTX 4090: which used card for local AI?
- RTX 3090 vs RX 7900 XTX: 24GB CUDA vs 24GB ROCm
- Mac Studio vs RTX 4090: which local AI box?
- RTX 3080 vs RTX 4070: cheap veteran or efficient successor?
- RTX 4080 vs RTX 4090: 16GB value or 24GB ceiling?
- RTX 5070 vs RTX 4070 Super: which 12GB card to buy used?