GeForce RTX 5090, 32GB
1.8TB/s of memory bandwidth, and 31GB of GDDR7 memory left for a model once the display buffer and the CUDA context have taken theirs. It holds 3 of the 11 models we track.
This is a card, not a computer. Every price on this page is the card plus $1,000 for the minimum host machine around it, so the comparison with a whole machine is fair.
NVIDIA's MSRP is $1,999 and no card sells at it. The price here is the US street median on 27 Aug 2026, plus $1,000 for the minimum host machine, because a card alone runs nothing.
Offers
checked 9 hours ago- NVIDIAout of stockchecked 9 hours ago$5,699.99
Timing
Collecting price history — 1 of 90 days. No timing verdict yet.
Models it runs — 8K context
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Won't fit— headroom
- Qwen3-32BQ6_KTight1GB headroomest. 44.4 tok/s
- Gemma-3-27BQ6_KFits5GB headroomest. 52.6 tok/s
- gpt-oss-20bQ8_0Fits7GB headroomest. 234 tok/s
| Model | Best quant | Fit | Headroom | Est. speed |
|---|---|---|---|---|
| Llama-4-Behemoth | — | Won't fit | — | — |
| DeepSeek-V3 | — | Won't fit | — | — |
| Llama-4-Maverick-17B-128E | — | Won't fit | — | — |
| Qwen3-235B-A22B | — | Won't fit | — | — |
| Mistral-Large-2411 | — | Won't fit | — | — |
| gpt-oss-120b | — | Won't fit | — | — |
| Llama-4-Scout-17B-16E | — | Won't fit | — | — |
| Llama-3.3-70B | — | Won't fit | — | — |
| Qwen3-32B | Q6_K | Tight | 1GB | 44.4 tok/s |
| Gemma-3-27B | Q6_K | Fits | 5GB | 52.6 tok/s |
| gpt-oss-20b | Q8_0 | Fits | 7GB | 234 tok/s |
Best quant is the highest-quality one that still loads with headroom. Every speed is an estimate; the formula is here.