• A Great friend to the HardForum with a great kid that he is trying to get a scholorship to continue his schooling. Please give hime a vote! Only 24 hours left! Thanks.
    If you have an VOTE FOR KEENAN!

RTX 6000 Blackwell vs. 3x 5090?

uOpt

2[H]4U
Joined
Mar 29, 2006
Messages
2,906
What happens when you compare LLM inference running on a Blackwell RTX6000 versus 3x RTX 5090 in one machine?

Would it be faster due to more GPU units? Or slower due to communication overhead? Or is it "it depends"?

Of course power consumption of the trio would suck (no pun intended).
 
Personally I would go RTX 6000 Pro Max-Q.

You are memory bandwidth and latency bound so crossing the PCIe bus will not speed things up.

Also you can actually fit the RTX 6000 in a normal case, 3x 5090 will need a mining type rig and extension cables.
 
Personally I would go RTX 6000 Pro Max-Q.
I don't understand why people nerf themselves and pay more money for the Max-Q. If you can get the Workstation edition for cheaper - get it.

You can easily set the max watts of the Workstation (default 600w) to whatever you want (300w for the Max-Q). IIRC, there is something like a 30% performance gain between 300W and 450W. It doesn't scale linearly, but if you are paying that much for hardware, why nerf it?
 
3x 5090 will prefill faster than 1x Pro 6000 and decode about the same. With 4x 5090 you can try something like PP=2 TP=2 which should give you ~2x decode at the expense of some prefill, but it also requires a 4x PCI-e 5.0 x16 board (expensive RDIMMS, $30/GB!) and patched Linux kernel modules to get P2P working.

96 GB is sort of silly though, because Qwen-27B only needs 64 GB and the next size up requires 256 GB. 96/128 is only useful if you have a specific reason to run Qwen-Flash-Next, like video annotation or web scraping.
 
Back
Top