What happens when you compare LLM inference running on a Blackwell RTX6000 versus 3x RTX 5090 in one machine?
Would it be faster due to more GPU units? Or slower due to communication overhead? Or is it "it depends"?
Of course power consumption of the trio would suck (no pun intended).
Would it be faster due to more GPU units? Or slower due to communication overhead? Or is it "it depends"?
Of course power consumption of the trio would suck (no pun intended).