• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Multi-GPU machine shuts down randomly

zakaryah

n00b
Joined
Jul 13, 2011
Messages
5
About 1.5 years ago I built two machines to do some image processing with GPUs. Both machines are identical except for PSU:

Mobo: ASUS P6T7
CPU: Intel i7-960
case: Lian Li Lancool PC-K58
PSU: Machine 1 had a Cooler Master 1200W, now replaced with a Rosewill Hercules 1600W, machine 2 has ThermalTake 1500W
GPUs: 4 nVidia GTX 570s
RAM: 16GB (CORSAIR XMS3 16GB (4 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333)
OS: OpenSUSE 12.1

I am running custom CUDA code on both machines, one instance per GPU. Everything was going fine for about 17 months, with both machines under full load almost continuously. Then machine 1 started hanging randomly - the machine would remain on (chassis lights, mobo lights, fans all on) but video output and network activity would go down, and logging to the hard drive would stop. To get the machine up I would have to power cycle. I immediately suspected the PSU, so I replaced it with a sturdier model - the Rosewill Hercules 1600W. This PSU has two +12V rails, one sourcing 50A and the other sourcing 110A. The 50A rail is dedicated to the CPU while the 4GPUs are on the 110A rail. Unfortunately, I am still having problems with apparently random shutdowns. The symptoms are the same: lights and fans remain on, video output, hard drive logging and network activity go down. I have monitoring of the hardware running via lm_sensors and nvidia-smi, and logs of voltages as well as mobo temp, CPU temp and GPU temp show no sign of increase prior to shutdown. The shutdowns occur even without the GPUs running any analysis. Now I am starting to have identical problems with the second machine! I would very much appreciate any help diagnosing this problem.
 
Remove all but the essential hardware necessary to run the machine and try again. If it works fine, start adding in hardware a piece at a time until the problem comes back. This process may take some time to identify the offending hardware, but that's the price you pay when running a fully loaded machine...;)
 
I immediately suspected the PSU, so I replaced it with a sturdier model - the Rosewill Hercules 1600W. This PSU has two +12V rails, one sourcing 50A and the other sourcing 110A. The 50A rail is dedicated to the CPU while the 4GPUs are on the 110A rail.
Not entirely accurate: Both rails draw from the same amount of amperage available just for the +12V rail. In this case, that Rosewill has 1560W available on the +12V rail or 130A.

Anyway, have you tried removing all the GPUs and trying a known working GPU?
Also, try testing the system outside of the PC case, i.e. the mobo on top of a cardboard box and starting the system up with a screw driver.

My hunch however is the motherboard.
 
I tried stressing the machine with all combinations of components, starting with the RAM and moving on to individual GPUs, then combinations of the GPUs. So far the only condition under which the machine fails is when 3 or 4 GPUs are running at the same time - with 1 or 2 GPUs, the stress test can go for over a week with no problems, while with 3 GPUs running the machine typically hangs after 3-4 days.

This makes me think that either something is overheating, or the PSU cannot supply enough power. I monitor mobo and GPU temperatures with lmsensors and nvidia-smi, and voltages with lmsensors. Nothing suspicious shows up in any measurement prior to hanging.

Any ideas?
 
You have already tried three different PSUs with the same four GPU system?

Heat is a possible contender considering it took four days for the three GPU setup to hang up. Leave the case side open and keep several 120mm fans pointed directly at the GPUs. Your case doesn't have side fan placement which is needed for optimal GPU cooling these days.
 
No, I only tried two PSUs (the Cooler Master 1200W, now replaced with a Rosewill Hercules 1600W). When I first had this problem with the machine hanging, I upgraded the PSU. This was because the other very similar machine had the 1500W ThermalTake, and I still have had no problems running that machine for long times with 4 GPUs.

I will try what you suggest with the fans, thanks. Should I do anything to monitor temperatures more carefully?
 
I am still having trouble with this machine, specifically the GPUs.

I set up some 120mm fans to blow cool air from the side. This lowered the GPU temps by a few degrees Celsius but they were not getting too hot anyway. So, I think I have ruled out problems with RAM, hard drives, temperature, and PSU.

Lately the GPUs have been failing frequently, either with messages like "GPU fell off the bus" or failing to execute code which is verified to be safe. In a sense, this is better from the old behavior, which was total freezing of the computer, but it is happening more often.

Due to the increased frequency of failures, I suspect either the GPUs or mobo. Because I am not sure whether some of the GPUs are defective, or if the mobo has a problem, I have tried several combinations of GPUs in various positions on the mobo. So far, the only conclusion I can reach is that in certain configurations, a card tests good in that it can compute properly for very long times (weeks), but when I change out cards in other PCI slots, the previously good card will start to have problems. I have ordered a couple new GPUs to see if this helps, but I am concerned about a problem with the mobo. I don't want to replace it unless I have to, and I don't have very strong evidence in my opinion. Can anyone recommend additional tests? Why would the mobo only have problems with respect to the GPUs? Is the Asus P6T7 prone to such problems? Of the mobos with many PCI x16 slots, is there one which is more reliable?
 
If the GPUs are exhibiting issues when additional GPUs are added or swapped, it really does sound like the motherboard. With that said, new GPUs would completely rule out the GPUs.
 
Back
Top