About 1.5 years ago I built two machines to do some image processing with GPUs. Both machines are identical except for PSU:
Mobo: ASUS P6T7
CPU: Intel i7-960
case: Lian Li Lancool PC-K58
PSU: Machine 1 had a Cooler Master 1200W, now replaced with a Rosewill Hercules 1600W, machine 2 has ThermalTake 1500W
GPUs: 4 nVidia GTX 570s
RAM: 16GB (CORSAIR XMS3 16GB (4 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333)
OS: OpenSUSE 12.1
I am running custom CUDA code on both machines, one instance per GPU. Everything was going fine for about 17 months, with both machines under full load almost continuously. Then machine 1 started hanging randomly - the machine would remain on (chassis lights, mobo lights, fans all on) but video output and network activity would go down, and logging to the hard drive would stop. To get the machine up I would have to power cycle. I immediately suspected the PSU, so I replaced it with a sturdier model - the Rosewill Hercules 1600W. This PSU has two +12V rails, one sourcing 50A and the other sourcing 110A. The 50A rail is dedicated to the CPU while the 4GPUs are on the 110A rail. Unfortunately, I am still having problems with apparently random shutdowns. The symptoms are the same: lights and fans remain on, video output, hard drive logging and network activity go down. I have monitoring of the hardware running via lm_sensors and nvidia-smi, and logs of voltages as well as mobo temp, CPU temp and GPU temp show no sign of increase prior to shutdown. The shutdowns occur even without the GPUs running any analysis. Now I am starting to have identical problems with the second machine! I would very much appreciate any help diagnosing this problem.
Mobo: ASUS P6T7
CPU: Intel i7-960
case: Lian Li Lancool PC-K58
PSU: Machine 1 had a Cooler Master 1200W, now replaced with a Rosewill Hercules 1600W, machine 2 has ThermalTake 1500W
GPUs: 4 nVidia GTX 570s
RAM: 16GB (CORSAIR XMS3 16GB (4 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333)
OS: OpenSUSE 12.1
I am running custom CUDA code on both machines, one instance per GPU. Everything was going fine for about 17 months, with both machines under full load almost continuously. Then machine 1 started hanging randomly - the machine would remain on (chassis lights, mobo lights, fans all on) but video output and network activity would go down, and logging to the hard drive would stop. To get the machine up I would have to power cycle. I immediately suspected the PSU, so I replaced it with a sturdier model - the Rosewill Hercules 1600W. This PSU has two +12V rails, one sourcing 50A and the other sourcing 110A. The 50A rail is dedicated to the CPU while the 4GPUs are on the 110A rail. Unfortunately, I am still having problems with apparently random shutdowns. The symptoms are the same: lights and fans remain on, video output, hard drive logging and network activity go down. I have monitoring of the hardware running via lm_sensors and nvidia-smi, and logs of voltages as well as mobo temp, CPU temp and GPU temp show no sign of increase prior to shutdown. The shutdowns occur even without the GPUs running any analysis. Now I am starting to have identical problems with the second machine! I would very much appreciate any help diagnosing this problem.