• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Extremely sporadic system turn offs

StoleMyOwnCar

2[H]4U
Joined
Sep 30, 2013
Messages
3,526
As stated in the topic, I'm having an issue with extremely sporadic system shut offs. When I say shut offs, I mean there is no blue screen or anything of the sort. The system literally acts like it has had the power switch on the back of the PSU shut off. At least as far as I can tell; this happens so suddenly that I can't even look inside of the case to see if the lights on the MB are on or not. The Windows event log just shows it as a "power kernel failure" or something.

My system is hooked up to an APC SUA1500 UPS. The UPS doesn't show anything unusual when this happens; the only light that is on says that it is online and working. I have in fact tried testing out my system running on pretty high settings with the UPS suddenly being forced to rely on battery power. The UPS had no issues sustaining the load on the battery... so I don't think it's the UPS. It would make no sense for it to suddenly have issues sustaining a load when plugged in that it could sustain under battery.

The frequency is... really random. I had it happen maybe 3 times total so far. Two of them were within 2-3 days of each other. I was on an short paid vacation from work and was playing WoW for a few days straight going into the weekend. It doesn't even make sense because I have played more demanding (GPU TDP wise) games than WoW. Granted I play on an RoG Swift so I can squeeze more out of my GPU's regardless of games due to the higher FPS. It's worth noting that I haven't had a crash since then (little over a week ago now). I've tried playing some very demanding games for a few hours but still can't seem to replicate the issue.

What tests should I perform? I tried running Furmark to see if it's TDP related, but I can't get it running on my SLI setup. It only stresses one card and not the other... I tried their guide to get it working on SLI, but I can't seem to.

Should I just send my PSU in for an RMA regardless? Is there something else that can cause this? I do have a backup system and backup 650 Corsair GS PSU (I'd probably have to disable one of the cards if I choose to use that).


Full (relevant) specs:
Corsair 850HX PSU
4770k@4.4Ghz (~1.29-1.3vcore, I turned off adaptive voltage and clocking so it's always at 4.4Ghz)
EVGA 780GTX SLI (both stock. They seem to boost to around 1.1-ish. Temps are normal, with one of them on an AIO)
MSI MPower Z87
Crucial Ballistix Tactical 16GB 1600

I have a bunch of HDD's and fans and whatnot if that matters.
 
Last edited:
Run Memtest, preferably for at least ten hours. If you get an error, especially one that doesn't happen on every pass, make sure your motherboard's BIOS/UEFI is updated.
 
-Don't you have a spare PSU to try it out? (*your current one must be near 5-6 years old i suppose ?. Not a very reassuring age especially if you run SLI )
-Also, maybe you should try to run your system without the UPS for few days, in order to exclude it out of the suspects, if the problem insists.
 
Run Memtest, preferably for at least ten hours. If you get an error, especially one that doesn't happen on every pass, make sure your motherboard's BIOS/UEFI is updated.

Can RAM just utterly cause a system to turn off?

I'm working on getting memtest going. I tried doing it this morning before I went to work, but my motherboard apparently really hates trying to run off of a USB thumbdrive. I even explicitly selected the thing out of my boot menu and it just went to Windows 8 anyway. Might just have to burn it to a disk. Why can't anything ever be simple? Anyway I'll try to get that memtest going soon enough. RAM was actually my second suspect, I just didn't know of a good program to test it with.

-Don't you have a spare PSU to try it out? (*your current one must be near 5-6 years old i suppose ?. Not a very reassuring age especially if you run SLI )

Hm? Why do you suppose it must be 5-6 years old? Not a single component in my system was released that long ago. This PSU about 2-3 years old. While it isn't utterly top of the line, I read it was a pretty good PSU. I've never had a PSU fail on me. Then again looking at Newegg reviews it's quite possible that it went bad. Anyway yes I do have a backup, but it's 650W, and it's a GS series. Not exactly prime material for running 780 SLI and overclocked CPU off of.



-Also, maybe you should try to run your system without the UPS for few days, in order to exclude it out of the suspects, if the problem insists.

The thing is, these shut offs are so sporadic that the tradeoff I'm looking at is possibly running without a UPS at all for the foreseeable future to try to test this. During which time the system might well shut off anyway because of an outage.... and that may be even worse for my system's health.

Same deal with swapping out the PSU. How long would I have to keep it like that?

That being said, I recently purchased an SUA1000 that I am using for my projector. I could switch over to using it for my PC instead (though it has less juice). It should be able to supply enough for my PC anyway as they are both server grade UPSs.



Either way I need more ideas on how to replicate/catch the issue, though (memtest is a good idea). Just swapping out components works fine when the issue happens reliably enough, but things get more tricky when it's like this. I guess I could do memtest and work off of my other UPS for a week or two and then if it never happens, I just chalk it up to the PSU and send it back to Corsair anyway. Better to eliminate one possible issue.
 
Hm? Why do you suppose it must be 5-6 years old? Not a single component in my system was released that long ago. This PSU about 2-3 years old.
---------------------------------
The thing is, these shut offs are so sporadic that the tradeoff I'm looking at is possibly running without a UPS at all for the foreseeable future to try to test this.

-About the age, i made an assumption based on the reviews about the HX line. [H]'s review about Corsair HX850 (*non-gold ) was released at 2009 :p ( http://www.hardocp.com/article/2009/05/27/corsair_hx850w_power_supply/#.VRLoduEghgw )
-As for the problem, if it appears so sporadically, then it might not be hardware problem, but it could be related with temperatures.
Some times during summer, when i was playing "heavy" games, my cpu was reaching 80-90C, and i had system shutdowns. :confused:
 
Can RAM just utterly cause a system to turn off?

I've seen it personally. In the first round of diagnostics, one of the RAM sticks was definitely bad and was replaced, but the shutdowns continued at a much lesser frequency. Memtest was getting an error consistently on every third pass which was unusual. Updating the motherboard's BIOS solved both the Memtest error and the shutdowns completely.

I do agree the UPS or PSU are the more likely suspects, especially if you never get any errors in Memtest.
 
Can RAM just utterly cause a system to turn off?

Bad memory can cause a machine to not POST, as well as random shutdowns. It can also cause other really bizarre issues, especially if the bad memory location is shared with the GPU where it can really cause havoc.

And other times you could have a bad memory location high in the memory map where it is rarely used and go years without noticing it since nothing ever tries to use it.
 
I expect this is a PSU issue. PSU's are rated not only for maximum power but consistency of power.

A bad PSU can run fine one moment then be under powering your system the next. The result is a total system shutdown.
 
I've seen it personally. In the first round of diagnostics, one of the RAM sticks was definitely bad and was replaced, but the shutdowns continued at a much lesser frequency. Memtest was getting an error consistently on every third pass which was unusual. Updating the motherboard's BIOS solved both the Memtest error and the shutdowns completely.

I do agree the UPS or PSU are the more likely suspects, especially if you never get any errors in Memtest.

Today I managed to get memtest up and running before I went to work. I'll let it run for 10 hours total (should be at about 9 by the time I get back, and then I'll be taking my hour long walk anyway). I did one pass of it yesterday and it didn't uncover any errors. I'll see if it's different this time. Is 10 hours a ballpark estimate or should I run it longer?

-About the age, i made an assumption based on the reviews about the HX line. [H]'s review about Corsair HX850 (*non-gold ) was released at 2009 :p ( http://www.hardocp.com/article/2009/05/27/corsair_hx850w_power_supply/#.VRLoduEghgw )
-As for the problem, if it appears so sporadically, then it might not be hardware problem, but it could be related with temperatures.
Some times during summer, when i was playing "heavy" games, my cpu was reaching 80-90C, and i had system shutdowns. :confused:

No, my top (technically not "top" considering the way my case is oriented) GPU is on an AIO water cooler and my CPU is also on an AIO water cooler. None of the temperatures are getting out of hand. Now, if I did Prime95 on my CPU, things would go out of hand. But I haven't had any temperature related issues during normal operation. I also haven't had any vcore-related BSODs since I've jacked up the vcore slightly a while back. So it isn't prime95 stable, but it's stable.

Bad memory can cause a machine to not POST, as well as random shutdowns. It can also cause other really bizarre issues, especially if the bad memory location is shared with the GPU where it can really cause havoc.

And other times you could have a bad memory location high in the memory map where it is rarely used and go years without noticing it since nothing ever tries to use it.

Well I knew it can cause it not to post. I've seen that happen when one of my sticks was slightly loose in the past. The BIOS made a different set of beeps than usual.

Totally off topic but that reminds me. It's a few years ago now (and not what I'm doing for my job so I forgot about a fair share)... but I've actually implemented a "simple computer" with VHDL and some MIPS as part of my course work in the past. Then my partner and I ran some simulations against that and watched the registers fill up and watched it store data and stuff.... implemented branch prediction, forwarding, pipelining, etc.

I never really thought about what would happen if any of those little logic elements were unreliable at that low of a level. Probably total frickin chaos. Modern PCs are so complex that it's frankly a miracle they work as well as they do.

I expect this is a PSU issue. PSU's are rated not only for maximum power but consistency of power.

A bad PSU can run fine one moment then be under powering your system the next. The result is a total system shutdown.

I suspect it is as well, which is why I posted it here. The idea was to see if it could be anything but a power issue. Thankfully someone pointed out an alternative (RAM) so I am testing that.
 
Last edited:
Ram's fine. Dunno what else to test. Sending back this PSU is gonna be such a pain. I thought these semi-expensive Corsair models were supposed to be good. Le sigh.
 
Back
Top