• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Stability issues with F@H & OC

jamsomito

2[H]4U
Joined
Aug 29, 2010
Messages
3,202
Ok, I wasn't sure if this should go here or in Overclocking & Cooling, so mods please move if you see fit. Also I'm just a little guy in the DC world, so go easy on me. :)

My goal now is to fold for a year. I'm half way through month 4 and I'm starting to see some issues. Since I started, I toyed with different versions of F@H control, removing background tasks (like system performance monitors with graphical readouts, etc), new CPU cooler which allowed me to up my CPU OC, and ultimately new folding cores which allowed power-realistic PPD out of my AMD GPU.

As you can see, I steadily increased my output... until recently.

Link to EOC stats.

I've been getting more and more random blue screens. Upping my CPU voltage seems to help for a little while, but the BSOD's keep coming back eventually, and with no particular pattern that I can tell (haven't checked specific WU's though). It's particularly frustrating because it will hang on the BSOD screen, then shut the system down. When it's in either of those states, I can't reset remotely through logmein, so it's usually down for at least half a day at a time, if not more before I realize it.

The only two things I can think of are CPU degradation from months of high(er) temps, and/or rising ambient temperature as we progress through springtime here in Chicagoland. Is 2 months straight @ 75C enough to hurt a 2500k? Is a difference between 58C and 66C enough to cause an overclock to be unstable?

Here's a screenshot of my settings.
Specs of rig in sig (TrueBlue).

Tips? Any help is appreciated.

TL;DR - getting more blue screens, don't know what from.
 
would be nice to know the BSOD codes, have you remembered/written any of them down?
 
Look in your event viewer logs to get the STOP codes. If they have been deleted, next time it happens, take a photo and post it.
 
I don't know why that never occurred to me... unfortunately I don't have any pics or numbers off of the BSOD. I'll probably catch one this weekend, so I'll let you know if I get one.

I just have too many variables at this point. It started happening when I started using Core 17 on my GPU... not sure if it's that or the extra power draw on my system, or my processor or what.

I'll let you know when I have more info. Thanks for the recommendation.
 
do you overclock your GPU? Core 17 may not like your OC and you may have to turn it down
 
The original intent was to OC the GPU, but I have not gotten around to it yet so it's still running stock.

It's an MSi Hawk though, so it's factory OC'd. Think that's messing with it?
 
probably not, it could have possibly been OC'ing but if you aren't yet, then we'll rule that out for now. get the stop error codes from your event log and that will give everyone more to go on to help out.

should be able to right click on "computer" or 'my computer', choose "Manage".

then open up the Event Viewer, then Windows Logs, then System. Look for red circles with exclamation points and look for the long 0x000000007e (or similar) code. there might be like 5 of them in a series, but i *think* the first one is all that is necessary. or you could paste all the text here with a
Code:
wrapper.
 
Alright, checked my system logs and found a few things. Here are the ones that look relevant

Event ID 1001
Code:
The computer has rebooted from a bugcheck.  The bugcheck was: 0x00000124 (0x0000000000000000, 0xfffffa800726f028, 0x00000000be200000, 0x000000000005110a). A dump was saved in: C:\Windows\MEMORY.DMP. Report Id: 041513-6879-01.

Event ID 1001
Code:
The computer has rebooted from a bugcheck.  The bugcheck was: 0x00000101 (0x0000000000000031, 0x0000000000000000, 0xfffff88003165180, 0x0000000000000002). A dump was saved in: C:\Windows\MEMORY.DMP. Report Id: 041413-14242-01.

Event ID 41 (this one marked Critical, not Error)
Code:
The system has rebooted without cleanly shutting down first. This error could be caused if the system stopped responding, crashed, or lost power unexpectedly.

Event ID 6008
Code:
The previous system shutdown at 8:01:06 AM on ‎4/‎15/‎2013 was unexpected.

Event ID 219 (this one only marked Warning)
Code:
The driver \Driver\WUDFRd failed to load for the device WpdBusEnumRoot\UMB\2&37c186b&0&STORAGE#VOLUME#_??_USBSTOR#DISK&VEN_MULTIPLE&PROD_CARD__READER&REV_1.00#058F63666438&0#.

Event ID 7000
Code:
The WinRing0_1_2_0 service failed to start due to the following error: 
The system cannot find the file specified.

There's also a few random hard disk controller errors on HD2 and HD3, but F@H runs on 1 I believe (OS installed on 0). These drives are also 8 years old, +/- a year, so I'm not too worried about it.
 
looks like error 0x...124 is fatal hardware error - could be driver issue. first google result indicates CPU.

looks like error 0x...101 is CPU related per google as well...
 
up your voltage or lower your OC, since you are running at 75, i would say lower the multi, do one at a time and fold, if it crashes again, lower the multi again. i can game and stress test and be stable at a lot higher frequencies with my 980x, but i can't reliably fold unless at a lower frequency, i chose the obvious choice ;)
 
Thanks Sc0tty8 and bigted, I haven't forgotten about you guys. I won't be able to work on my system until Sunday coming up here. Just wanted to say thanks for the help.

Also, the PC hasn't crashed since last weekend. I really think the ambient temperature plays a role in it... do borderline OC's get unstable at higher temps?
 
i remember using this guide when i was pushing my 920: http://www.xtremesystems.org/forums/showthread.php?266589-The-OverClockers-BSOD-code-list

it lets you know what to do when you get certain bsods when overclocking. both of the codes you mention say to raise vcore, so methinks your OC has become unstable just enough to have f@h crash it every now and then. you could try raising vcore, but like i said, your temps are a little high, so maybe try lowering multi one first
 
I personally prefer configuring my client to Uniprocessor sense it seems that most PPD comes from the GPU anyway. (I have a 560ti for reference) Plus it allows me to use the extra three cores for Universal Media Server Video file processing+streaming to other devices in the house.

I personally downclock my GPU to 700 core (from 880 stock), but have my CPU OC'd slightly to 3.75 from 3.3 stock. I pull 20k average no problem, and I think it's worth the reduced stress&temps on my CPU and the lowered temps on the GPU. It also lowers the wattage from the wall by close to 100 watts when I go from 4 cores maxed and 800+MHz on the GPU. fyi

Overclocking to attain incremental increase in PPD+Heat+substantially increased wattage (imo) seems like it's barely worth it imho. Longevity of parts could also be a concern for some folks as well, increased heat output in warmer months might also be undesirable. But if you want to OC and be stable u have to understand your chip_heatsink temp limitations along room temperature variances. So yeah just keep testing different variables until it clicks. But like I said I would actually downclock the GPU and leave fan on auto, but if ur trying to squeeze every last PPD then you will have to find the temp ceilings and voltage ceilings that will keep u stable. And incresed heat output does lower the lifespan of electronics. Peace.
 
Thought I'd update my post with a pic for factual proof. I do electric readings with a kill-awatt as well for personal reference...no pic of that but you can find them online around 15-20 usd quite easily for personal reference when overclocking or underclocking your system. Peace.
qry97a.png
 
@ bigted - where do you dig this stuff up?? That BSOD code list went right in my bookmarks.

@ sc0tty8 - checked my chipset drivers and they were the latest. However, I did not do research into if they were the most stable or not...

@ teletran8 - Thanks for the input! I really don't have any way of checking power consumption, but I think I saw some tangible gains in PPD with my overclock from 3.3 GHz to 4.5 GHz. It went from ~10-12k to ~16-18k. I have another rig that does 6k PPD for 250W, so even if OCing my main rig added 100W, it was still more efficient, and I was happy with the gains. Also, this rig sits idle about 90% of the time (only use for an hour a day on weekdays, and depending on the weekends, 0-10 hrs on the weekends) so if I really need the extra processing power while I'm using it, I can just back off the F@H for that short time and letter rip again afterwards.

I haven't gotten any BSOD's in a week, which I think is attributed to the lower ambient temperatures as of late. I'll keep an eye on it. If I get another one, I'll back down from 4.5 to 4.4 GHz :(. In my limited experience, it seems that ambient deltas (rising room temps) have an almost degree-for-degree impact on operating temps (i.e. room temp rises 5 deg, CPU temp rises 5 deg). As we get through spring here, I expect a 15deg shift, so I may have to back off a couple CPU multis.
 
OP you will see a valid increase, teletran runs uniprocessor, which will not see nearly the gain you do with smp units.
 
Back
Top