• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

-bigadv: what am I doing wrong?

Kendrak

[H]ard|DCer of the Year 2009
Joined
Aug 29, 2001
Messages
21,141
Code:
Launch directory: /usr/local/fah
Executable: ./fah6
Arguments: -bigadv -smp 7 

[23:26:51] - Ask before connecting: No
[23:26:51] - User name: kendrak (Team 33)
[23:26:51] - User ID: 59588A326783A152
[23:26:51] - Machine ID: 1
[23:26:51] 
[23:26:51] Loaded queue successfully.
[23:26:51] 
[23:26:51] + Processing work unit
[23:26:51] Core required: FahCore_a2.exe
[23:26:51] Core found.
[23:26:51] Working on queue slot 01 [December 9 23:26:51 UTC]
[23:26:51] + Working ...
[23:26:51] 
[23:26:51] *------------------------------*
[23:26:51] Folding@Home Gromacs SMP Core
[23:26:51] Version 2.10 (Sun Aug 30 03:43:28 CEST 2009)
[23:26:51] 
[23:26:51] Preparing to commence simulation
[23:26:51] - Ensuring status. Please wait.
[23:26:51] Files status OK
[23:26:56] - Expanded 30328091 -> 159726549 (decompressed 101.8 percent)
[23:26:57] Called DecompressByteArray: compressed_data_size=30328091 data_size=159726549, decompressed_data_size=159726549 diff=0
[23:26:57] - Digital signature verified
[23:26:57] 
[23:26:57] Project: 2681 (Run 8, Clone 5, Gen 44)
[23:26:57] 
[23:27:01] Assembly optimizations on if available.
[23:27:02] Entering M.D.
[23:27:07] Using Gromacs checkpoints
[23:27:13] 
[23:27:14] Entering M.D.
[23:27:20] Using Gromacs checkpoints
[23:27:39] Resuming from checkpoint
[23:27:40] Verified work/wudata_01.log
[23:27:41] Verified work/wudata_01.trr
[23:27:42] Verified work/wudata_01.xtc
[23:27:42] Verified work/wudata_01.edr
[23:27:44] Completed 30750 out of 250000 steps  (12%)
[23:57:03] Completed 32500 out of 250000 steps  (13%)
[00:23:47] 
[00:23:47] Folding@home Core Shutdown: INTERRUPTED
[00:23:54] CoreStatus = FF (255)
[00:23:54] Sending work to server
[00:23:54] Project: 2681 (Run 8, Clone 5, Gen 44)
[00:23:54] - Error: Could not get length of results file work/wuresults_01.dat
[00:23:54] - Error: Could not read unit 01 file. Removing from queue.
[00:23:54] - Preparing to get new work unit...
[00:23:54] Cleaning up work directory
[00:23:54] + Attempting to get work packet
[00:23:54] - Connecting to assignment server
[00:23:56] - Successful: assigned to (171.67.108.22).
[00:23:56] + News From Folding@Home: Welcome to Folding@Home
[00:23:57] Loaded queue successfully.

I'm lost, what am I doing wrong, I always seem to crash a bit above 10%
 
unstable overclock? I assume by your 7 cores that you're using an oc'd i7?
 
It seems that T4rd was having the same issue. Looks like he had to play around with voltages to get his OC stable.

I'm pretty sure it's stable now @ 3.8 with a 1.26875v vCore. It's been crunching on a -bigadv for almost 6 hours now at 10.6k PPD and temps maxing around 78°C :D. I hope it doesn't need any more adjustment. If it does, I'll try playing with the uncore as you suggested (everything besides vCore is on Auto right now).
 
I agree with the above two posts. Milder OC and resume to see if there's a continuing problem.
 
I've got the cores under 70C while under 100% load. It's not heat.

I'm cranking up the volts and we will see where that does.

The thing is I would assume it is stable, it ran normal smp units on all 8 cores for over a week.
 
I had that problem at one time and I had to increase the QPI voltage rather than the core voltage to get stable.
 
K

Are you running your OC with Turbo enabled?
Ive noticed some people turn that off when they OC so the clocks and multiplier dont fluctuate causing inconsistencies.
 
I've got the cores under 70C while under 100% load. It's not heat.

I'm cranking up the volts and we will see where that does.

The thing is I would assume it is stable, it ran normal smp units on all 8 cores for over a week.
Kendrak, it might not be directly related to heat, but the OC might still be a problem. I have a system with very cool running processors that never get beyond slightly warm to the touch of their HSFs, but crank the OC to a few MHz higher, and they will cause errors. I imagine the -bigadv WUs that utilize so many cores and threads would probably be even more sensitive to stress.
 
How high can you up the Core Voltage before it starts to become a concern....?

:D

I'll try the sledge hammer way before I start dropping the BLK any more.
 
K on gigabyte boards it turns the text red in BIOS

But i think its like 1.45?

Also make sure your RAM and QPI and uncore are getting plenty of voltage.
I dont think it needs much extra vcore.
 
K on gigabyte boards it turns the text red in BIOS

But i think its like 1.45?

Also make sure your RAM and QPI and uncore are getting plenty of voltage.
I dont think it needs much extra vcore.

What are good settings on those?

Just bump them up a few from default or what?

I'm a bit simple on OC. Up Vcore and the FSB/BLK and let 'er rip.
 
I've got the cores under 70C while under 100% load. It's not heat.

I'm cranking up the volts and we will see where that does.

The thing is I would assume it is stable, it ran normal smp units on all 8 cores for over a week.

I've seen a lot of people say F@H is not a good test to see if an OC is stable. Try running prime95 or orthos and see what you come up with.
 
What are good settings on those?

Just bump them up a few from default or what?

I'm a bit simple on OC. Up Vcore and the FSB/BLK and let 'er rip.

Really im not sure myself.
Not a big OC guy.

All i did was set the OC in the windows gigabyte config utility.
I didnt mess with voltages.

Between running SMP and GPU clients there are enough variables i dont want to troubleshoot an OC
 
I've seen a lot of people say F@H is not a good test to see if an OC is stable. Try running prime95 or orthos and see what you come up with.
Hmm, that's interesting. A lot of people here in the past have posted the opposite, that success in programs like prime95, etc., aren't indicative of a successful F@H OC. Many folders believe that nothing is more stressful on their rigs than folding. :confused:
 
Hmm, that's interesting. A lot of people here in the past have posted the opposite, that success in programs like prime95, etc., aren't indicative of a successful F@H OC. Many folders believe that nothing is more stressful on their rigs than folding. :confused:

This

There is no way to know if your rig is F@H stable unless you run F@H on it.

Ever couple of months we get someone that says it's prime 95 stable, and every time we tell them to back off the OC and it works.
 
I'm pretty sure that nitro is right about the voltage. Everywhere I've looked 1.45v is the max safe voltage. I'm guessing that the key word there is "safe" :)
 
I'm pretty sure that nitro is right about the voltage. Everywhere I've looked 1.45v is the max safe voltage. I'm guessing that the key word there is "safe" :)

Wait.... are you suggesting we go over the safe voltage....

evilscientist.jpg


:)
 
Thanks for the laugh Kendrak!
I needed it after the day that I've had.

"safe" is kinda like the pirate code if you ask me
more of a guideline :)
 
Sorry, I haven't been on this evening or I'd have posted sooner because I was experiencing the same thing you are for the first week or two, as Doozer mentioned.

LOL, I wouldn't recommend going over 1.4v unless you got some serious (water) cooling or an AC unit blowing over the heatsink, hah. When I had my vCore set to Auto it wanted to give mine 1.43v and it instantly hit the TJ max when it loaded up a WU, lol. I definitely wouldn't go past 1.45v, If you can't hit a good OC with less, then I'm afraid you're just not going to get a good OC at all :(. But I'm sure you can with some more tinkering as I did.

I could have swore I was stable for a couple months because I was folding normal big WUs in notfreds for a couple months with no issues at all @ 3.36 GHz (160 blck, 21x multi, vcore "auto"), but my RAM was at 1604 MHz (rated for 1600). I think this was the culprit for my -bigadv crashing because I was utilizing > 80% of my RAM with a -bigadv VM, that's just my best guess though. I would crash anywhere from 10%-40% into a WU constantly. So I decided to run Prime95 to see what happend, and sure enough it error'd out on most of my cores/threads within 5 mins.

So I looked up Kyle's review of my mobo and in the OC section he said he got his 920 stable at 190 BLCK 20x multi. So I figured I should be able to do that too.

Here's how I got my rig stable:

I put in a 190 BLCK and 20x multi in the BIOS for a 3.8 GHz OC. This put my RAM at 1522 (I think) MHz, which was the fastest speed increment available under its 1600 MHz.

vCore was left on Auto.

Booted to windows, loaded up Realtemp, Prime95, and my mobos OC utility (TurboV on Asus boards, I'm sure your mobo has a similar utility) for adjusting voltages and bus speeds on the fly while benching it.

Ran Prime95, temps shot up to 100°C instantly due to the mobo giving a 1.43 vCore on Auto. I immediately dropped down the vCore manually in the software OC utility to 1.3 and temps dropped accordingly to low 80°s.

Prime never error'd out during all that, so I would wait a couple mins, drop vCore, wait a couple more mins, drop vCore again. Repeat until it errors/crashes, or BSODs as mine did finally at 1.25 vCore.

So I set the vCore in the BIOS to 1.25625v (next stop up from 1.25v) after the BSOD and started Prime95 again. Got errors in a few mins. Bumped vCore another notch in OC utility, and it didn't error after 30 mins, so I hoped it was stable. But after 12 hours of crunching a -bigadv WU, it "interrupted" again. So I bumped the vCore up another notch again. Next time it took a day to get "interrupted". Then I bumped it to 1.275v and I've been stable through the last 2.7 -bigadv WUs so far.

I know some people say Prime95 isn't the best to test stability, but I'll tell you this; when it took my VM to error out after 12-24 hours of crunching, Prime95 error'd after only 5 mins or so. Also, my temps were 3-5° higher on average over what I see folding. So from what I saw, Prime was stressing my CPU harder than folding was. Just my experience.

So although it was slightly a PIA, at least I didn't have to mess with anything besides vCore, BLCK, the multiplier, and RAM speed. And still a relatively easy OC compared to older CPU architectures and mobos I've had. Sorry for dragging it out a bit, I just wanted to let you know my whole experience with getting my i7 stable for -bigadv WUs.

Cliffs:
-Stay under RAM spec'd speeds.
-Keep vCore <= 1.4v on air.
-Use Prime w/ OC utility to easily test vCore/stability.
-Keep bumping vCore up one notch until stable.

I hope this helps, K. Here's a screenshot I just took of my current settings to verify.

myoc.jpg


PS: Here's where you can get that GPU/CPU temp monitoring gadget I have there. You need Rivatuner installed to use it though. Once you have that installed, just open on the hardware monitor and click on the little red button at the bottom-left of the window to "enable background monitoring" and you'll be set. To add CPU temps to the gadget, download/open Real Temp, click "settings" > "Rivatuner", browse to rivatuners .exe > follow onscreen directions from there.
 
Last edited:
Thanks.

I'm at 17% now. Farther than I have been before. My hope is that is gets though work today.

I'm sitting at 75c so I should be good.
 
Back
Top