• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Lost a 8102 at 80%

thinklet

Limp Gawd
Joined
Jul 4, 2012
Messages
142
Hi,
I lost an 8102 today, I copied the message below, what does it mean? How do I avoid a repeat? PPD was 560k with a 382k credit. Unit has been folding stable for 4 days.

[
00:52:30] Completed 192500 out of 250000 steps (77%)
[01:02:14] Completed 195000 out of 250000 steps (78%)
[01:11:58] Completed 197500 out of 250000 steps (79%)
[01:21:44] Completed 200000 out of 250000 steps (80%)

Step 7700434, time 30801.7 (ps) LINCS WARNING
relative constraint deviation after LINCS:
rms 15535.776003, max 541260.562500 (between atoms 163348 and 163347)
bonds that rotated more than 90 degrees:
atom 1 atom 2 angle previous, current, constraint length
163353 163352 90.0 0.1530 129.6753 0.1530
163339 163338 94.5 0.1530 0.9781 0.1530
163340 163339 92.0 0.1430 5.1608 0.1430
163341 163340 125.8 0.1610 22.3396 0.1610
163344 163341 91.0 0.1610 694.3965 0.1610
163345 163344 90.0 0.1430 3761.4143 0.1430
163370 163369 90.6 0.1530 510.2241 0.1530
163371 163370 96.5 0.1530 84.2371 0.1530
163338 163333 100.6 0.1583 0.1107 0.1583

Step 7700435:
The charge group starting at atom 163353 moved than the distance allowed by the domain decomposition (1.480467) in direction X
distance out of cell 32.483398
Old coordinates: 25.125 10.930 4.468
New coordinates: 55.460 5.451 7.729
Old cell boundaries in direction X: 17.619 19.830
New cell boundaries in direction X: 17.614 19.830

-------------------------------------------------------
Program Gromacs, VERSION 4.5.3
Source code file: /vspm58/VM/fah-converted/mnt/fah_windows_build/LinuxBuilds/gromacs-4.5.3/src/mdlib/domdec.c, line: 4117

Fatal error:
A charge group moved too far between two domain decomposition steps
This usually means that your system is not well equilibrated
For more information and tips for troubleshooting, please check the GROMACS
website at http://www.gromacs.org/Documentation/Errors
-------------------------------------------------------

Thanx for Using GROMACS - Have a Nice Day

[01:23:25] mdrun returned 255
[01:23:25] Going to send back what have done -- stepsTotalG=250000
[01:23:25] Work fraction=37.4188 steps=250000.
[01:23:29] logfile size=149586 infoLength=149586 edr=25 trr=1
[01:23:29] logfile size: 149586 info=149586 bed=25 hdr=1
[01:23:29] - Writing 150124 bytes of core data to disk...
[01:23:29] Done: 149612 -> 16141 (compressed to 10.7 percent)
[01:23:29] ... Done.
[01:23:34]
[01:23:34] Folding@home Core Shutdown: UNSTABLE_MACHINE
[01:23:35] CoreStatus = 7A (122)
[01:23:35] Sending work to server
[01:23:35] Project: 8102 (Run 0, Clone 9, Gen 30)


[01:23:35] + Attempting to send results [September 2 01:23:35 UTC]
[01:23:35] - Reading file work/wuresults_01.dat from core
[01:23:35] (Read 16653 bytes from disk)
[01:23:35] Connecting to http://128.143.231.201:8080/
[01:23:35] Posted data.
[01:23:35] Initial: 0000; Conversation time very short, giving reduced weight in bandwidth avg
[01:23:35] - Uploaded at ~34 kB/s
[01:23:35] - Averaged speed for that direction ~367 kB/s
[01:23:35] + Results successfully sent

Best regards, Charlie
 
It might just be an intrinsic problem of the WU by chance, I think.
 
Oh man, that's sad :(
It could be just a bad WU, but gotta ask, is there any OC involved?
 
Oh man, that's sad :(
It could be just a bad WU, but gotta ask, is there any OC involved?

Hi Gryphon,

Yes, it is a 6174 running at 2500 MHz. Great cooling and been stable since I powered it up.
Best regards, Charlie
 
Not saying it's the OC, but it could be related. Some WUs are more sensitive to those tiny bit of instabilities that are not revealed 99.99% of the time... Or, it could just be a bad WU.

Edit: I guess I would keep running it as it is, but watch for any other lost WUs. In case you have another, it may be time to tune the OC down a bit.

Hi Gryphon,

Yes, it is a 6174 running at 2500 MHz. Great cooling and been stable since I powered it up.
Best regards, Charlie
 
One thing about it you should get some credit for it. It sent something back to stanford. try droping the OC by 1 just to be on the safe side but my guess is it was an unstable WU from reading up on the error in the link provided in your log. If you have been completing 8101's with no problem I doubt the 8102 would cause your OC to crash the 8102, the 8102's are easier WU's to process than the 8101's I can run the 8102's at 3.1Ghz but I have to drop the OC to 3Ghz with the 8101's on my ES chips.
 
Back
Top