• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Doesn't this just piss you off?

Axdrenalin

[H]ard|DCer of the Month - Nov. 2009
2FA
Joined
Jan 28, 2004
Messages
6,301
I checked one of my Borgs at work today that ws folding, only to find yet another failed WU. Keep in mind this is on a bone stock 3.4 Ghz P4 with 2Gigs of Corsair DDR2- no overclocking or anything. It was from Project: 2653 (Run 19, Clone 50, Gen 61). This is what I had in the log:

Code:
[05:28:39] Completed 465000 out of 500000 steps  (93 percent)
[05:43:40] Timered checkpoint triggered.
[05:54:07] Writing local files
[05:54:08] Completed 470000 out of 500000 steps  (94 percent)
[06:09:07] Timered checkpoint triggered.
[06:19:33] Writing local files
[06:19:33] Completed 475000 out of 500000 steps  (95 percent)
[06:20:46] Gromacs error.
[06:20:46] 
[06:20:46] Folding@home Core Shutdown: UNKNOWN_ERROR
[06:20:46] 
[06:20:46] Folding@home Core Shutdown: UNKNOWN_ERROR
[09:21:37] CoreStatus = 7B (123)
[09:21:37] Client-core communications error: ERROR 0x7b
[09:21:37] Deleting current work unit & continuing...
[09:23:59] - Warning: Could not delete all work unit files (6): Core returned invalid code
[09:23:59] Trying to send all finished work units
[09:23:59] + No unsent completed units remaining.
[09:23:59] - Preparing to get new work unit...
[09:23:59] + Attempting to get work packet
[09:23:59] - Will indicate memory of 2023 MB
[09:23:59] - Connecting to assignment server
[09:23:59] Connecting to http://assign.stanford.edu:8080/
[09:24:00] - Couldn't send HTTP request to server
[09:24:00] + Could not connect to Assignment Server
[09:24:00] Connecting to http://assign2.stanford.edu:80/
[09:24:00] Posted data.
[09:24:00] Initial: 40AB; - Successful: assigned to (171.64.65.64).
[09:24:00] + News From Folding@Home: Welcome to Folding@Home
[09:24:00] Loaded queue successfully.
[09:24:00] Connecting to http://171.64.65.64:80/
[09:24:04] Posted data.
[09:24:04] Initial: 0000; - Receiving payload (expected size: 2427552)
[09:24:07] - Downloaded at ~790 kB/s
[09:24:07] - Averaged speed for that direction ~607 kB/s
[09:24:07] + Received work.
[09:24:07] + Closed connections
[09:24:12] 
[09:24:12] + Processing work unit
[09:24:12] Core required: FahCore_a1.exe
[09:24:12] Core found.
[09:24:12] Working on Unit 07 [May 14 09:24:12]
[09:24:12] + Working ...
[09:24:12] - Calling 'mpiexec -channel auto -np 4 FahCore_a1.exe -dir work/ -suffix 07 -checkpoint 15 -forceasm -verbose -lifeline 3604 -version 591'

[09:24:13] 
[09:24:13] *------------------------------*
[09:24:13] Folding@Home Gromacs SMP Core
[09:24:13] Version 1.74 (March 10, 2007)
[09:24:13] 
[09:24:13] Preparing to commence simulation
[09:24:13] - Ensuring status. Please wait.
[09:24:30] - Assembly optimizations manually forced on.
[09:24:30] - Not checking prior termination.
[09:24:37] - Expanded 2427040 -> 12916317 (decompressed 532.1 percent)
[09:24:37] - Starting from initial work packet
[09:24:37] 
[09:24:37] Project: 2653 (Run 31, Clone 3, Gen 60)
[09:24:37] 
[09:24:37] tray files
[09:24:37] - Starting from initial work packet
[09:24:37] 
[09:24:37] Project: 2653 (Run 31, Clone 3, Gen 60)
[09:24:37] 
[09:24:38] Assembly optimizations on if available.
[09:24:38] Entering M.D.
[09:24:44] Rejecting checkpoint
[09:24:46] Protein: Protein in POPC
[09:24:46] Writing local files
[09:24:48] Extra SSE boost OK.
[09:24:48] Writing local files
[09:24:48] Completed 0 out of 500000 steps  (0 percent)

I've seen several WU's end up like this over the past month, and all on different machines. It wouldn't be so bad if they failed within the first 10% of the WU, but to get all the way to 95% and then crash.....:mad: I'm 13% into the new one it picked up now, so I'm hoping since this is a different run it will go all the way through.

Just thought I'd share in the agony.....:rolleyes:

 
Yeah, 0x7b errors is the most dumbest one in existence, mainly due to the fact that it will delete the WU before th client had a chance to upload a partial result !!!

I even had 2-3 WU which erroned out with a valid error then prepare to upload the partial results. In that time, a 0x7b error popped and deleted before it's uploaded :mad:
 
Yeah, 0x7b errors is the most dumbest one in existence, mainly due to the fact that it will delete the WU before th client had a chance to upload a partial result !!!

I even had 2-3 WU which erroned out with a valid error then prepare to upload the partial results. In that time, a 0x7b error popped and deleted before it's uploaded :mad:

My guess is if you look in your work directory it’s full of old junk. Delete it. Delete que.dat and log files then restart the program.

The smp client isn’t the best at cleaning up after itself so sometimes a little sweeping is needed;)
 
God it would be nice if they fixed the issue about it cleaning up after itself! That little bit would be a nice change... :) hehehehe somethign so small but yet so nice.

I don't know exactly what I can and cannot delete in the F@H while an SMP is running. Is there documentation on what is old stuff versus current?

 
Maybe it would be a good addition to notfreds diskless client. Maybe even a program that deletes it every few days or so.
 
BillR, don't worry about me, I'm great at cleaning the poo mess the client left (thank you qfix for the help). Also, this didn't happen for a few months now (knock on wood).

Ax, you can download qfix and run it. I bet it will find the result file, rebuild the queue then use -send all to send the partial work for partial credit (which should be a nice chunk with the 95% of the WU done).

 
BillR, don't worry about me, I'm great at cleaning the poo mess the client left (thank you qfix for the help). Also, this didn't happen for a few months now (knock on wood).

Ax, you can download qfix and run it. I bet it will find the result file, rebuild the queue then use -send all to send the partial work for partial credit (which should be a nice chunk with the 95% of the WU done).


On the latest LinSMP beta, the queue.dat file is no longer recognized by qfix, may be the same for the latest WinSMP beta as well.
 
untitled-1.jpg
 
Back
Top