- Joined
- Jan 28, 2004
- Messages
- 6,301
I checked one of my Borgs at work today that ws folding, only to find yet another failed WU. Keep in mind this is on a bone stock 3.4 Ghz P4 with 2Gigs of Corsair DDR2- no overclocking or anything. It was from Project: 2653 (Run 19, Clone 50, Gen 61). This is what I had in the log:
I've seen several WU's end up like this over the past month, and all on different machines. It wouldn't be so bad if they failed within the first 10% of the WU, but to get all the way to 95% and then crash.....
I'm 13% into the new one it picked up now, so I'm hoping since this is a different run it will go all the way through.
Just thought I'd share in the agony.....

Code:
[05:28:39] Completed 465000 out of 500000 steps (93 percent)
[05:43:40] Timered checkpoint triggered.
[05:54:07] Writing local files
[05:54:08] Completed 470000 out of 500000 steps (94 percent)
[06:09:07] Timered checkpoint triggered.
[06:19:33] Writing local files
[06:19:33] Completed 475000 out of 500000 steps (95 percent)
[06:20:46] Gromacs error.
[06:20:46]
[06:20:46] Folding@home Core Shutdown: UNKNOWN_ERROR
[06:20:46]
[06:20:46] Folding@home Core Shutdown: UNKNOWN_ERROR
[09:21:37] CoreStatus = 7B (123)
[09:21:37] Client-core communications error: ERROR 0x7b
[09:21:37] Deleting current work unit & continuing...
[09:23:59] - Warning: Could not delete all work unit files (6): Core returned invalid code
[09:23:59] Trying to send all finished work units
[09:23:59] + No unsent completed units remaining.
[09:23:59] - Preparing to get new work unit...
[09:23:59] + Attempting to get work packet
[09:23:59] - Will indicate memory of 2023 MB
[09:23:59] - Connecting to assignment server
[09:23:59] Connecting to http://assign.stanford.edu:8080/
[09:24:00] - Couldn't send HTTP request to server
[09:24:00] + Could not connect to Assignment Server
[09:24:00] Connecting to http://assign2.stanford.edu:80/
[09:24:00] Posted data.
[09:24:00] Initial: 40AB; - Successful: assigned to (171.64.65.64).
[09:24:00] + News From Folding@Home: Welcome to Folding@Home
[09:24:00] Loaded queue successfully.
[09:24:00] Connecting to http://171.64.65.64:80/
[09:24:04] Posted data.
[09:24:04] Initial: 0000; - Receiving payload (expected size: 2427552)
[09:24:07] - Downloaded at ~790 kB/s
[09:24:07] - Averaged speed for that direction ~607 kB/s
[09:24:07] + Received work.
[09:24:07] + Closed connections
[09:24:12]
[09:24:12] + Processing work unit
[09:24:12] Core required: FahCore_a1.exe
[09:24:12] Core found.
[09:24:12] Working on Unit 07 [May 14 09:24:12]
[09:24:12] + Working ...
[09:24:12] - Calling 'mpiexec -channel auto -np 4 FahCore_a1.exe -dir work/ -suffix 07 -checkpoint 15 -forceasm -verbose -lifeline 3604 -version 591'
[09:24:13]
[09:24:13] *------------------------------*
[09:24:13] Folding@Home Gromacs SMP Core
[09:24:13] Version 1.74 (March 10, 2007)
[09:24:13]
[09:24:13] Preparing to commence simulation
[09:24:13] - Ensuring status. Please wait.
[09:24:30] - Assembly optimizations manually forced on.
[09:24:30] - Not checking prior termination.
[09:24:37] - Expanded 2427040 -> 12916317 (decompressed 532.1 percent)
[09:24:37] - Starting from initial work packet
[09:24:37]
[09:24:37] Project: 2653 (Run 31, Clone 3, Gen 60)
[09:24:37]
[09:24:37] tray files
[09:24:37] - Starting from initial work packet
[09:24:37]
[09:24:37] Project: 2653 (Run 31, Clone 3, Gen 60)
[09:24:37]
[09:24:38] Assembly optimizations on if available.
[09:24:38] Entering M.D.
[09:24:44] Rejecting checkpoint
[09:24:46] Protein: Protein in POPC
[09:24:46] Writing local files
[09:24:48] Extra SSE boost OK.
[09:24:48] Writing local files
[09:24:48] Completed 0 out of 500000 steps (0 percent)
I've seen several WU's end up like this over the past month, and all on different machines. It wouldn't be so bad if they failed within the first 10% of the WU, but to get all the way to 95% and then crash.....
Just thought I'd share in the agony.....


