• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Constant Windows SMP Crash

sachem87184

[H]ard|Gawd
2FA
Joined
Nov 6, 2007
Messages
1,593
This is a log of the error. The only way I can get it to continue is to delete the queue.dat file and the work folder. The problem is thought that even after that the new WU will go to 17% again and crash. Any help will be greatly appreciated.

[02:31:25] Writing local files
[02:31:26] Completed 40000 out of 250000 steps (16 percent)
[03:18:28] Writing local files
[03:18:28] Completed 42500 out of 250000 steps (17 percent)
[03:40:06] Warning: long 1-4 interactions
[03:40:06] Gromacs cannot continue further.
[03:40:06] Going to send back what have done.
[03:40:06] logfile size: 41984
[03:40:06] - Writing 42520 bytes of core data to disk...
[03:40:06] ... Done.
[03:40:07] - Failed to delete work/wudata_01.sas
[03:40:07] - Failed to delete work/wudata_01.goe
[03:40:07] Warning: check for stray files
[03:42:07]
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:07]
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:11] CoreStatus = 7B (123)
[03:42:11] Client-core communications error: ERROR 0x7b
[03:42:11] This is a sign of more serious problems, shutting down.


Log upon restart of app:

[13:02:25] + Processing work unit
[13:02:25] Work type a1 not eligible for variable processors
[13:02:25] Core required: FahCore_a1.exe
[13:02:25] Core found.
[13:02:25] Using generic mpiexec calls
[13:02:25] Working on queue slot 01 [January 23 13:02:25 UTC]
[13:02:25] + Working ...
[13:02:25]
[13:02:25] *--------- Looking at optimizations...
[13:02:43] - Working with standard loops on this execution.
[13:02:43] - Previous termination of core was improper.
[13:02:43] - Going to use standard loops.
[13:02:43] - Files status OK
[13:04:43]
[13:04:43] Folding@home Core Shutdown: MISSING_WORK_FILES
[13:04:43] Finalizing output
[13:04:45] CoreStatus = 1 (1)
[13:04:45] Client-core communications error: ERROR 0x1
[13:04:45] This is a sign of more serious problems, shutting down
 
Is this a new SMP install? It looks like something isn't setup right. It looks like you are doing Uniprocessor work for one thing. Give us a little more of the log file. What work unit are you working on here?
 
Can you paste the start of the log before the crash so I can see the project number and run,clone,gen to see if it's a reported bad unit.
 
This is the log from after the 1st time it crashed and I deleted the work and queue data. (Modified so its not really long)




Launch directory: C:\Program Files\Folding@Home Windows SMP Client V1.01
Executable: C:\Program Files\Folding@Home Windows SMP Client V1.01\Folding@home-Win32-x86.exe
Arguments: -smp

[13:54:51] - Ask before connecting: No
[13:54:51] - User name:
[13:54:51] - User ID:
[13:54:51] - Machine ID: 1
[13:54:51]
[13:54:51] Could not open work queue, generating new queue...
[13:54:51] - Preparing to get new work unit...
[13:54:51] + Attempting to get work packet
[13:54:51] - Connecting to assignment server
[13:54:52] - Successful: assigned to (171.64.65.64).
[13:54:52] + News From Folding@Home: Welcome to Folding@Home
[13:54:52] Loaded queue successfully.
[13:55:15] + Closed connections
[13:55:15]
[13:55:15] + Processing work unit
[13:55:15] Work type a1 not eligible for variable processors
[13:55:15] Core required: FahCore_a1.exe
[13:55:15] Core found.
[13:55:15] Using generic mpiexec calls
[13:55:15] Working on queue slot 01 [January 22 13:55:15 UTC]
[13:55:15] + Working ...
[13:55:15]
[13:55:15] *------------------------------*
[13:55:15] Folding@Home Gromacs SMP Core
[13:55:15] Version 1.74 (March 10, 2007)
[13:55:15]
[13:55:15] Preparing to commence simulation
[13:55:15] - Files status OK
[13:55:24] - Expanded 4803458 -> 24810145 (decompressed 516.5 percent)
[13:55:24] - Starting from initial work packet
[13:55:24]
[13:55:24] Project: 2665 (Run 2, Clone 89, Gen 86)
[13:55:24]
[13:55:25] Assembly optimizations on if available.
[13:55:25] Entering M.D.
[13:55:47] ial work packet
[13:55:47]
[13:55:47] Project: 2665 (Run 2, Clone 89, Gen 86)
[13:55:47]
[13:55:53] check for stray files
[13:55:53] - Starting from initial work packet
[13:55:53]
[13:55:53] Project: 2665 (Run 2, Clone 89, Gen 86)
[13:55:53]
[13:55:55] Entering M.D.
[13:56:04] Rejecting checkpoint
[13:56:07] Protein: HGG with glycosylations
[13:56:07] Writing local files
[13:56:22] Extra SSE boost OK.
[13:56:23] Writing local files
[13:56:24] Completed 0 out of 250000 steps (0 percent)
[14:43:18] Writing local files
[02:31:26] Completed 40000 out of 250000 steps (16 percent)
[03:18:28] Writing local files
[03:18:28] Completed 42500 out of 250000 steps (17 percent)
[03:40:06] Warning: long 1-4 interactions
[03:40:06] Gromacs cannot continue further.
[03:40:06] Going to send back what have done.
[03:40:06] logfile size: 41984
[03:40:06] - Writing 42520 bytes of core data to disk...
[03:40:06] ... Done.
[03:40:07] - Failed to delete work/wudata_01.sas
[03:40:07] - Failed to delete work/wudata_01.goe
[03:40:07] Warning: check for stray files
[03:42:07]
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:07]
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:11] CoreStatus = 7B (123)
[03:42:11] Client-core communications error: ERROR 0x7b
[03:42:11] This is a sign of more serious problems, shutting down.

Folding@Home Client Shutdown at user request.

Folding@Home Client Shutdown.
 
No report of this as bad unit, I'll make one and let you know when I get a confirmation about the status of this. For now, delete the work folder and queue.dat then restart (if you get the same workunit, rinse and repeat until you get a different one).
 
I deleted the work queue and the core files 3 times and got a new unit. I will update if this one also crashes. Thanks you for such a quick response because no PPD on this boxen has been killing me. :)
 
Always check in your work folder that theres not an ophan results file in there.
From part of your logfile
Code:
[03:18:28] Completed 42500 out of 250000 steps (17 percent)
[03:40:06] Warning: long 1-4 interactions
[03:40:06] Gromacs cannot continue further.
[03:40:06] Going to send back what have done.
[03:40:06] logfile size: 41984
[03:40:06] - Writing 42520 bytes of core data to disk...
[03:40:06] ... Done.
[03:40:07] - Failed to delete work/wudata_01.sas
[03:40:07] - Failed to delete work/wudata_01.goe
[03:40:07] Warning: check for stray files
[03:42:07] 
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:07] 
[03:42:07] Folding@home Core Shutdown: EARLY_UNIT_END
[03:42:11] CoreStatus = 7B (123)
[03:42:11] Client-core communications error: ERROR 0x7b
The line at 03:40:06 says its writing data to disk so there could of been a results file from that.
But a 7b error will delete the work-unit from the queue.dat file which stops it being sent in.
To fix that, you need to run qfix.exe to repair the queue.dat file, so you can send the partial result back.
How to use Qfix
If you can send a partial result in it will give you some points for the work done, plus it tells Stanford there may be a bad work-unit in the wild.

Luck ........... :D
 
Update:

I cleared everything out, got a new WU, and everything seems to be working fine. It completed and is working on its 3rd new WU right now. "keeps his fingers crossed"
 
Back
Top