• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

EUEs on GPU

Mr. Pedantic

[H]ard|Gawd
Joined
Sep 19, 2009
Messages
1,707
I'm having a huge number of these happening:

Code:
[05:08:07] Folding@Home GPU Core - Beta
[05:08:07] Version 1.24 (Mon Feb 9 11:00:12 PST 2009)
[05:08:07] 
[05:08:07] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[05:08:07] Build host: amoeba
[05:08:07] Board Type: AMD
[05:08:07] Core      : 
[05:08:07] Preparing to commence simulation
[05:08:07] - Assembly optimizations manually forced on.
[05:08:07] - Not checking prior termination.
[05:08:07] - Expanded 68535 -> 357580 (decompressed 521.7 percent)
[05:08:07] Called DecompressByteArray: compressed_data_size=68535 data_size=357580, decompressed_data_size=357580 diff=0
[05:08:07] - Digital signature verified
[05:08:07] 
[05:08:07] Project: 5745 (Run 4, Clone 46, Gen 360)
[05:08:07] 
[05:08:07] Assembly optimizations on if available.
[05:08:07] Entering M.D.
[05:08:13] Tpr hash work/wudata_02.tpr:  3093445337 740017563 3891334339 2079336442 2375787072
[05:08:13] Working on Protein
[05:08:13] Client config found, loading data.
[05:08:13] Starting GUI Server
[05:08:16] mdrun_gpu returned 
[05:08:16] Nonzero force sum on GPU
[05:08:16] 
[05:08:16] Folding@home Core Shutdown: UNSTABLE_MACHINE
[05:08:19] CoreStatus = 7A (122)

I haven't changed anything with regards to the GPU, I'm still running the 4870 at stock, and it's staying at about 50C load on 33% fan.

However, I did increase my Vcore by 0.02V, to increase CPU stability (I kept getting random restarts when my VMs were up). Good news is that the reboots and BSODs have stopped. Bad news is I get UNSTABLE_MACHINE errors from the GPU client.

Are these in any way related? Or is it just a coincidence? If anything, the machine is more stable now than it was about 3 days ago, but the errors from the GPU client have increased, and I should be getting more ppd overall from not losing SMP WUs every so often. However, having to restart the GPU client every time the EUE limit is reached (which seems to be about every second WU) is a bit annoying and calls for a level of micromanagement I'm not prepared to give.

Any advice, people?
 
only suggestion would to be delete the entire F@H GPU folder.. and make sure you delete the folder in your applications data folder.. then reinstall the gpu client.. see if that works.. the only reason that comes to mind in why this would be happening if its client related is that the BSOD's eventually corrupted something.. i find it hard to believe a simple cpu voltage change would actually cause the gpu client to all of a sudden go unstable..
 
It's the 57xx WU's, they do it all the time.
Every supposed fix I've heard about/read about to fix this, I've used and not had any luck. I've tried enabling the 2nd display port on the gpu (despite not having a monitor attached to it), I've put in the -forcegpu nvidia_g80 flag (on both of mine, since they're both nV cards)....nothing. Neither of my cards are oc'd, and I get this problem on both systems (my quad is oc'd and stable @ 3.6, the E8400 runs @ stock). I still get these stupid EUE "unstable machine" 'pausing 24 hours' messages. :mad: The only thing you can really do is keep an eye on the client and close it then reopen it when you have this happen.

I check my clients every few hours during the day and then right before I go to bed and right when I wake up in the morning, to make sure they're running.

Lots of people have this problem...seems like (going off the Stanford forums) it's the 57xx series WU's, something about them (no one knows). :(
 
I wanted to bump this because I just had the same issue as the OP. But got another question about it;

This is my first encounter with any "UNSTABLE_MACHINE" core errors. I was just browsing the web, then my display froze for a second, then the nVidia tray said it just did a VPU recovery. Saw my GPU2 client stalled in FahMon and here's what the log said:

Code:
[04:50:35] Completed 87%
[04:51:15] Completed 88%
[04:51:35] Run: exception thrown during GuardedRun
[04:51:35] Run: exception thrown in GuardedRun -- Gromacs cannot continue further.
[04:51:35] Going to send back what have done -- stepsTotalG=10000000
[04:51:35] Work fraction=0.8845 steps=10000000.
[04:51:39] logfile size=0 infoLength=0 edr=0 trr=23
[04:51:39] - Writing 642 bytes of core data to disk...
[04:51:39] Done: 130 -> 127 (compressed to 97.6 percent)
[04:51:39]   ... Done.
[04:51:40] 
[04:51:40] Folding@home Core Shutdown: UNSTABLE_MACHINE
[04:51:44] CoreStatus = 7A (122)
[04:51:44] Sending work to server
[04:51:44] Project: 5761 (Run 11, Clone 588, Gen 2)
[04:51:44] - Read packet limit of 540015616... Set to 524286976.


[04:51:44] + Attempting to send results [November 19 04:51:44 UTC]
[04:51:45] + Results successfully sent
[04:51:45] Thank you for your contribution to Folding@Home.

So did I still get credit for that incomplete core? If so, that's still cool with me, heh.

I just noticed that it was a P57xx core as you mentioned, zero. I just pulled another 57xx Wu though and my GPU1 is 50% through a P57xx also. Hopefully they get through these WUs though. For some reason after that VPU recovery, I can't get my GPU2 clocks to go back up to performance 3D clocks, it's stuck in low-power 3D performance clocks which is getting me half the PPD it's supposed to be doing. Is there a way to force the GPU back to performance 3D clocks without restarting the PC? I'm using Rivatuner to monitor all current clock speeds and temps.

Thanks!
 
Last edited:
No, I don't think you do. In any case, at least your client tells you something happened. If my 4870 does a VPU recover, the client just stalls, with no indication that it's stopped folding or why. It's only if you look in the CCC or on Rivatuner and you see that GPU load and clocks are at idle, then you know something's wrong.

In any case, I found a solution. Deleting the Work folder solves the problem. Possibly a corrupted queue log or something that was causing errors.

Is there a way to force the GPU back to performance 3D clocks without restarting the PC? I'm using Rivatuner to monitor all current clock speeds and temps.
I get that sometimes too. If you use FahSpy, since it has the ability to restart or stop services, you can use that to restart the GPU client service. That usually does it. Or you can start the service manually...or just quit it and run the GPU client executable again.
 
Well it seemed to finish the other P57xx's over the night. So I guess it was just a bad WU.

I tried restarting the GPU client like 5 times and it wouldn't go back to the high performance 3D clock speed. I think the VPU recovery locked it down somehow. I had to just restart the PC. Weird stuff. I'm having a hard enough time keeping my linux VM SMP client finishing WUs, I'm gonna be mad if these GPU WUs keep stalling the clients too :(.
 
^^^ Yeah, I've had a few 57xx's finish fine.
Lately I haven't gotten this error as much, so hopefully the buggy WU's are finished.
 
I'm having a hard enough time keeping my linux VM SMP client finishing WUs,

Yeah I've got a Win7 64 box running NFreds/ VMWare that will apparently complete the unit to 100% and then just sit there and not do another thing unless I remove all files in it's folder and start it over again, but then it repeats. :confused:
 
Yeah I've got a Win7 64 box running NFreds/ VMWare that will apparently complete the unit to 100% and then just sit there and not do another thing unless I remove all files in it's folder and start it over again, but then it repeats.
You do know that after the WU finishes it takes a while for the cores to synchronize? Maybe that's when you delete the WU? And it also takes quite a while to upload the unit; it's about 50MB up.
 
You do know that after the WU finishes it takes a while for the cores to synchronize? Maybe that's when you delete the WU? And it also takes quite a while to upload the unit; it's about 50MB up.

Yeah, I'm coming back and checking the box hours later and it's doing nothing. My other box with VMware is doing fine, always running.
 
How come? I haven't lost an SMP WU for weeks and weeks.

I have to keep babysitting it because it stops in the middle of a WU and says this:

Code:
[01:41:26] Completed 115000 out of 250000 steps  (46%)
FahCore_a2.exe[2380]: segfault at b6f93eo ip 00000000006a22027 sp 0000000040abe9a0 error 4 in
[01:41:40] 
[01:41:40] Folding@home Core Shutdown: INTERRUPTED application called MPI_Abort(MPI_COMM_WORLD, 102) - process 0
[01:41:48] CoreStatus = 66 (102)
[01:41:48] + Shutdown requested by user. Exiting.
Folding@Home Client Shutdown.

I dunno wtf is "interrupting" it, but it's pretty annoying and I'm about to go back to Notfreds if I don't figure it out soon. This is using the linux image from this guide here (that Kendrak recommended) running -smp 8 in low priority.

Anyone have any ideas? I'm about to make a thread on it if no one sees it here.

Thanks!
 
Last edited:
Time to delete the VM and make a new one.

Good suggestion. I did this last night, new iso and vmx and so far the box is continuing on vs just completing one set and halting. Unfortunately I didn't read your post till this morning, but you still get some credit because it appears to be correct.
thumbsup.gif
 
Good suggestion. I did this last night, new iso and vmx and so far the box is continuing on vs just completing one set and halting. Unfortunately I didn't read your post till this morning, but you still get some credit because it appears to be correct.
thumbsup.gif
Better late than never :cool:.
 
Hey, I just wanted to add that I found a way to force 3D clocks on the GPUs through Rivatuner. Just follow this guide here. It seems to be working so far for me because I had this problem randomly. It should prolly be noted somewhere like the GPU guide because I think this would be an issue with a decent amount of people.
 
Back
Top