• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Problems with new nVidia GPU WUs

APOLLO

[H]ard|DCer of the Month - March 2009
Joined
Sep 17, 2000
Messages
9,089
I'm having frequent problems with the P10xxx WUs recently launched by Stanford. Seems they are crashing frequently and many times this leaves me with no alternative but to reboot my system. Fed up with restarting systems just to get my GPU clients working again that should work without a restart. Anyone else running into problems with the new WUs?
 
i had some problems with the 10105 WU i believe.. but havent had a problem since i was forced to start using windows xp again.. but then again because of nvidia's pos drivers ive been stuck having to run a lower overclock then i was at the time when i was using windows 7.. when i was running my shaders at 1512 i started seeing the problem more often.. running at 1458 it hasnt been an issue but i havent figured out if its OS related or because i dropped the clocks.. temps are actually higher now then they were when i had the shaders at 1512..

if you dont want to keep restarting just leave a shortcut to the nvidia drivers extracted folder.. believe its C:\nvidia\driverversion then just reinstall the drivers.. and boom you have 3D clocks again..
 
i had some problems with the 10105 WU i believe.. but havent had a problem since i was forced to start using windows xp again.. but then again because of nvidia's pos drivers ive been stuck having to run a lower overclock then i was at the time when i was using windows 7.. when i was running my shaders at 1512 i started seeing the problem more often.. running at 1458 it hasnt been an issue but i havent figured out if its OS related or because i dropped the clocks.. temps are actually higher now then they were when i had the shaders at 1512..
Come to think of it, most of my problems are occurring on my Win 7 machines which now compose half my farm. Definitely my OCs are not the issue since sometimes the problems affect cards that are only seeing mid-80s core temps, whereas cards in neighboring slots will not EUE even if their temps are pushing 100... This is a software issue for sure, it's a little unclear where the center of the issue is, though. :confused:

if you dont want to keep restarting just leave a shortcut to the nvidia drivers extracted folder.. believe its C:\nvidia\driverversion then just reinstall the drivers.. and boom you have 3D clocks again..
I was just about to reinstall drivers however my issue is not the clocks because my cards stick to 3D, and performance is normal when there's no EUEs. It's just the bloody 24h pause that I know cannot be resolved without a restart. Now with a system that has 4 GPUs and a CPU client, well you can see what I mean about frequently restarting...
 
Come to think of it, most of my problems are occurring on my Win 7 machines which now compose half my farm. Definitely my OCs are not the issue since sometimes the problems affect cards that are only seeing mid-80s core temps, whereas cards in neighboring slots will not EUE even if their temps are pushing 100... This is a software issue for sure, it's a little unclear where the center of the issue is, though. :confused:

I was just about to reinstall drivers however my issue is not the clocks because my cards stick to 3D, and performance is normal when there's no EUEs. It's just the bloody 24h pause that I know cannot be resolved without a restart. Now with a system that has 4 GPUs and a CPU client, well you can see what I mean about frequently restarting...


out of curiousity.. next time it happens.. try closing explorer.exe in the task manager then going to run.. and open explorer.exe again and see if the client allows you to run it again without restarting your system.. by default the client should see that as a restart since thats what its technically doing.. but its definitely a software problem.. either the way the WU's using the gpu and conflicting with windows using the gpu as well to run aero or something else.. but its only the 101xx WU's that have the problem.. and mostly the 10105 WU..
 
out of curiousity.. next time it happens.. try closing explorer.exe in the task manager then going to run.. and open explorer.exe again and see if the client allows you to run it again without restarting your system.. by default the client should see that as a restart since thats what its technically doing.. but its definitely a software problem.. either the way the WU's using the gpu and conflicting with windows using the gpu as well to run aero or something else.. but its only the 101xx WU's that have the problem.. and mostly the 10105 WU..
Yes, I totally agree. Actually, I had restarted my system and the problem is still plaguing one of my GPUs. I am in the middle of reinstalling the drivers now, and will be using a driver cleaner to make certain of a clean install. The WU that is giving me big problems as of this writing is a P10503. I don't have any other GPU client running this one and no matter how many times it errors, Stanford still sends it back to this client... :(

PS. I don't have Aero running on this system, should I enable it?
 
Yes, I totally agree. Actually, I had restarted my system and the problem is still plaguing one of my GPUs. I am in the middle of reinstalling the drivers now, and will be using a driver cleaner to make certain of a clean install. The WU that is giving me big problems as of this writing is a P10503. I don't have any other GPU client running this one and no matter how many times it errors, Stanford still sends it back to this client... :(

PS. I don't have Aero running on this system, should I enable it?


you could try enabling it.. really wont effect your folding unless you are running a card with 512mb or less video memory.. since aero uses about 100-150mb of video memory then 350-385mb for the WU.. just check on evga precision to see how much memory is being used on the card with it off and with it on..
 
I've not seen any issues with the 10xxx series WUs, but I'm also no longer running GPU2 on any of my daily drivers (laptop and AMD rig w/8800GTX). I turn everything off when running my sig rig as I don't want to have to deal with any contention.

 
I haven't had any crashes, but I've had EUE's lately (can't recall if it's a specific WU# or not though). Moreso on my 8800GT, par for the course there - but I've had a few on my 260's also.
 
I haven't had any crashes, but I've had EUE's lately (can't recall if it's a specific WU# or not though). Moreso on my 8800GT, par for the course there - but I've had a few on my 260's also.
I'm at my rope's end. All my nVidia cards are 8800GT. It could just be a bad card but the temps are OK, even tried it at stock, and everything was working fine until I installed Win 7 on this machine about a week ago. I keep on getting the "Nonzero force sum on GPU" message when the client errors. Anyone know what this means?

BTW, I'm using the v181.20 nVidia drivers and this is Win 7 Ultimate x64. The machine has four 8800GT cards.

I might try swapping slots with another card later tonight, and see if it's the card or possibly something else that is causing the problem. Fortunately, I have spare 8800GT cards not being used so if worse comes to worst I can always replace it. Somehow though, I don't think the card is at fault...
 
I'm at my rope's end. All my nVidia cards are 8800GT. It could just be a bad card but the temps are OK, even tried it at stock, and everything was working fine until I installed Win 7 on this machine about a week ago. I keep on getting the "Nonzero force sum on GPU" message when the client errors. Anyone know what this means?

BTW, I'm using the v181.20 nVidia drivers and this is Win 7 Ultimate x64. The machine has four 8800GT cards.

I might try swapping slots with another card later tonight, and see if it's the card or possibly something else that is causing the problem. Fortunately, I have spare 8800GT cards not being used so if worse comes to worst I can always replace it. Somehow though, I don't think the card is at fault...

Are you sure that the 181.2 drivers are win 7 compatable?? Try 191.70 - they seem to be quite a stable set - quite a few of the guys over on FF use them
 
All I get on my 8800GT are EUE errors. Thank God for HFM.net reporting them (emailing me) now when they happen. I have to shutdown/reopen my gpu client on that machine several times a day because of it.

(still running 195.62 on everything)

I mutter under my breath practically every day that 8800GT's are $#!t for folding. :mad:
 
Are you sure that the 181.2 drivers are win 7 compatable?? Try 191.70 - they seem to be quite a stable set - quite a few of the guys over on FF use them
Great, I just realized I don't have the correct driver set. I have the 191.7s running on my other Win 7 machines. Somehow I DL the wrong drivers for this system when I switched to Win 7 last week. :(

OK, I'll look to see if I still have the 191s and if not I'll just DL them again. Hopefully, that is at the crux of the problem. Can't believe I didn't notice this before... :rolleyes:
 
Well, I spoke too soon.

8800GTX @ stock clocks @ 100% fan just up and EUE'd on me. It was running the 10xxx series WU. :(

I am currently running Win7 x64 and latest nVidia drivers. Grabbed them off of eVGA website.
 
Now I'm getting those EUE's.

It's a conspiracy to rob us nVidia owners of PPD to make the ATI folks feel better. ;)
 
I've had problems on and off with the 10xxx units since they were released on my 8800GT. I do not have anything overclocked really high (core: 675, shader: 1728, memory: 1000) and my temps usually don't go much above 70C even with a single slot cooler. Well, that's as long as the room temp doesn't go above 80F. I have actually lowered the shader clock considerably from around 1800 just to make sure I had no problems. The RAM and core clocks have never had an effect on stability so they stay where they are.

This is the card in my main system so I'm not going to be lowering the clocks on the core or RAM as they help a bit for gaming. Also, I've never noticed more than a 1C drop in temps after dropping both core and RAM speeds to about as low as they can go.

I've run these work units on Vista and 7 with different revisions of drivers and I get the same behavior no matter what.

Since I've had the EUE problems with the 10xxx units since they were released, my guess is that they are rather unstable work units. They don't always EUE and sometimes I'll go through plenty of them without seeing an EUE. I may miss some as I don't go back and check to see if any have EUE'd unless I see some EUE. I don't have a problem with any other work units which is why I think the problem resides with these work units.

 
I shouldn't have opened my big mouth. Now my 8800GTX @ stock clocks will NOT do a 10xxx WU to save its life :(

It'll do any other WU, but not them.
 
OK, time for an update. I'm more confused than before. The install to get my system running the 191.70 drivers has had mixed success. Gone are the continuous crashes on GPU #2 with every WU but the P10503 WUs will not complete worth a ****. Why GPU #2 and not the other 3 GPUs on this system is way beyond me. Any other WUs will have no problem completing and any other GPU client has no problem at all.

So, what happens in a nutshell is everything runs smoothly until it's time for Stanford to release a run of P10503s when I experience continuous errors, and the 24h timeout is imposed. :confused:
 
Count me in the confusion club. :confused: When I went to check my PC after I got home from work today I was greeted by both of my GPUs having EUE's. :mad::mad: I restarted both of them and then I had to go back to work, so I'll get to see what they are up to when I get home.
 
I've had issues with them, but no EUEs, but I just got back from funeral detail and I guess my quadrig (295s) locked up on these some time yesterday. Was 105xx unit, and restart the rig, 3 cards run full blast, one card, number 3 always for me, goes to half speed on these units. Not 2d mode, just half speed. I can run multiple of the same runs and the rest run fine, but whenever 3 pulls that unit now, half speed.

I'm down to 2 cards running in there anyways, my AC has never been run at this new house and it's not running now, so when it gets 80+ outside, these cards run hot, and make it hotter inside! Slowly moving my way to pure SMP servers in the rack in the much cooler folding room.
Posted via [H] Mobile Device
 
I checked through my logs and lo and behold, the WU that has been giving me all these problems is the SAME ONE!! Stanford keeps on sending me this same WU over and over regardless how many times it never completed and now it has been days of this. How do I stop a corrupt WU from being returned to the same client? Surely, there must be some way?? :mad:
 
I delete the work directory and several of the other files in the main gpu directory like unitinfo, queue.dat, the core...basically everything but the client config. When you restart the client it generates all new files and has always downloaded a new wu. Sorry, I am not at home to see exactly what files to delete but it works.
 
I delete the work directory and several of the other files in the main gpu directory like unitinfo, queue.dat, the core...basically everything but the client config. When you restart the client it generates all new files and has always downloaded a new wu..
Yep, I tried this yesterday and it seemed to work for the rest of the day. I didn't see that WU come up. When I checked the logs today however, I noticed the same WU had been DL again several times early in the morning with the same inevitable results, except the client did not timeout this time. Fortunately, another WU was eventually DL and the client had recommenced work.

The fact that this specific client is attracting the exact same corrupt WU over and over again for days points to some identifier, and that leads me to conclude it can only be the client.cfg file. Assuming this supposition is correct, will altering the details of the .cfg file prevent this client from DL that same corrupt WU again?
 
I too have been having MASSIVE issues with my second card crashing from these units, and twice it has caused a bsod lockup, making my system worthless usually right after I go to bed. It was enough that I was set on quitting folding until I saw this thread. Weird that only my second 260 shows these problems, much like your second gpu is only showing them. I will try running just one gpu and cpu for a while then, until these units are gone. If it bluescreens again though I quit.
 
I too have been having MASSIVE issues with my second card crashing from these units, and twice it has caused a bsod lockup, making my system worthless usually right after I go to bed. It was enough that I was set on quitting folding until I saw this thread. Weird that only my second 260 shows these problems, much like your second gpu is only showing them. I will try running just one gpu and cpu for a while then, until these units are gone. If it bluescreens again though I quit.
Check your logs. Is it the same WU that is crashing your client over and over again? If so, do this: delete everything in the client's folder except the executable, client.cfg and cudart.dll files. In the client.cfg file, edit and change the machine id of the client to any other ID. Not sure this will work, but it's worth a shot if you keep on receiving the same corrupt WU that is causing crashes or timeouts.
 
I've been having fits with my GPUs lately too. It seems to be primarily happening on my secondary cards in my systems. Seems that the primary GPUs are unaffected. ANybody else seeing this?
 
I was seeing exactly that as well. Always my second gpu. I'm on 197.13 now and I've only had one driver crash, no bsod. I think the driver crash was my own fault too, so otherwise they have been stable. Something wonky with these new units indeed...
 
Running 2 gx2s in 2 Seperate rigs, each one is doing just fine
when I had my dual 295's, always card 3..
Posted via [H] Mobile Device
 
Well for anyone who has been having the 'display driver has stopped responding and has recovered' error with these, I think I found a solution... Using rivatuner you can set some advanced options and set the low power 3d clocks as well. I had to crank up the max clock value setting as well to allow me to set the clocks the proper speed, but now with the driver crashed it appears I can continue to fold and use my computer. Under the Power user tab, set EnableLowPower3DControl to 1 in Rivatuner\nvidia\Overclocking, and set MaxClockLimit to 200 (higher if needed once you see the low power 3d page) in Rivatuner\Overclocking\Global. Source: http://forums.vr-zone.com/overclockers-hideout/56017-guide-how-overclock-rivatuner.html#post2403191

 
Back
Top