• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Strange performance on nVidia clients

APOLLO

[H]ard|DCer of the Month - March 2009
Joined
Sep 17, 2000
Messages
9,089
One of my systems is running four nVidia 8800GT cards and lately I've been noticing a bizarre performance phenomenon. On all the WUs, I see the expected performance but the 353pt WUs have dropped a massive 1600 PPD on each card!! That comes out to 6400 PPD drop on that machine when all the cards are processing these WUs! :eek:

Now, what the heck could be causing such a disparity between the performance of a single WU compared to all others? Why would one WU type drop and the others are unaffected? None of my other machines are displaying this disparity. They all show a much higher PPD with the 353pt WUs like before. :confused:
 
are you running SMP along? Maybe they need more CPU power?
Not yet on this system. I have 8 standard clients and isolated an entire core for the GPUs. This system has been running fine since the beginning of the year with the same setup. Nothing has been changed so there really isn't a reason why one type of WU is undergoing a massive drop in performance. All the other WUs, 787pt, 1888pt, etc., are processing at normal speed.
 
Are there different 353 pewnters that could just be more chewy than other 353 pewnters? Just guessing here as I haven't done any GPU folding in months.
 
Are there different 353 pewnters that could just be more chewy than other 353 pewnters? Just guessing here as I haven't done any GPU folding in months.
No, they're all pretty much the same and haven't seen a huge variance between 353pt WUs, maybe 100-200 PPD variance but that is about it.

OK, one of the four clients finished a WU and DL a 472 pointer. Now, the other three 353pt WUs just shot up to regular speed!! Why is this happening? As long as I don't have more than three 353 pointers, everything is normal. This never happened before. Freaking weird.
 
My 353s vary slightly, but it's because I'm also running SMP on the systems.

P.S. Plus I'm not running more than 3 gpu clients per box.
 
My 353s vary slightly, but it's because I'm also running SMP on the systems.

P.S. Plus I'm not running more than 3 gpu clients per box.
Well, it's been mine and others' experience that as long as you allocate sufficient CPU power to the GPU clients, they should all process at full speed no matter how many GPUs or cards you install, and that is what I have experienced. There's one full core devoted to all the GPU clients and there was never a problem until recently...

Anyway, I think I can come up with one probable cause, but not a reason. A couple of days ago, I swapped one 8800GT from another system and took one out of this one. They aren't the same models. One is a BFG and the other an XFX. This is the only hardware change that was made, but it doesn't really qualify as a change since they're the same GPU architecture. :confused:
 
I've seen some wiggle on some of the WU for my GPU clients.

It seems that it ebbs and flows when WU are rock solid on the ppd or if they wiggle.

For the most part I just let them run. Many a time when I check on them in an hour or so they are back to where I expect them.
 
I've seen some wiggle on some of the WU for my GPU clients.

It seems that it ebbs and flows when WU are rock solid on the ppd or if they wiggle.

For the most part I just let them run. Many a time when I check on them in an hour or so they are back to where I expect them.
This phenomenon is beyond the normal variance. I'm talking about a 1.5k drop on every client when all are running 353s. Besides, I know something is out of whack because they're all processing at full speed now that one client is not folding a 353 pointer. I'm going to try swapping PCIE slots and hope I can come up with a combination that won't afflict performance. There's something up with brand mixing.
 
I'm going to try swapping PCIE slots and hope I can come up with a combination that won't afflict performance. There's something up with brand mixing.
I'd put the odd-ball device in the first PCIE slot.
 
I just got an oddball 353 WU, too. It's going to take 11 hours! *8500GT on a 'normal' 353 takes 11 hours, to...)
 
I just got an oddball 353 WU, too. It's going to take 11 hours! *8500GT on a 'normal' 353 takes 11 hours, to...)
Stop it and restart it. If that doesn't work, restart the system. My problem is definitely a hardware conflict. I checked the Device Manager and I have an exclamation mark on an unknown PCI device...
 
Project: 5791 (Run 5, Clone 558, Gen 4)

it's at 71% and I'm just going to sit through it... only 1 more hour to go (though that is way long...)
 
Project: 5791 (Run 5, Clone 558, Gen 4)

it's at 71% and I'm just going to sit through it... only 1 more hour to go (though that is way long...)
I would press Ctrl-C and restart the client. You have nothing to lose and possibly multiply the speed of the WU. Sometimes Stanford sends slowed-down WUs to alleviate server backups.
 
apollo are both cards running the same speeds? check and see in gpu-z or something if one of the cards all of sudden downclocked.. maybe its overheating or something..
 
apollo are both cards running the same speeds? check and see in gpu-z or something if one of the cards all of sudden downclocked.. maybe its overheating or something..
I checked everything and there doesn't seem to be any reason I could find. For most of the year, I had identical performance to my other systems with the same GPU clocks. When I recently swapped one card for another one of different make into this system, the 353pt WUs started processing slow when all clients had them. To be honest, I'm not certain if the problem arose immediately after I made the swap. It could have started a few days later. I only noticed the disparity a couple of days ago.

What we have to understand here is that the slowdown only happens with the 353pt WUs, and only when all four clients are processing them. Any other WU combination sees full speed. No way this is a clock frequency issue.

I will change the setup later this evening and post what happens.
 
yeah!!!! Turns out, the slow one I got was a massive, 600,000 thing job. Now I get these petite 15,000 little zappers!


yes! The way it's meant to be done!

edit: oops... that was borderline thread crap :(

EDIT2: okay... I've done at least 4WU today... including that monster... what's up with this!?

EDIT3: heh. I move up over 1600 places
 
Last edited:
OK, did a lot of fiddling last night to try and get some answers for the anomalies I was seeing. Well, I couldn't get it resolved was forced to swap back the old card that was in the system originally. Now, everything is back to normal performance-wise.

The only thing I can think of that was causing a drop in performance the different manufacturers. Brand does make a difference sometimes in some systems. These were nearly identical cards in all respects except for manufacturer. For those who intend to build multi-GPU rigs, especially if you're thinking of 3 or more GPUs, try as much as you can to stay with the same make and model. You could save yourself a lot of potential issues
 
Very weird indeed Apollo. Very weird. *scratches head*
 
Different work units use system resources differently. It could be that the bandwidth required from the PCI-e slots with 4 353-pointers is just too much for the system to handle efffectively, and that's bottlenecking your ppd, whereas the 472-point ones may use less bandwidth and more floating point performance. Just my guess.
 
Very weird indeed Apollo. Very weird. *scratches head*
Tell me about it. It's too bad because this card is a dual-slot card and there's very little room in the case. That's why I wanted to use the other card which is single-slot. Unfortunately the single-slot card was affecting all my clients. Alone or in another system it works fine, but not in this box with 4 GPUs. :rolleyes:
 
Different work units use system resources differently. It could be that the bandwidth required from the PCI-e slots with 4 353-pointers is just too much for the system to handle efffectively, and that's bottlenecking your ppd, whereas the 472-point ones may use less bandwidth and more floating point performance. Just my guess.
No. The old card was in my system since spring and I was seeing 4x 5600 PPD when I had 353-pointers across all clients. Everything was fine until I swapped the original card out. This is a Skulltrail motherboard with quad16x PCI-E slots, thus no possible bandwidth issues with four 8800GT cards, even OC'd.
 
seems like nvidia and the board partners have frequently released updated revisions of certain GPU cores or video card models. I'm suspecting if you really boiled the problem card down, you would find a different revision or something that wouldn't normally be apparent.

THat's just my BS answer as I'm sitting here, anyway.
 
seems like nvidia and the board partners have frequently released updated revisions of certain GPU cores or video card models. I'm suspecting if you really boiled the problem card down, you would find a different revision or something that wouldn't normally be apparent.

THat's just my BS answer as I'm sitting here, anyway.

It's not a BS answer.
 
seems like nvidia and the board partners have frequently released updated revisions of certain GPU cores or video card models. I'm suspecting if you really boiled the problem card down, you would find a different revision or something that wouldn't normally be apparent.

THat's just my BS answer as I'm sitting here, anyway.
It's very possible, and I have no doubt there might be a shred of truth in what you're saying and my issues. I'm not done testing. For now the old card stays but might locate yet another card for further testing because I really want this card out of the system since it's taking two slots and I need them.
 
Back
Top