• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Ati GPU client woes

APOLLO

[H]ard|DCer of the Month - March 2009
Joined
Sep 17, 2000
Messages
9,089
I have a single Ati card running the GPU2 client. Since installing it a couple of months ago, it has failed to upload ~20% of its completed work, or roughly one in five WUs. The WUs remains in queue but never succeed to upload and eventually expire. Without actually moving the card to another system which would be difficult ATM, I cannot verify if this an issue on my end or Stanford's. I recall someone else posting about the same problem a while back, so I know that there are others who had similar issues. I tried reinstalling the client with no improvement.

Here is a typical log:

[16:36:31] Project: 5746 (Run 0, Clone 42, Gen 1)

[16:36:31] + Attempting to send results [November 30 16:36:31 UTC]
[16:36:33] - Server reports packet it received an incomplete payload.
[16:36:33] (May be due to packet loss during network transmission or a corrupted file.)
[16:36:33] - Error: Could not transmit unit 06 (completed November 30) to work server.

[16:36:33] + Attempting to send results [November 30 16:36:33 UTC]
[16:39:11] - Couldn't send HTTP request to server
[16:39:11] + Could not connect to Work Server (results)
[16:39:11] (171.67.108.17:8080)
[16:39:11] + Retrying using alternative port
[16:41:48] - Couldn't send HTTP request to server
[16:41:48] + Could not connect to Work Server (results)
[16:41:48] (171.67.108.17:80)
[16:41:48] Could not transmit unit 06 to Collection server; keeping in queue.
[16:41:48] - Preparing to get new work unit...
[16:41:48] + Attempting to get work packet
[16:41:48] - Connecting to assignment server
[16:41:49] - Successful: assigned to (171.64.65.102).
[16:41:49] + News From Folding@Home: GPU folding beta
[16:41:49] Loaded queue successfully.
[16:41:50] Project: 5746 (Run 0, Clone 42, Gen 1)

 
Have you checked your firewall settings to be sure it is letting them out? (Windows/AntiVirus/Router)
Possible network driver, router or network cable issue?
 
Have you checked your firewall settings to be sure it is letting them out? (Windows/AntiVirus/Router)
Possible network driver, router or network cable issue?
I haven't checked these because I never had such issues with the CPU clients I used to run on the same machine and the fact the majority of WUs actually succeed in uploading to Stanford. It would be relatively easy to change network cards but I think the problem lies elsewhere.
 
Just try to eliminate the basics first. Learned trouble shooting is easier from the bottom up rather than the top down.
 
Try to ping the server the client is trying to connect to. If you don't get a response, it means that you have a connection problem, or that the server is down.
 
Just try to eliminate the basics first. Learned trouble shooting is easier from the bottom up rather than the top down.
I'm going to start trying some of the other possibilities later on tonight.

Try to ping the server the client is trying to connect to. If you don't get a response, it means that you have a connection problem, or that the server is down.
Then the server has been down very frequently (several times a week) since the end of summer, if that's the case. I find that to be strange because I have had situations where more recently completed WUs succeed to upload and then the queued WUs attempt to upload immediately after but fail.

From the logs, it appears like there might be some corruption but if so it is affecting only a minority of the WUs. :confused: In case anyone inquires, the card is running on stock clocks.
 
I have a single Ati card running the GPU2 client. Since installing it a couple of months ago, it has failed to upload ~20% of its completed work, or roughly one in five WUs. The WUs remains in queue but never succeed to upload and eventually expire. Without actually moving the card to another system which would be difficult ATM, I cannot verify if this an issue on my end or Stanford's. I recall someone else posting about the same problem a while back, so I know that there are others who had similar issues. I tried reinstalling the client with no improvement.

Here is a typical log:

[16:36:31] Project: 5746 (Run 0, Clone 42, Gen 1)

[16:36:31] + Attempting to send results [November 30 16:36:31 UTC]
[16:36:33] - Server reports packet it received an incomplete payload.
[16:36:33] (May be due to packet loss during network transmission or a corrupted file.)
[16:36:33] - Error: Could not transmit unit 06 (completed November 30) to work server.

[16:36:33] + Attempting to send results [November 30 16:36:33 UTC]
[16:39:11] - Couldn't send HTTP request to server
[16:39:11] + Could not connect to Work Server (results)
[16:39:11] (171.67.108.17:8080)
[16:39:11] + Retrying using alternative port
[16:41:48] - Couldn't send HTTP request to server
[16:41:48] + Could not connect to Work Server (results)
[16:41:48] (171.67.108.17:80)
[16:41:48] Could not transmit unit 06 to Collection server; keeping in queue.
[16:41:48] - Preparing to get new work unit...
[16:41:48] + Attempting to get work packet
[16:41:48] - Connecting to assignment server
[16:41:49] - Successful: assigned to (171.64.65.102).
[16:41:49] + News From Folding@Home: GPU folding beta
[16:41:49] Loaded queue successfully.
[16:41:50] Project: 5746 (Run 0, Clone 42, Gen 1)



I posted about this problem a couple of months ago. Unfortunately I wasn't able to
get the problem corrected. There were a couple posts over on folding forum but nothing
solid that nailed the problem down. I usually uninstalled the client and re-installed, to
get it working 100%. What I mean by 100% is even after having that error the next WU
the client crunches may be fine or the next 5 but it would do it again. I did find the
more I had the card overclocked that more apt I was to run into this problem. Knock
on wood I haven't seen the error in 2 month or so. Good Luck
 
That is a really weird error.I have not seen that one error on any of my nine ATI GPU's.

I have not had a failed upload in about 3 months and my GPU hit the servers every 30 min or so due to the amount of GPUs I have.

From what I remember, the units should upload to 171.64.65.102 or 171.64.65.103 for ATI.
It should not be asking for another server like that

I have never seen another server listed in my logs, but I am at work so I will have to double check that in the morning.

What client version are you running?

Also try running with the -gpu 0 flag in the shortcut.

When I get home, I will look at my logs, because you should not be connecting to that server at all.
 
From what I remember, the units should upload to 171.64.65.102 or 171.64.65.103 for ATI. It should not be asking for another server like that
That's what I thought as well.

I have never seen another server listed in my logs, but I am at work so I will have to double check that in the morning.
Is this what is meant by a ghost server?

What client version are you running?
I'm not sure of the client version but the core version is 1.18.

Also try running with the -gpu 0 flag in the shortcut.
I tried it and restarted the client but didn't make a difference. I still have a WU in queue that won't upload.
 
open http://171.67.108.17:8080/ ... should reply OK ... works for me.
Yes, it reports OK.

171.67.108.17 is a collection server. you're not connecting to the proper GPU server for some reason.

try opening 171.64.65.102 ...
It works OK.

The client attempts to connect to server 171.67.108.17 when it cannot upload to server 171.64.65.102. Here's another portion of the log:

[21:44:56] + Attempting to send results [November 30 21:44:56 UTC]
[21:44:56] Working on queue slot 09 [November 30 21:44:56 UTC]
[21:44:56] + Working ...
[21:45:01] - Couldn't send HTTP request to server
[21:45:01] + Could not connect to Work Server (results)
[21:45:01] (171.64.65.102:8080)
[21:45:01] + Retrying using alternative port
[21:45:01] - Couldn't send HTTP request to server
[21:45:01] + Could not connect to Work Server (results)
[21:45:01] (171.64.65.102:80)
[21:45:01] - Error: Could not transmit unit 06 (completed November 30) to work server.

[21:45:01] + Attempting to send results [November 30 21:45:01 UTC]
[21:45:01] - Couldn't send HTTP request to server
[21:45:01] + Could not connect to Work Server (results)
[21:45:01] (171.67.108.17:8080)
[21:45:01] + Retrying using alternative port
[21:45:01] - Couldn't send HTTP request to server
[21:45:01] + Could not connect to Work Server (results)
[21:45:01] (171.67.108.17:80)
[21:45:01] Could not transmit unit 06 to Collection server; keeping in queue.
 
Based on your logs, i believe it is a network error on your end.

I just pulled my logs from my systems and I can't locate a server connection error anywhere over the past two weeks ...
 
Based on your logs, i believe it is a network error on your end.
Then it's a network issue that seems to be very selective and only affecting 1 in 5 GPU2 WUs. Console client WUs and the GPU1 WUs of the past were invisible to this networking issue.

I just pulled my logs from my systems and I can't locate a server connection error anywhere over the past two weeks ...
Most folders seem to be unaffected but that should not automatically imply that the problem is on my end. Months ago, I was one of the very few that experienced the Linux SMP A2 hanging client bug and actively complaining about it. At the time, my issue was determined by others on this forum and Stanford to be a user-end problem, that is, until other people started reporting it in larger numbers. As it turns out from the posts above, I'm not the only one to have experienced the queued-up expired WU syndrome despite its rarity.

Besides, I already underwent preliminary system checks by replacing cables and network cards, etc. and doubt with near certainly that this is on my end. The only thing I can do to verify it with a 100% confirmation is to remove the Ati card and install it in another system, but that would require a lot of work including new OS install which I'm not about to do unless I ascertain a higher probability there's something amiss with my setup.
 
After a couple of months not having any issues, this problem is back worse than ever. I had at least one expired WU and 3 more WUs queued up in only a few days. If this problem doesn't get resolved by the end of the week, I'm going to shut down my Ati client. There's no point of continuing to fold on this card. Has anyone else experienced similar problems with expired backed up WUs or know of others who have??

 
I've honestly never had any issues like that at all. Have you tried completely wiping the client off of your machine and setting up a fresh copy? If that doesn't work, then it could be a hardware or OS issue.
 
I've honestly never had any issues like that at all. Have you tried completely wiping the client off of your machine and setting up a fresh copy? If that doesn't work, then it could be a hardware or OS issue.
Hi, thanks for the quick reply. I reinstalled the client sometime during the end of the year holidays, IIRC, because I had to reinstall the OS for unrelated reasons. This is a two month old client/OS installation, and I don't do any other work on this particular system. At this point I'm open to any suggestion and might wipe everything again if there's a chance it could work. I'd rather try another suggestion first before doing something as drastic as a new OS, though.
 
Back
Top