• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

GPU servers down?

Joined
Jun 25, 2004
Messages
663
7PM eastern

Servers down?

8 GPU clients all waiting for new GPU WU's

That's a real kick in PpD ! :D
 
One of the servers (65.71) has been having some issues off and on all day and the assignment server has been acting stupid about it. However, I just was assigned to 108.21 and downloaded work on four GPUs.
 
Last edited:
Earlier I was having trouble getting work and it had gone through 10 attempts. I restarted the client and it immediately picked up a work unit. After it finished that one, it went to not being able to connect again. I restarted the client again and I'm still not able to get any work.

 
Ah crap, here we go again. :mad:
My gpu1 is twiddling its thumbs, I'm sure once gpu2 is done it will join gpu1 and twiddle its thumbs as well. :(
 
2 down now, a third in about 20 minutes....
 
I wasn't able to get any WU's either. I restarted my machine and it's crunching again.

just a heads up, but it could of been a coincidence.
 
From Vijay regarding this:

Joe has been continuing to pound on this and he thinks he's found and fixed several bugs. I'm optimistic, but the history shows that this may not be all. The WS code on .71 was updated ~5pm pacific time.

I'm sorry for all this mess. Once we get back to normal, we will have a very careful rollout plan for new WS changes, since this should never be happening.
 
Gotta hand it to him for staying in contact..
 
A little more detail from Vijay's blog:

Joe has been pounding on the v5 WS trying to shake it out from the recent disaster with problems returning NVIDIA GPU WUs. The upshot of all of this is that the v5 server code was pushed hard in many ways and several issues have now been found. Joe is testing them, but we're hopefully that beyond the initial good news we had a few days ago, that several additional issues may now be fixed.

It's too early too tell since we're still testing, but I'm optimistic. This only affects particular servers (vsp07b, vspg10a, vsp11a) and the vsp09a CS.
 
mine have been going just fine, last pickup was 10 minutes ago
 
I have a question regarding servers but it's not about the small outage we had today. This afternoon, I had shut down one of my GPU clients in oder to restart a system. I then saved the folder because the client was at 88%. After restarting the system, the client crashed when I executed it. So I went into my backup folder and copied it over thinking it was going to be fine. Well, once it completed, I get a message that the server already received this unit, and from appearances it didn't upload the completed results. So, not only was the WU lost but also the time it worked through the last remaining percentage to complete it once I restored the backup. Is there a proper way to backup our work or is it not possible anymore??
 
I have a question regarding servers but it's not about the small outage we had today. Well, once it completed, I get a message that the server already received this unit, and from appearances it didn't upload the completed results.
This outage, the big outage a week ago, and what you just experienced are actually all related, it's all been one big problem related to multiple bugs in the new server code. I actually received that same message myself on one unit and these messages, judging by reports posted on FF, aren't unique to us.
 
I have a question regarding servers but it's not about the small outage we had today. This afternoon, I had shut down one of my GPU clients in oder to restart a system. I then saved the folder because the client was at 88%. After restarting the system, the client crashed when I executed it. So I went into my backup folder and copied it over thinking it was going to be fine. Well, once it completed, I get a message that the server already received this unit, and from appearances it didn't upload the completed results. So, not only was the WU lost but also the time it worked through the last remaining percentage to complete it once I restored the backup. Is there a proper way to backup our work or is it not possible anymore??
Sometimes when a unit crashes, the client uploads what had already been completed up to that point. That may have been what happened; if the server already received partial results for that unit, it could then have rejected it when you sent it again fully completed.
 
Sometimes when a unit crashes, the client uploads what had already been completed up to that point. That may have been what happened; if the server already received partial results for that unit, it could then have rejected it when you sent it again fully completed.
It's very possible. What can be done to avoid this if we wish to use a backup?
 
I have a question regarding servers but it's not about the small outage we had today. This afternoon, I had shut down one of my GPU clients in oder to restart a system. I then saved the folder because the client was at 88%. After restarting the system, the client crashed when I executed it. So I went into my backup folder and copied it over thinking it was going to be fine. Well, once it completed, I get a message that the server already received this unit, and from appearances it didn't upload the completed results. So, not only was the WU lost but also the time it worked through the last remaining percentage to complete it once I restored the backup. Is there a proper way to backup our work or is it not possible anymore??

With certain errors the client probably assumes that the entire WU is un-usable or that they won't be getting it in a usable fashion, but not the fault of the user. They award partial credit for work done, but they are better off re-issuing the entire unit than to count on a [H]ardass donor resurrecting a WU and sending it back to them.
 
It's very possible. What can be done to avoid this if we wish to use a backup?
You can try restoring the backup right away, but if the unit crashes and it uploads a partial workunit, there's nothing you can do about it.
 
Back
Top