• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Time to kick the server

Oh god don't say that, I have one at 99%.

Edit: I got a WU, no problem
 
It does appear that 171.64.65.106 is in reject mode at this time (nVidia Server)

Stanford was just made aware and attempting to resolve situation at this time.

I am not sure why the AS is not pushing nVidia cleints to the 2 other nVidia servers.

I can see one Alt nVidia server is a little low on work but the other one has plenty.

Issue should be resolved soon.

Xilikon, we might want to mention this issue to the higher ups ..
The AS should attempt to get work from the other two servers but is not!
 
It looks like nVidia 171.64.65.106 is back up and running at this time.

I hope all of you got your cards back up and folding again!!!!
 
Someone should really explain the concept of load balancing to the folks @ Stanford. It'll alleviate their need to constantly babysit the servers (assuming of course that the servers still functioning can handle the increased load)
 
Someone should really explain the concept of load balancing to the folks @ Stanford. It'll alleviate their need to constantly babysit the servers (assuming of course that the servers still functioning can handle the increased load)

There are three nVidia GPU servers and Two ATI servers.

In Theory, they are designed to load balance, but sometimes it does not work.

I know the Pande Group is working & testing some new server code last week so they are aware of that situation.

The main issue is the EUE's. When a core is unstable, and tons of clients keep asking and submitting work units, it slows the server and sometimes crashes it.

Here is an example:

If the server *Normally* gets 500 work units requests per hour, but a new core comes out and now the server is getting 2000 work unit requests per hour (Due to Instant EUE's--> there is no way for it to handle that type of load increase on such short notice.

It is getting better with time, and the newer server side code will help.

 
500 per hour?!

I must be 1 percent of the load of one of their server. I'd hate to think what ROC is. :eek:

Shoot with 300+ GPU for the horde we would give one of thier servers most of its work just on our own.
 
I totally understand having a non constant load on the Stanford servers. I like the explanations by EvilAlchemist about how the server system could be fubared (over stressed) with the load going from say 500 requests an hour (IMO probably low) to a rate of over 2000 requests an hour. It's not like the servers screw up once in a while, but regularly. (like every week end :mad:)

I keep coming back to the old expressions like "if you can't handle the heat get outa' the kitchen", "sometimes your best is not good enough" and the one that seems to apply most to this server situation "excuses are like a$$holes, everyone has one". I'm also reminded of the expression "talk is cheap". ;)

To me it's very sad that Stanford and the F@H project has got in such dire straights.(I'm not lucky enough to be an "old timer" at folding, so maybe this has happened before) Just to name a few examples, the Winsmp client is fubared, the WU's are unstable, the point system is out of wack, they have made too many systems obsolete (like lower end CPU's and I don't want to hear the improvement crap) and their own forum having a bunch of "arrogant dildo's" on it. We're just lucky to have people like brother Xilikon that have good relations with the Stanford "higher up muckity mucks" and is quick to post the releases of new clients, drivers for GPU2, new WUs and any other help in the F@H arena. (I'd hate like hell to depend on the F@H forum):D

A few things that seem abundantly clear to me with all this shite. We may see an improvement and "attitude check" (like the donors are important) in the F@H project and even if it we don't, we still have the BOINC WCG program to fall back on. I have two quads doin' both the Stanford F@H GPU2 and the USC BOINC WCG program (I'm just waiting on some gear for a third boxen). Until Stanford et al gets their "head outa' their a$$"" and comes up with some solutions for the CPU clients and WUs there is no way in hell I'd ever fold with a CPU.

Folding and WCGing for the CURE

 
Considering how much money is being spent by individuals, and how important this research is, it IS very surprising that the Pande group doesn't have powerful enough servers to handle the WU load. I mean, I bet thousands of the people folding for them have more than enough computing power to handle the load, even when clients start EUEing.

It's a very odd imbalance.




 
First off, I just made up the (500) number as an example.
I am sure it is much more then that ....

There are so many nVidia clients asking the server for work at one time.
Think about it, your GPU needs work every couple hours.

So 1 GPU x 10 Work units per day ---> now multiply that by how my nVidia GPU in the world ... it is a freaking huge number

I think the Pande Group gets good funding, but with a project this large it costs ever more to keep it running and buy new systems ..

I bet thousands of the people folding for them have more than enough computing power to handle the load, even when clients start EUEing.

I think you are under-estimating how much CPU & network power it actually takes.

Go read up on the [H]ardOCP & [H]ardForum servers and what load they are on ...

They have almost 70 servers for Folding@Home .. so don't say they don't buy hardware.

jws2346r said:
I keep coming back to the old expressions like "if you can't handle the heat get outa' the kitchen", "sometimes your best is not good enough"

Sounds good to me .. I will ask Dr Pande to shut down the nVidia side of things so ATI can have more work and servers .. good suggestion.
 
Hey, I almost fell out of the chair laughing when I read that statement "Sounds good to me .. I will ask Dr Pande to shut down the nVidia side of things so ATI can have more work and servers, good suggestion " I like that suggestion, although I feel it's a little on the sarcastic side. :rolleyes:

It would really be halarious if it wasn't so serious. If ole' doc V ain't careful and he doesn't "keep his eye on the ball" concerning F@H's clients, drivers, programs, sarcastic reps statements, etc he may find himself not only shutting down the nVIDIA side, but the whole freaking F@H side altogether.The way things stand at the moment a lot of people have switched to WCG for CPU DC'ing because it's much more stable than F@H. I haven't tried the GPUGRID program yet, but it may be a good program. :confused:

To get one thing clear. I'm going to use my computers to help medical science cure some very fu*cked up diseases. Now, which program I use is dependant on which program is the most stable, which programs administration is more solicitous to it's donors (in other words which program gives some respect for the investments and time expended in donations) and which program doesn't do the "bait and switch" deal. (like GPU1 to GPU2, regular clients no pernts, etc) Now I understand progress in IT, but when a lot of people expend a lot of time and money for naught I find it a litthe concerning. :(

Folding and DCing for a CURE

 
I like that suggestion, although I feel it's a little on the sarcastic side. :rolleyes:

LOL .. it was ......

I have been on the defensive latley because it seems F@H has become the new whipping boy to some ...

No matter what is done .. it is not good enough. Some of those posts, I just ignore.
Other I chime in on but it really just depends.

Being on the Beta side of things, I have learned a great deal and I have come to truly understand most of the issues and the sides people tend to take.

No matter what choice Stanford makes, one group is happy - one is mad.
Problem is, the mad ones hoop and hollar .. the happy ones just keep folding.

It is a never ending cycle. I can tell you things look better from what I have seen.
There are somany people working everyday to fix issues and make it a little better.

A year ago, linux would crash (and people yell) but now that is mostly fixed, now the GPU crash (and people yell) but it will get fixed .... the cycle goes on ...

Some people quit F@H, we gain new members ... but we will always be [H]
 
To continue from my last post ... let me make up an example for the group to ponder on.

I must stress ... these are fake numbers - but the situations are real

Lets just take the GPU2 side of things in this post ....
---------------------------------------------------------------------------------------------
We all know that the GPU2 work units are benchmarked on a ATI 3850.

Project : 0001 - Core Version 1
Nvidia 9800 = 5000 PPD on this project
ATI 4850 = 3000 PPD on this project

Now, lets say the ATI core was improved by 20% --> so they re-benchmark the Work Unit.

Project : 0001 - Core Version 2
Nvidia 9800 = 4000 PPD on this project
ATI 4850 = 3000 PPD on this project

Which Core 2, ATI clients can run the unit faster but since the benchmark system is ATI the PPD remains the same on their side, but causes a drop on the nVidia side.

What do you do when nVidia users get mad?

If you give a bonus to nVidia clients to bring the PPD back to the same level, ATI users get upset because they can't catch nVidia.

If you don't give a bonus and keep the lower PPD , nVidia users get mad at the PPD drop

---------------------------------------------------------------------------------------------
Now that is just one situation. Now add in the Quad Uni-Processors Clients vs SMP PPD difference . the SMP vs GPU PPD difference - Uni-Processors - GPU PPD difference ..
The ATI vs nVidia PPD difference ...

The list just goes on and on .. and are some extremely complex issues....
 
I have been on the defensive latley because it seems F@H has become the new whipping boy to some ...

It always has been the whipping boy, and for good reason. WCG just works. UD just worked. Something breaks in F@H with amazing regularity. And its not just all these issues that keep coming back, its Stanford's attitude about it. Instead of screaming STFU and GTFO, a little appreciation towards those of us who buy the hardware and pay the electric bill would go a long way.

After the PS3 situation back in Feburary, the repeated GPU2 and SMP issues since, I've earned the right to bitch. It's the only thing I get in return for the 3+ years I've been doing this project.
 
Back
Top