• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Dumb SMP vs GPU tradeoff question

jebo_4jc

[H]ard|DCer of the Month - April 2011
2FA
Joined
Apr 8, 2005
Messages
14,572
So if you have a dual core CPU and a X1900XT, should you run the GPU client + SMP? Or should you go with one normal client + GPU? Or SMP only?

Is there a list somewhere of what hardware gets what in terms of PPD?
 
If you are after points alone then the SMP client atm is pretty much guaranteed to get you the most PPD. The GPU client seems to return more work faster (based on what I read but then again I am from the south so that there readin shore is hard) but isn't quite as many PPD. Running the GPU client and one standard client (assuming you have dual core) seems to give you about 75% PPD of the SMP client
 
If you're after more points-per-day then go with SMP no question about it at the moment. (things always change though keep in mind both of these clients are beta)... If you want to do more science quickly then go with 1 GPU + 1 standard client (assuming dual core). GPU's are much more powerful than even quad core cpus in terms of their floating point math power, but stanford's points system is a little bias against GPUs atm. Either way though, you'll get good PPD and all types of contributions are appreciated.;)
 
I really wish Stanford would boost the point reward for GPU WUs. I'm nearly ready to deploy my dual GPU folding box (need another GPU and a hard drive), but considering the power consumption and possibly problems with such a machine, I've hesitated. Especially since the Windows SMP clients have hit.

Considering that the scientific value of the GPU work units surely must be higher than for the regular and SMP CPU clients, I wonder why they're still sitting at 330 points each. I'm still going to deploy my GPU box, eventually, but I'm going to put the money that would have gone into it next month into a 64bit Linux box I'm building for it's SMP client. Instead of a dual core, I'll be buying a Q6600.
 
I contacted Mr. Pande himself about the points situation with GPU vs. SMP.. he basically said not to worry, the GPU client is being updated shortly.. he said they're back-porting the PS3 code to run on the GPU which will make the GPU more flexible in terms of the types of simulations it can run.. and at that time they'll be making a big push for the GPU client.. He also noted that the SMP client has such high points value because the particular experiments being run at this moment require SMP..

GPUs do have more FLOPS (and once made more flexible, more potential) for computational science than SMP, but the points will come in time.
 
If you're after more points-per-day then go with SMP no question about it at the moment. (things always change though keep in mind both of these clients are beta)... If you want to do more science quickly then go with 1 GPU + 1 standard client (assuming dual core). GPU's are much more powerful than even quad core cpus in terms of their floating point math power, but stanford's points system is a little bias against GPUs atm. Either way though, you'll get good PPD and all types of contributions are appreciated.;)

Can't you run a combination?!

Like use a quad core chip running the SMP on 3 cores, then use a 4th core to run the app for the GPU client ? I'm completely "new" to this (havent run F@H in a long time), but all this talk of PS3 & GPU clients have gotten me interested in running F@H again.
 
Is there a list somewhere of what hardware gets what in terms of PPD?

I second this question.

If such a list isn't available, having seen that the PS3 is churning out 900 PPD, what kind of PC system(s) is/are necessary to put out that kind of production? And what kind of production are the various X1900 series GPUs capable of?
 
Can't you run a combination?!

Like use a quad core chip running the SMP on 3 cores, then use a 4th core to run the app for the GPU client ? I'm completely "new" to this (havent run F@H in a long time), but all this talk of PS3 & GPU clients have gotten me interested in running F@H again.
Well you can't run three SMP threads. AFAIK right now all the SMP clients spawn 4 processes.

Please forgive me as I confused and interchange process and thread all the time. I understand there is a difference, but in the context right now I don't think it matters :(

If you have a quad core chip -- I would run the SMP client.
If you have a GPU -- I would run the GPU client.
If you have a GPU + quad -- I would run BOTH! :)

From what I understand -- the CPU client hogs up 1 core. Well, there are people running the SMP client on dual core machines, so obviously a three core machine would work just as fine as long as the processes can end someone concurrently (if I understand how the things operate, the threads need to catch up).

:)
 
I contacted Mr. Pande himself about the points situation with GPU vs. SMP.. he basically said not to worry, the GPU client is being updated shortly.. he said they're back-porting the PS3 code to run on the GPU which will make the GPU more flexible in terms of the types of simulations it can run.. and at that time they'll be making a big push for the GPU client.. He also noted that the SMP client has such high points value because the particular experiments being run at this moment require SMP..

GPUs do have more FLOPS (and once made more flexible, more potential) for computational science than SMP, but the points will come in time.

Very interesting.

The deal is currently, as confusing as it is, the GPU client needs a core of anything to feed it for max output (~ 660 PPD +/- some depending on the card). The GPU does not care if it gets 100% of a P3 @ 1.0 GHz core, or 100% of a C2D core @ 3.6 GHz. The GPU client scales way back if you cut back on the amount of the core you feed it with-- makes no difference (~ 5% max hit) on what actual core you feed it with. Lower the P3 by ~ 25% total feeding it the PPD hit will be same as lowering the C2D core by ~25%.

Makes no sense does it? Need to think on GPU in terms on cores feeding it, not % of a given core. Second you run the SMP client it's competing for CPU ticks which slows done the GPU a bunch. Generally the GPU won't make up the PPD loss from the SMP side of things-- or it's a wash in PPD terms. Science terms I'd assume it's doing way more since well it is doing way more.

Confused? SMP + GPU generally blow short hand version of things. This is all based on current client offerings/point values.

edit: on a quad you can't limit (currently) the SMP to only spawn 3 threads and use the 4th to feed GPU. You end up with the GPU thread and SMP thread fighting on the same core = GPU being fead by 1/2 core.
 
Can you set affinity to have the SMP client using 3 cores and leave one open for GPU?>
 
I might try that on my dual-dual box.
At the moment each core gives me around 500 PpD.
My vid cards are pulling close to 800 PpD.
So if dropping a core from the SMP and giving it to the GPu works then I'llbe up by 300 PpD.
The down side is that box will probably pull around another 80 watts out of the wall.

Luck ......... :D
 
Under nix, odds are. Under win, no clue. No GPU under nix so....

I know you can set affinity IN windows I just dont know if you can set affinity for F@H specifically. Someone send me a quad core and I will let you all know :) Oh yeh I need an ATI video card too so that we can fully test the theory


 
Tried it on my C2D @ 2.75 Ghz ....... :D
Before just running the SMP client on both cores.

[11:01:27] Project: 2610 (Run 0, Clone 175, Gen 1)
snip
[11:01:44] Writing local files
[11:01:45] Completed 0 out of 500000 steps (0 percent)
[11:17:44] Writing local files
[11:17:44] Completed 5000 out of 500000 steps (1 percent)
[11:33:43] Writing local files
[11:33:43] Completed 10000 out of 500000 steps (2 percent)
[11:49:41] Writing local files
[11:49:41] Completed 15000 out of 500000 steps (3 percent)

So looking at 16 mins per frame for ~1570 PpD.

After manualy locking all 4 theads to core 1 and starting a GPU client and locking that to core 0.
I get .................

[19:12:37] Writing local files
[19:12:37] Completed 150000 out of 500000 steps (30 percent)
[19:34:36] Writing local files
[19:34:36] Completed 155000 out of 500000 steps (31 percent)
[19:56:44] Writing local files
[19:56:44] Completed 160000 out of 500000 steps (32 percent)
&
[18:39:29] Completed 52
[18:45:50] Completed 53
[18:52:12] Completed 54
[18:58:30] Completed 55

So looking at ~22:05 per frame for the SMP client & ~6:20 per frame for the GPU one.

Looks like the SMP client runs ~2/3 speed and the GPU at full speed.
So for this combo of CPU & protien it works and I'm now looking at ~1900 PpD.

Stanford wont like it but I'm now wondering with a fast C2D CPU could you run 2 SMP clients, each locked to its own core and get then back before the deadlines.
If my C2D could do it I'd be looking ~2200 PpD off this one box .......... :eek:

Luck ............ :D
 
So looking at ~22:05 per frame for the SMP client & ~6:20 per frame for the GPU one.

Looks like the SMP client runs ~2/3 speed and the GPU at full speed.
So for this combo of CPU & protien it works and I'm now looking at ~1900 PpD.

That's quite interesting. Perhaps there's a memory bandwidth issue or something with the SMP client? Two clients with all their threads locked to one core each will probably suck - anyone wanna try it? :D
 
Stanford wont like it but I'm now wondering with a fast C2D CPU could you run 2 SMP clients, each locked to its own core and get then back before the deadlines.
If my C2D could do it I'd be looking ~2200 PpD off this one box .......... :eek:

Interesting numbers, not what I'd have expected. My playing has been w/ not locking the threads but with the C2D I'd expect a larger hit giving up a C2D core.

Actually... now this may be odd since your the last person at the [H] I'd question their numbers but this can't be right. The SMP never meets deadlines on single core boxes.... then again no one ever tried on a fast C2D core...

How do you lock threads to a core?

zim01 has played a ton with this stuff in nix land, at the end of the dat running 2x clients on a 8 way box ends up better PPD than 4 or more (3x clients corrupt each others threads I guess).
 
Interesting, I wonder if my E6600 @ 3.0 could run the SMP fast enough to make the deadlines and push both X1950XTXs that are in the box. I've pretty much given up on GPU folding because of it needing a dedicated core and I could never get both cards folding at the same time, but if I can do both, it might be worth revisiting.
 
Interesting, I wonder if my E6600 @ 3.0 could run the SMP fast enough to make the deadlines and push both X1950XTXs that are in the box. I've pretty much given up on GPU folding because of it needing a dedicated core and I could never get both cards folding at the same time, but if I can do both, it might be worth revisiting.

dual GPU results are all over the map, no first hand experience here. About the only constant is you take a sizable hit with dual GPU's. Of note, the CPU usage is not 2x more like ~ 135% or something goofy as heck with dual GPU (instead of needing 200% for example).

Try it... I think you might need a display on each card for context crap, I've not followed dual GPU much. To my knowledge no one has done quad GPU in a box, Macholic pussed out on it!
 
Times have settled down a bit.
Its now running 24:30 & 6:16.
But still a gain.

I opened Windozes task manager and locked the treads to a core in there.
The down side is that every time you start a new protien, you will need to re-lock the client to its core.

Also the numbers are a little off.
Was useing 1760 points not 1523 points which p2610 is worth.
So the PpD has droped to only ~1640.

Luck ............ :D
 
Times have settled down a bit.
Its now running 24:30 & 6:16.
But still a gain.

I opened Windozes task manager and locked the treads to a core in there.
The down side is that every time you start a new protien, you will need to re-lock the client to its core.

Also the numbers are a little off.
Was useing 1760 points not 1523 points which p2610 is worth.
So the PpD has droped to only ~1640.

Luck ............ :D

Is this still a gain in points running both?

And I'd imagine that box is pretty much useless for doing anything else with both clients open, yes?
 
Is this still a gain in points running both?


This is C2D @2.76 Ghz with 2 gig of Memory & dual X1950xtx's in.
Dual GPU clients was giving me ~1350 PpD pulling ~400 watts.
Single SMP client was giving me around the same with this protien but only pulling ~290 watts.
So running both clients like this is giving me ~1640 PpP pulling ~350 watts.
So its a gain of around 300 PpD.

And I'd imagine that box is pretty much useless for doing anything else with both clients open, yes?

No.
If anything its more responsive with these two clients than it is with two plain GPU ones.
Windozes task manager shows the GPU core is loaded ~95% CPU & ~90% kernel and the SMP core is loaded 100% CPU & ~10% kernel.

Luck .............. :D

 
I tried this on a P4D/x1950 Pro. The XTX obviously outproduced the Pro so +200 PPD.

Still nuts your can pull that much PPD off one C2D core/SMP not scaling on 2x cores for poo.
 
Back
Top