• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

New GPU client - Open Beta

I was actually wonder if it has to do with the pre handling and post handling and the P4 not being able to load the gpu fast enough because other machines with AGP are getting over 1200PPD

Yeah, we know the P4 is a POS compared to any A64 CPU on nForce, but it's up to Stanford to work around this and max out the GPU regardless. The HD 3850 should be doing better than just 100 more PPD than your 1950! That's just silly. It's a way better card on a P4:

http://www.hardforum.com/showpost.php?p=1032113072&postcount=63

Although I’ve moved on to PCIe, I still maintain and run a 3.2 Northwood on an Intel D865 PERLL motherboard.

In my own testing, the 3850 outscores the 1950 by 88% in 3DMark03, and by 26% in 3DMark06. So for now, it would appear that the 3850 is a worthy successor to the 1950 and in reality offers more ‘bang for the buck’.

http://www.theinquirer.net/gb/inquirer/news/2008/02/14/agp-nth-coming

The HD 3850 AGP gives your midget Pentium something to stand on in a world of Core2/Athlon/Phenom giants. If you are one of the latter, do not mock the 3850 on AGP, it scores 900% better than our old Geforce FX5900 and almost 800% better than our Radeon 9700Pro. You can rest assured your AGP rig has a year's worth of gaming left in it.

They'll fix this. They have to if they want all those thousands of worthy P4s to fold with HD 3850s.


 
I haven't given up on it yet today I'm going to max my OC on the P4 it can hit 3.9Ghz and I am going to try tightening the mem timings and do anything I can to eek out a bit of performance!!!

 
This gave me a nice idea for my HTPC box project. I will have a HD3870 card in it for media streaming, gaming on the tv and folding with a Q6600. Give me about 1 month to setup something like that.

It's open beta so it's still room for improvement since we now see it's not being used to the maximum.

 
With HT off I still see 43% usage which means I'm going to turn it back on and try to run 2 clients shortly

Results with 2 GPU clients on one card with HT running? Does it max out the GPU?

Well, as many people noticed, on most configs, running F@H-GPU2 on a single 38xx and dualcore A64 X2 or Core2 results in about 1300 - 1500 ppd and 65 to 75% GPU load. HD3650 is at 100% though.

I tried to run 2nd GPU client on another GPU (1st is 3850 at 750 MHz, 2nd is 3650 at 920 MHz), but the 2nd didn't start saying it supports only 2xxx and 3xxx (and what else do I have?!). I tried to remove -gpu 1 from the command prompt, it runs, but... also on the 1st GPU (3850).

Well, it isn't what I wanted. But... They load my 3850 at 100% and get an average of 950 ppd each (1900 ppd for two)... While running SMP on single A64 X2 core doesn't do any good (not enough horsepower to get there before deadline...) and old CPU client is nothing like point-effective, dual F@H-GPU on a single GPU are the best solution - until further client improvements, at least.

Why not? Is there any science problems? No way... deadlines are very far away, as 3850 isn't 2400pro and it does about 20 WU's per day with 2 "threads". I suppose, to utilize GPU completely seems to be more important. Also, you can see clearly, what hardware to upgrade first, and if buing some 3xxx worth it with your current CPU. :)



 
Excellent indeed, this really make sense...

I will buy a HD3870 card for my HTPC box soon. I like to vary the production a bit and if I can manage that without losing the total PPD of the box, it's better.

 
Excellent indeed, this really make sense...

I will buy a HD3870 card for my HTPC box soon. I like to vary the production a bit and if I can manage that without losing the total PPD of the box, it's better.

Yeah, if I can find a good excuse to run GPU2 on a slow P4, I'll be contributing to just about every kind of project (Windows Uniprocessor, Windows SMP, Linux SMP, PS3, GPU). That's important to me.



 
for everyone's info (particularly metallicafan's)

i have:
a64 x2 5000+ black oc'd to 3GHz (240x12.5)
2gb DDR2-857, cas4
hd 3870, stock clocks

and my gpu utilization according to ATI catalyst hovers around 70%
 
Lol, this is coming from someone who bash on those who like to tweak for the most PPD. My comment is also made with the power efficiency in mind.

Keep your opinion to yourself and let us decide with numbers, not ramblings.



Ramblings, huh? Opinions, huh? Try facts. I guess you're wrong again, twice. Maybe I do know what I am talking about...

FAH NEWS: How the GPU2 core works

Vijay said:
There have been some misunderstandings on how the GPU2 core works. In particular, for small proteins like villin on GPU's with large number of stream processors (SP's) like the 3850 or 3870, the protein is too small to use a larger number of SP's unless the CPU is very fast.

...In parallel, Mike Houston at AMD is working to optimize CAL such that it has lower CPU overhead.


...
97 point WUs are very small. Small chunks of data are proabably inefficient to parallize between 320 Stream Processors. They likely won't saturate a high end GPU, so you won't know the full performance of the GPU2 client until we get larger WUs.

And as always, this is an EARLY look at an early beta client. There is a lot of room for tweaking and performance improvements...


For the second point, I DON'T bash anyone. I don't bash people for optimizing for points. I DO, however, go to great lengths to dispell the myth that higher PPD is automatically more helpful to the project. Running 2 clients more slowing is not has helpful to the project as running 1 client much faster. Vijay has said so himself. However, I have absolutely NO problems with people optimizing for more points. Just don't claim it's better for the project because it is not. It is simply better for the points, but still GOOD for the project (just slightly less good). All donations are welcomed, even Xilikon's. ;)
 
A data point, moving from 3.6Ghz to 3.9Ghz on my P4 netted me about 776 vs 722 so a 54pt gain for 300mhz. I think as the cpu overhead is reduced this 3850 with 320 stream processors is going to pull 1200 pts or more. Making it a lot more attractive and being able to allow us to breath life back into those old AGP systems.
 
From the horses mouth as to why the current work units aren't maxing out the GPU

http://folding.typepad.com/news/2008/04/how-the-gpu2-co.html

This is very encouraging news indeed. I'm glad there is a reasonable explaination and I look forward to larger WU's!

On a good note I have been running the GPU client for a few days now (not 24/7) and have completed several WU's successfully with no problems or issues. Seems to be fairly stable in my experiences. Sunin and master381 - Thanks for the info.:)
 
Yeah, if I can find a good excuse to run GPU2 on a slow P4, I'll be contributing to just about every kind of project (Windows Uniprocessor, Windows SMP, Linux SMP, PS3, GPU). That's important to me.
For me as well. Besides GPU2, the only client I haven't tried yet is the PS3 because I don't have one. My next video card will likely be a new gen ATI card, and will hopefully fold on that eventually.

BTW, anyone know why some stated that the HD2900XT cards might have higher potential than the newer 3000-series cards? I thought they were essentially the same architecture but on a different die process.
 
I know that some post i saw, pointed out that fact that the 2900's have a 512bit bus and the 3800s only have a 256bit bus. What effect this could have on performance I dont think really anyone knows yet, obviously.
 
I know that some post i saw, pointed out that fact that the 2900's have a 512bit bus and the 3800s only have a 256bit bus. What effect this could have on performance I dont think really anyone knows yet, obviously.

Probably about as much difference as it makes in games: just about zippo.


 
On-board memory performance has little impact on folding speed until you get way high or way low. Note that memory is generally less tolerant to overclocking than the core, and that overclocking that is "safe" for graphics may cause issues with Folding.
 
Probably about as much difference as it makes in games: just about zippo.
It makes little difference now because the standard amount of memory on most new cards is 512MB or less. Many people have speculated that in the near future when 1GB cards become the norm, a 512-bit bus will be required for optimal bandwidth. Some companies are already manufacturing such cards but crippled them because they're only 256-bit wide. I have no idea how this may affect F@H, though, but I can't imagine it will hurt with the much larger WUs slated for release next week.
 
Folding isn't affected by memory speed very much. Overclocking the RAM does little.


My 3850 gets ~1650PPD running on a 3GHz Q6600. GPU utilization is about 92%. Interestingly enough, the load on the CPU is spread evenly across the four cores. I'm just glad that Vista64 isn't messing anything up.
 
Okay, I have a Sapphire 3850 AGP in hand. I'm running the drivers from the CD with the 8.3 hotfix drivers installed over the top (yes the driver install should be less of a pain...). With an Athlon 3200+ (@2.2GHz stock) running on a MSI K8N Neo (Nforce 3 stock settings) and with the 3850 AGP running at stock (stock: 669c/829m), I'm seeing 80-85sec per frame and GPU utilization of 65%. We are working on updates to the Brook and CAL code to reduce CPU load which should raise the GPU utilization.

I haven't yet been able to reproduce Sunin's low numbers. Tomorrow morning I can run on an old Dell box with an 865 chipset (it's a G instead of a PE like Sunin has) to see if it's a chipset issue. But, that old Athlon is pretty slow and old but is still doing okay. So this is looking like a chipset or driver issue.

Sunin: are you running the default windows chipset drivers? Do you have Microsoft Defender or an anti-virus package running? What video drivers exactly are you running? I'm using the 8.471 drivers from the 8.3 hotfix Sapphire has posted with the 2D driver version being 6.14.10.6783. I first installed the Sapphire 8.452 drivers (the whole thing) and then did a "update driver" and pointed it at the 8.3 hotfix drivers.
 
On-board memory performance has little impact on folding speed until you get way high or way low. Note that memory is generally less tolerant to overclocking than the core, and that overclocking that is "safe" for graphics may cause issues with Folding.

The 2900XT has a theoretical peak of 105.6 GB/s, the 3870 has a theoretical peak of 72 GB/s and the 3850 has a theorectical peak of 53.1 GB/s. Will this difference in memory bandwidth have no affect on the F@H client?
 
You can test this for yourself. When not CPU bound, the client is bound by GPU core clock much more so than memory speed. This may change over time as more develop happens and optimizations change the bottlenecks.
 
No microsoft defender ro anti-virus. IJ am running the default because asus has no chipset driver for the board. I checked their website. I have used both 8.1 and 8.3 with the hotfix.

The real issues I think is that the P4 is much slower than an athlon. When I go from 3.6 to 3.9 I get nearly 10% increase in PPD its sitting at 784 atm from 722... I think your cpu load adjustments along with more complicated gpu intensive WU's might result in some increases. We will see. I can try to find the 8.471 drivers and do the hot fix just like you if you wish for me to.

 
I'm seeing 80-85sec per frame and GPU utilization of 65%. We are working on updates to the Brook and CAL code to reduce CPU load which should raise the GPU utilization.

For folks that don't have the client, what does this translate to in PPD?




 
for everyone's info (particularly metallicafan's)

i have:
a64 x2 5000+ black oc'd to 3GHz (240x12.5)
2gb DDR2-857, cas4
hd 3870, stock clocks

and my gpu utilization according to ATI catalyst hovers around 70%

More fun stats, fyi:

I oc'd my 3870 10% core and mem, 855/1241
and my gpu utilization dropped 10%, to about 60%.
 
More fun stats, fyi:

I oc'd my 3870 10% core and mem, 855/1241
and my gpu utilization dropped 10%, to about 60%.

I assume PPD stayed level, this really does show the CPU is holding back the GPU... /crosses fingers for improvements! If they do happen this can be HUGE for folding

 
ok after running the GPU2 client for 5 days nonstop and playing with card clocks on my 3870...

1. GPU utilization plays the biggest impact on folding speed (kinda a no brainer)

2. Core clock plays the next most important part. However, going from 750MHz @ 90% utilization folds almost the same speed as 875MHz @ 70%. Only gain about 1-2 seconds per frame.

3. CPU speed only affects utilization %. However this plays the biggest impact on overall performance.

4. Card memory speed plays almost no part. With the core at 850MHz, ram at 1000MHz (down-clocked), I only gained 1 second per frame going to ram @ 1325MHz.

On average I am getting 1746 PpD with 17-18 WUs turned in per day. This is with a Q6600 @ 3.42GHz running WinXP Pro. The GPU2 client is super stable for an early beta... way more so then the SMP client. You can actually close it without losing work. However hitting the Reset button did cause me to lose a WU.. once out of 5 times. Can't say the same for the SMP client, as it lost the WU 4 times out of the same 5 times. With 2 SMP clients I was averaging 4200PpD with Affinity Changer. With one SMP and GPU2 I'm averaging 4500PpD without Affinity Changer.

 
[H]ugh_Freak;1032372778 said:
ok after running the GPU2 client for 5 days nonstop and playing with card clocks on my 3870...

1. GPU utilization plays the biggest impact on folding speed (kinda a no brainer)

2. Core clock plays the next most important part. However, going from 750MHz @ 90% utilization folds almost the same speed as 875MHz @ 70%. Only gain about 1-2 seconds per frame.

3. CPU speed only affects utilization %. However this plays the biggest impact on overall performance.

4. Card memory speed plays almost no part. With the core at 850MHz, ram at 1000MHz (down-clocked), I only gained 1 second per frame going to ram @ 1325MHz.
This client seems to behave very differently from the first one, unless it's the architecture difference of the video subsystems that are playing a major role. With the original client, the biggest increase I found is with increasing card memory speed, then comes GPU clocks and last is CPU frequency. This is on an AGP card in XP.

On average I am getting 1746 PpD with 17-18 WUs turned in per day. This is with a Q6600 @ 3.42GHz running WinXP Pro. The GPU2 client is super stable for an early beta... way more so then the SMP client. You can actually close it without losing work. However hitting the Reset button did cause me to lose a WU.. once out of 5 times. Can't say the same for the SMP client, as it lost the WU 4 times out of the same 5 times. With 2 SMP clients I was averaging 4200PpD with Affinity Changer. With one SMP and GPU2 I'm averaging 4500PpD without Affinity Changer.
I'm sure your power consumption went way up though.
 
This client seems to behave very differently from the first one, unless it's the architecture difference of the video subsystems that are playing a major role. With the original client, the biggest increase I found is with increasing card memory speed, then comes GPU clocks and last is CPU frequency. This is on an AGP card in XP.

I'm sure your power consumption went way up though.

I'm sure GPU clocks and memory speed will play a bigger role once they get the CPU usage better. As for the power consumption.. yeah it probably did go up but my headaches with having to watch over the buggy as hell SMP client went waaaaaaayy down.

 
I'm sure your power consumption went way up though.

I see no difference in power consumption between the 1950 and the 3850. Not sure why I expected to increase to about 250, but nope. still around 210ish.
 
I see no difference in power consumption between the 1950 and the 3850. Not sure why I expected to increase to about 250, but nope. still around 210ish.
Actually, I was referring to the increase from running two SMP clients vs 1 SMP+1 GPU2 client.
 
Actually, I was referring to the increase from running two SMP clients vs 1 SMP+1 GPU2 client.

Ahh yes that should be a jump. A jump that might not net adequate P/watt
 
How did that affect your PPD?

Not sure... I'm new to folding actually, I just wanted to put my 3870 to use for something. I used to do a lotta seti@home, but this is my first experience doing f@h.

So, I dunno how to measure PPD... will that Fahmon program I've heard about tell me that, or do I have to calculate it myself?
 
Okay, swapped the 3850 AGP into an old Dell box with the same chipset as Sunin. I can match Sunin's poor performance. The Intel 865 chipset/processor seems to not be doing well:

CPU: Pentium 4 3GHz
Motherboard: Dell GX270 (Intel 865G chipset)
GPU: Sapphire 3850 AGP (669c/829m stock)
Driver: 8.3 hotfix driver (package 8.471.1.1-080324b-061196E-ATI), 2D Driver 6.14.10.6783
OS: Windows XP 32-bit SP2 (fully patched)

GPU utilization: 42% (CCC and Rivatuner)
130 seconds per %

Hrmm. Either the P4 is running the code way slower than the Athlon, or this is a chipset issue. Sunin reports that overclocking the CPU seems to help, but a 2.2GHz Athlon is able to keep up, so I would expect a 3GHz P4 to be able to do the same...

Anyone else have a AGP system with an Intel processor and a Nvidia chipset or different Intel chipset?
 
I wonder if the card being AGP has anything to do with it. Do you have hyperthreading turned on?
 
The same board in two different configs has different performance. Nforce3+Athlon 3200+ seems okay (not as fast as PCIe equivalent), Intel 865+Pentium4 3GHz is much slower. Same performance (within a few percent) on the Intel box with HT on and off.
 
Yeah I've done alot of testing with this and nothign seems to impact performance. I've turned on and off fast writes, I've decreased and increased aperature size from 64 to 256mb no changes. The only thing I did not try it to OC the AGP bus. I remember back in the day fears about corrupting the HD because that bus speed also somehow tied into your sata controller.

I tried the following:

OC CPU from 3.6 to 3.9 had some impact
OC'n the 3850agp no impact just dropped utilization
Different Drivers 8.1, 8.3 with hotfix
HT on and off
Dual GPU Clients boost gpu % by about 10-12%
Tighten memory timings
Overclock memory (adjust ratio)
Aperature changes
Fast Writes

Man I know there were a few more in there... I spent probalby a full 15hrs tweaking or more with little enhancement.

I'm just happy the lack of performance could be recreated and its not just me :)
 
I've been poking at this for awhile as well with no real performance improvement. That's why I'm hoping that someone has another Intel based machine with a Nvidia/ATI/or different series Intel chipset to test with so we can narrow things down.

I'll keep poking at this, but most of the knobs have been covered and being on a Dell box means I'm limited in tweaks and overclocking. For the current WUs, a 2600/3600 series GPU seems to be the best bet. As proteins get larger, the 3850 AGP may look better on this system. On AMD+Nforce3, things look pretty good with the 3850.
 
Back
Top