SR-2 & L5640

Uncore sets the frequency of the on-die memory controller and the L3 cache. That is all, QPI frequency is different.
 
Ah crap, wait...yeah, I got confused there for a minute.
Yep...QPI frequency and uncore frequency...now I remember seeing both in my GB's. D'oh! Sorry guys :(
 
Retail and B0 ES chips have uncore locked, and are not changeable on the SR2, only the B1 ES and A0 ES have working unlocked uncore on the SR2

Yes, and before anyone brings it up, I know the BIOS lets you change it with the retail chips, but it does not work, as verified by me and other users.
 
Uncles settings are rumored to be in the future bios'es (what's plural bios lol, I'm drunk)
Posted via [H] Mobile Device
 
Uncles settings are rumored to be in the future bios'es (what's plural bios lol, I'm drunk)
Posted via [H] Mobile Device
lol its "bios" not bioseseseses.... :p yeah sadly im finally sober.. :(

oh and its uncore not uncles :p (waits for the "i swear i typed uncore and the phone changed it to uncles" comment)
 
I swear I typed uncles and the phone changed it to uncles comment
Posted via [H] Mobile Device
 
Damn it iPhone!!! Can't even make a funny...

UNCORE.... there we go, stupid iPhone
Posted via [H] Mobile Device
 
So what's looking like a peak OC and a peak ppd for these chips on an SR2? if I'm 205 stable on my 5530's could I expect the 5640's to match that?
Posted via [H] Mobile Device
 
So what's looking like a peak OC and a peak ppd for these chips on an SR2? if I'm 205 stable on my 5530's could I expect the 5640's to match that?
Posted via [H] Mobile Device

From my experience, no. I was 208 bclk stable with Gainstowns, but I am lucky to get 200 bclk with these chips. I would say you can count on 3420 (190 x 18), and anything over that is a bonus. I really hope with the number of these coming on-line that someone figures out the secret. If the single processor boards can push these chips well over 200 bclk, the SR-2 should as well. From my experience so far, that has not been the case.
 
I think I've got 190 stable, that said, I might be able to get 195 with a bump in voltage (I'm not at 1.4 yet) I'm not sure that little bit is worth it.

I'm seeing ~100k ppd @ 190 blkc and 430 watt pull. I'm happy with that.
 
I think I've got 190 stable, that said, I might be able to get 195 with a bump in voltage (I'm not at 1.4 yet) I'm not sure that little bit is worth it.

I'm seeing ~100k ppd @ 190 blkc and 430 watt pull. I'm happy with that.
about 100kppd at 190x18 = 3420mhz??

Hmm I'm gaining interest in sr-2. Looks like it's not that hard to overclock to decent clocks, and it's producing more points than I thought.

Btw you guys tried running smp 23/ smp 22 instead of 24? There was reports of better performance using less threads on 4P servers on foldingforum.
 
I'm interested in the math or logic behind that... I've heard the sme thing, even a guy with 64 cores was getting better times from like 53 or something like that

What's the reasoning?
 
I think I've got 190 stable, that said, I might be able to get 195 with a bump in voltage (I'm not at 1.4 yet) I'm not sure that little bit is worth it.

I'm seeing ~100k ppd @ 190 blkc and 430 watt pull. I'm happy with that.
 
I'm interested in the math or logic behind that... I've heard the sme thing, even a guy with 64 cores was getting better times from like 53 or something like that

What's the reasoning?

I think there's a limit to the parallelism of most programs. The SMP client seems very good, but apparently it doesn't scale well to 64 threads. 24 threads are no problem. The only reason to run with 23 is if you are also running GPU clients (or something else that constantly needs processor time).
 
I'm interested in the math or logic behind that... I've heard the sme thing, even a guy with 64 cores was getting better times from like 53 or something like that

What's the reasoning?
I don't understand really why, but here's some quotes:
Punchy wrote:lizard, if you still have the system, I'd suggest running with a lower SMP number than 64. With a set of slower CPUs and an asymmetric memory population, I went from 19:00 at SMP 64 down to 12:06 at SMP 52. 52 works out nicely with 13x1x3 with 39 p-p nodes and 13 pme nodes.
Have you tried running with -smp 11? (I have found it better for my 980x)
OK, with -smp 52 I got first frame in 13:46 (22h56m40s) for project 2686 - PPD 120492.

-smp 7 -smp 8
TPF = 3.59 TPF = 3.56
PPD = 11808 PPD = 11965
CPU = 86-88% CPU = 96-98%
 
This is a little OT, but has anyone run multiple GPU clients specifically with Fermi cards (460 or faster) with the SR-2 and two L5640s? I am wondering what PPD hit there would be on the SMP client if say two GPU clients are running on a dedicated 'core' and we used the -smp 23 flag? With the new GPU3 WUs, two Fermi cards would add a massive amount of PPD and wondering if there would also be a corresponding massive drop with -bigadv?
 
This is a little OT, but has anyone run multiple GPU clients specifically with Fermi cards (460 or faster) with the SR-2 and two L5640s? I am wondering what PPD hit there would be on the SMP client if say two GPU clients are running on a dedicated 'core' and we used the -smp 23 flag? With the new GPU3 WUs, two Fermi cards would add a massive amount of PPD and wondering if there would also be a corresponding massive drop with -bigadv?

My guess is it wouldn't affect the tpf by more than 0x seconds. Running 1 or 2 gpu client uses so much little cpu (what, 1% of a quad core so 0.3% of 24threads) and uses the unused cpu cycles. I wouldn't set affinity on the gpus at all.
 
My guess is it wouldn't affect the tpf by more than 0x seconds. Running 1 or 2 gpu client uses so much little cpu (what, 1% of a quad core so 0.3% of 24threads) and uses the unused cpu cycles. I wouldn't set affinity on the gpus at all.
The new GPU3 WUs for Fermi consume a lot of CPU cycles. So much in fact that PG recommended not running these GPU3 WUs at all in conjunction with -bigadv. Of course, their concerns are solely centered on the quickest returns possible for WUs not total PPD
 
The new GPU3 WUs for Fermi consume a lot of CPU cycles. So much in fact that PG recommended not running these GPU3 WUs at all in conjunction with -bigadv. Of course, their concerns are solely centered on the quickest returns possible for WUs not total PPD

I was not aware of that. Should I be running -smp 11 then?
 
I was not aware of that. Should I be running -smp 11 then?
I honestly don't know. I don't have a S-1366 system and inquiring here because if I ever build one it would be dual-sockets and likely with a pair of L5640s.
 
They are correct. I have to set the slider down to about 80% cpu on the GPU client to get a good balance of -bigadv and gpu wu PPD.
 
They are correct. I have to set the slider down to about 80% cpu on the GPU client to get a good balance of -bigadv and gpu wu PPD.
Interesting, so how much does that affect the GPU production with your setup?
 
I lose about 2k ppd per 480, but I pick up 10k ppd on the-bigadv.
 
They are correct. I have to set the slider down to about 80% cpu on the GPU client to get a good balance of -bigadv and gpu wu PPD.

Yes, you have to reduce it to 70-80% cpu on the GPU client.
rolleyes.gif
 
I lose about 2k ppd per 480, but I pick up 10k ppd on the-bigadv.

Yes, you have to reduce it to 70-80% cpu on the GPU client.
rolleyes.gif
OK, understood how an appreciable balance can be struck with a little fine tuning. Is there a consensus that this method of manipulating the GPU CPU utilization is superior to dedicated core access for GPU clients?
 
I briefly ran four GPU clients (GTX 295) on my SR-2 without reducing SMP cores, and as you might expect it had a big negative effect on PPD. I would try setting core affinity on the SMP and GPU clients to make sure they don't conflict. That worked nicely on another box.
 
I briefly ran four GPU clients (GTX 295) on my SR-2 without reducing SMP cores, and as you might expect it had a big negative effect on PPD. I would try setting core affinity on the SMP and GPU clients to make sure they don't conflict. That worked nicely on another box.
Even with regular SMP without configuring priority for the GPU clients, they will drop production. Four GPU clients with the SR-2 is excessive, I was thinking more in line of two cards. I had 4 GPU clients in my Skulltrail and was compelled to remove a card. They were cannibalizing the -bigadv client horribly with the P2684. Three clients is acceptable but I want to reduce this machine down to 2 clients. Obviously, I don't see anywhere near the reduction compared to a dual S1366 system running concurrent GPU clients because my TPFs are so much lower, and thus I don't receive a high bonus to begin with. Moreover, with only 8 cores to work with, I cannot dedicate a core to GPU. That is precisely why I raised the question with dual hex-cores and 24 available threads. I'm wondering if using the -smp 23 flag would see that significant of a drop in -bigadv if there are GPU clients also present.
 
Even with regular SMP without configuring priority for the GPU clients, they will drop production. Four GPU clients with the SR-2 is excessive, I was thinking more in line of two cards. I had 4 GPU clients in my Skulltrail and was compelled to remove a card. They were cannibalizing the -bigadv client horribly with the P2684. Three clients is acceptable but I want to reduce this machine down to 2 clients. Obviously, I don't see anywhere near the reduction compared to a dual S1366 system running concurrent GPU clients because my TPFs are so much lower, and thus I don't receive a high bonus to begin with. Moreover, with only 8 cores to work with, I cannot dedicate a core to GPU. That is precisely why I raised the question with dual hex-cores and 24 available threads. I'm wondering if using the -smp 23 flag would see that significant of a drop in -bigadv if there are GPU clients also present.

I actually just did this:

Dual L5640s @ 3564, 12 Gb memory @ 1584 triple channel
P2686 (Run 3, Clone 19, Gen 15)
-smp 24 - 14:34/frame ~110K ppd
-smp 23 - 14:58/frame ~107K ppd
-smp 22 - 15:03/frame ~105.5K ppd

So -smp 24 is by far the best configuration. The difference between -smp 23 and -smp 22 is minimal, but at these speeds, the ppd difference more than you would expect.
 
I actually just did this:

Dual L5640s @ 3564, 12 Gb memory @ 1584 triple channel
P2686 (Run 3, Clone 19, Gen 15)
-smp 24 - 14:34/frame ~110K ppd
-smp 23 - 14:58/frame ~107K ppd
-smp 22 - 15:03/frame ~105.5K ppd

So -smp 24 is by far the best configuration. The difference between -smp 23 and -smp 22 is minimal, but at these speeds, the ppd difference more than you would expect.

Thanks for posting those...

I can't find my test figures, but I gave up on GPU on the SR2 altogether, as the loss in PPD was more than the gain. From memory it was something like taking PPD on a good bigadv from 140K ppd to 128K ppd - more than my GTX460 could do. (and more to the point, using an extra 100w to do the same thing)

That wasn't -smp 23, just letting them duke it out.
 
That wasn't -smp 23, just letting them duke it out.

Yeah, that doesn't work so well, and it's what I was talking about above. I have screenshots of HFM.net somewhere, but it was pretty ugly with the GPUs running.

On my K9A2 quad GX2 box, I finally decided to try manually setting affinity, with Muon1 on three cores and the eight F@H GPU clients on the fourth core. The PPD increase was pretty noticeable (from about 34K to 40K overall), even though the CPU usage on the dedicated core was low (5% or less). Muon1 performance went down of course, but I think the tradeoff was worth it. I'll try it on the SR-2 again eventually.
 
So -smp 24 is by far the best configuration. The difference between -smp 23 and -smp 22 is minimal, but at these speeds, the ppd difference more than you would expect.
-smp 24 is the best option relatively speaking in comparison to the other flags. With ~3k PPD advantage when a system is capable of well over 100k, it's not significant at under 3% difference of the total PPD. I would definitely consider adding video cards if it were mine. I understand that SR-2 owners are generally averse to running GPU clients due to power consumption concerns.

I guess no GPUs are installed in your system. Hypothetically adding one or two Fermi cards each achieving north of 10k and assuming (I realize a BIG assumption) the SMP client will not suffer additional preying from the GPU client(s), is something worthwhile considering, IMO. It's one big reason why I would never build a dual S-1366 system unless I had two hex-cores. With the L5640s it becomes an attainable possibility for many people, whereas previously it proved formidably expensive. 24 available threads permit all kinds of possibilities that systems with less cores/thread numbers have a harder time delivering.
 
Yeah, I just did some quick theoretical calcs using a points calculator, and losing 1 of 12 cores takes a 2686 project ppd from 142K to 125K for me. That 17K shortfall would need at least 2 GTX-460s to make up, and at higher power draw.

Assuming linear scaling, and no interference from whatever you might want to run on the freed cores (As Sazaneyes points out, duking it out is bad, but even setting affinity can't protect from shared memory issues):

24/24 threads = 12:20/frame = 142K ppd
23/24 threads = 12:52/frame = 133K ppd = 4.1% slower, but 6.3% less ppd
22/24 threads = 13:12/frame = 125k ppd = 6.5% slower, but 12% less ppd

This is way worse % loss than Musky measured, so worth practical testing. But HT makes everything more complicated.

When I get time I do want to test the difference between say -smp 22 or 23 and using WinAFC to manually change the affinity. I am loathe to run SMP flags all the time.
 
SR-2 #2 is running. This one looks like it is going to be a little tougher to get the memory set up correctly, although a managed to screw up the overclock on SR-2 #1 last night as well by switching back to matching memory.

SR-2 #1 - 18 x 180
SR-2 #2 - 18 x 175

I know what I will be doing this weekend...
 
Back
Top