• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

IvyBridge Xeons: Folding performance

Folding is significantly faster with NUMA on on all my systems

I can't speak for IB-E but if you're claiming "NUMA off is faster" on SB-E or
older intel chips or any AMD chip, there must be a bug in your firmware/OS
or problem with your method.
 
Thanks tear, interesting.

Can't talk for AMD based systems (don't have one).

Shared the observation with my systems
I used 2 motherboards: ASUS Z9PE-D8 WS, ASUS Z9PE-D16
CPUs: SB XEON E5-2687W, IB XEON E5-2697v2

Which components (mb, cpu) did you use to check the behavior of your Intel systems?
What was the result (difference in perf)?
 
I have two 2.3g v2 running on Supermicro's X9DAi board. Last night, I turned the NUMA to OFF and the same unit ran 10% more time, from 7'26" to 8'10" than the ON option.
 
I do not know about v2 I do not have any yet and do not plan on having any until there are some 46xx chips available, But I can confirm what tear says about the SB-E chips. When I first got my 4650's I tested with both Numa enabled and disabled and Numa enabled /on my SM XRQRI-F+ is quite a bit faster than Numa disabled / off iirc 8102 0% to 10% enabled was tpf 6:40 the same WU 0% to 10% with Numa disabled was tpf 7:20

Perhaps Intel did make some kind of change in the IB-E or perhaps there is something wrong with testing method. It might be something worth investigating.
 
thanks owc and Grandpa.

All of you used SM motherboards.
Anybody out there with Asus Z9PE boards?
 
I do, and I know Patriot does as well.

I do also, always ran it with NUMA enabled. Though, since my July PCIe BIOS trouble were solved, I have been pretty reluctant to reboot. I would be grateful if you or Patriot could carry out the experiment suggested by AndyE :D
 
I do also, always ran it with NUMA enabled. Though, since my July PCIe BIOS trouble were solved, I have been pretty reluctant to reboot. I would be grateful if you or Patriot could carry out the experiment suggested by AndyE :D

There is a long Asus Z9PE-D8 thread at overclock.net. I asked the guys there to cross check my observations.
http://www.overclock.net/t/1261060/asus-z9pe-d8-owners-thread/720#post_20872007

First responses seems to support the behavior of this setting.
Would be interesting if this whole thing is mobo manufacturer dependent.

We will see,
Andy
 
Or they just have the setting reversed :D

I'll give it a try now.
 
Or they just have the setting reversed :D
This is what I meant.
If this is the case, then the next question will be even more interesting to find out: How to test which implementation is "right" or "wrong" ?

Thanks for checking.
 
Unfortunately my results were inconclusive. It actually slightly hurt performance (11:42 vs 11:35 on an 8104) however my chips are a bit strange. They ignore the BIOS power limit settings and limit themselves to 95W no matter what. So it is possible turning "off" NUMA increased compute performance which caused the power consumption to increase and caused the chips to drop to a lower clock, thus negating the gains.
 
Thanks tear, interesting.

Can't talk for AMD based systems (don't have one).

Shared the observation with my systems
I used 2 motherboards: ASUS Z9PE-D8 WS, ASUS Z9PE-D16
CPUs: SB XEON E5-2687W, IB XEON E5-2697v2

Which components (mb, cpu) did you use to check the behavior of your Intel systems?
What was the result (difference in perf)?

I have hard numbers for Dual E5645s in SM X8DTi; each CPU loaded with 3GB
of memory (total of 6GB) -- gathered them as part of another exercise.

Test P6901 w/NUMA on: 15m50s TPF
Test P6901 w/NUMA off: 16m04s TPF

Also, to verify NUMA was working , I ran C version of STREAM benchmark 5.10
(http://www.cs.virginia.edu/stream/FTP/Code/) via taskset -c 1 (bound to second core).
Naturally, tests were conducted on an idle system.

When built with gcc -O2 (default):
NUMA off (best score out of 20 runs):
Code:
horde@fah-003048cbc23e:~/stream-5.10$ taskset -c 1 ./stream_c.exe
-------------------------------------------------------------
STREAM version $Revision: 5.10 $
-------------------------------------------------------------
This system uses 8 bytes per array element.
------------------------------------------------------------- 
Array size = 10000000 (elements), Offset = 0 (elements)
Memory per array = 76.3 MiB (= 0.1 GiB).
Total memory required = 228.9 MiB (= 0.2 GiB).
Each kernel will be executed 10 times.
 The *best* time for each kernel (excluding the first iteration)
 will be used to compute the reported bandwidth.
-------------------------------------------------------------
Your clock granularity/precision appears to be 1 microseconds.
Each test below will take on the order of 15724 microseconds.
   (= 15724 clock ticks)
Increase the size of the arrays if this shows that
you are not getting at least 20 clock ticks per test.
------------------------------------------------------------- 
WARNING -- The above is only a rough guideline.
For best results, please be sure you know the
precision of your system timer.
-------------------------------------------------------------
Function    Best Rate MB/s  Avg time     Min time     Max time
Copy:            8008.0     0.019991     0.019980     0.020009
Scale:           7781.4     0.020586     0.020562     0.020606
Add:             8117.2     0.029594     0.029567     0.029611
Triad:           8076.2     0.029746     0.029717     0.029762
-------------------------------------------------------------
Solution Validates: avg error less than 1.000000e-13 on all three arrays
-------------------------------------------------------------
horde@fah-003048cbc23e:~/stream-5.10$

NUMA on (best score out of 20 runs):
Code:
horde@fah-003048cbc23e:~/stream-5.10$ taskset -c 1 ./stream_c.exe
-------------------------------------------------------------
STREAM version $Revision: 5.10 $
------------------------------------------------------------- 
This system uses 8 bytes per array element.
-------------------------------------------------------------
Array size = 10000000 (elements), Offset = 0 (elements)
Memory per array = 76.3 MiB (= 0.1 GiB).
Total memory required = 228.9 MiB (= 0.2 GiB).
Each kernel will be executed 10 times.
 The *best* time for each kernel (excluding the first iteration)
 will be used to compute the reported bandwidth.
-------------------------------------------------------------
Your clock granularity/precision appears to be 1 microseconds.
Each test below will take on the order of 15354 microseconds.
   (= 15354 clock ticks)
Increase the size of the arrays if this shows that
you are not getting at least 20 clock ticks per test.
-------------------------------------------------------------
WARNING -- The above is only a rough guideline.
For best results, please be sure you know the
precision of your system timer.
-------------------------------------------------------------
Function    Best Rate MB/s  Avg time     Min time     Max time
Copy:            8700.9     0.018419     0.018389     0.018431
Scale:           8523.3     0.018804     0.018772     0.018899
Add:             9375.4     0.025628     0.025599     0.025652
Triad:           9317.1     0.025784     0.025759     0.025823
-------------------------------------------------------------
Solution Validates: avg error less than 1.000000e-13 on all three arrays
------------------------------------------------------------- 
horde@fah-003048cbc23e:~/stream-5.10$


When built with gcc -O9 -DSTREAM_ARRAY_SIZE=80000000 -DNTIMES=100:
NUMA off (best score out of 20 runs):
Code:
horde@fah-003048cbc23e:~/stream-5.10$ taskset -c 1 ./stream_c.exe
-------------------------------------------------------------
STREAM version $Revision: 5.10 $
-------------------------------------------------------------
This system uses 8 bytes per array element.
-------------------------------------------------------------
Array size = 80000000 (elements), Offset = 0 (elements)
Memory per array = 610.4 MiB (= 0.6 GiB).
Total memory required = 1831.1 MiB (= 1.8 GiB).
Each kernel will be executed 100 times.
 The *best* time for each kernel (excluding the first iteration)
 will be used to compute the reported bandwidth.
-------------------------------------------------------------
Your clock granularity/precision appears to be 1 microseconds.
Each test below will take on the order of 107101 microseconds.
   (= 107101 clock ticks)
Increase the size of the arrays if this shows that
you are not getting at least 20 clock ticks per test.
-------------------------------------------------------------
WARNING -- The above is only a rough guideline.
For best results, please be sure you know the  
precision of your system timer.
-------------------------------------------------------------
Function    Best Rate MB/s  Avg time     Min time     Max time
Copy:            8126.6     0.157614     0.157507     0.161598
Scale:           8116.7     0.157810     0.157699     0.159920
Add:             8261.2     0.232485     0.232412     0.232542
Triad:           8246.4     0.232900     0.232828     0.232990
------------------------------------------------------------- 
Solution Validates: avg error less than 1.000000e-13 on all three arrays
-------------------------------------------------------------
horde@fah-003048cbc23e:~/stream-5.10$

NUMA on (best score out of 20 runs):
Code:
horde@fah-003048cbc23e:~/stream-5.10$ taskset  -c 1 ./stream_c.exe
-------------------------------------------------------------
STREAM version $Revision: 5.10 $
-------------------------------------------------------------
This system uses 8 bytes per array element.
-------------------------------------------------------------
Array size = 80000000 (elements), Offset = 0 (elements)
Memory per array = 610.4 MiB (= 0.6 GiB).
Total memory required = 1831.1 MiB (= 1.8 GiB).
Each kernel will be executed 100 times.
 The *best* time for each kernel (excluding the first iteration)
 will be used to compute the reported bandwidth.
-------------------------------------------------------------
Your clock granularity/precision appears to be 1 microseconds.
Each test below will take on the order of 101459 microseconds.
   (= 101459 clock ticks)
Increase the size of the arrays if this shows that
you are not getting at least 20 clock ticks per test.
-------------------------------------------------------------
WARNING -- The above is only a rough guideline.
For best results, please be sure you know the
precision of your system timer.
-------------------------------------------------------------
Function    Best Rate MB/s  Avg time     Min time     Max time
Copy:            8851.6     0.144708     0.144606     0.148615
Scale:           8663.4     0.147841     0.147748     0.151805
Add:             9272.9     0.207140     0.207055     0.209344
Triad:           9209.8     0.208530     0.208473     0.209325
------------------------------------------------------------- 
Solution Validates: avg error less than 1.000000e-13 on all three arrays
------------------------------------------------------------- 
horde@fah-003048cbc23e:~/stream-5.10$

It would be very interesting to see if your machine yields similar STREAM results
-- they would tell us if OS+firmware are doing their jobs (natural assumption is
that each CPU is optimally populated with memory).

EDIT: quad SB-E system behaved similarly folding-wise but I don't have numbers recorded;
     I will gather them at next opportunity (maintenance/power interruption).
 
Last edited:
An idea - why not use numactl/numastat to show data/check if NUMA is working as expected?

Code:
indrek@2419x2:~$ numastat
                           node0           node1
numa_hit               266457046       215030938
numa_miss               36019018        97972016
numa_foreign            97972016        36019018
interleave_hit              3292            3182
local_node             266457007       215027747
other_node              36019057        97975207
indrek@2419x2:~$ numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 1 2 3 4 5
node 0 size: 2946 MB
node 0 free: 91 MB
node 1 cpus: 6 7 8 9 10 11
node 1 size: 3022 MB
node 1 free: 85 MB
node distances:
node   0   1
  0:  10  20
  1:  20  10
 
OS visibility of NUMA is not a problem (we confirmed that with fahdiag).

What doesn't fit is NUMA-off configuration yielding better folding performance on AndyE's machine.

Purpose of STREAM test is ensuring that NUMA actually *works* as advertised == memory nodes
are associated with respective CPUs [in SRAT].

In suggested test, NUMA-on configuration should yield visibly higher numbers than NUMA-off.
If it doesn't, something is b0rked.

Does that make sense?
 
some quick 1p v2 stats, all run on a single 12c/24t at 2.0-2.2 (not sure of speed at moment - need to reinstall i7z)

6099:- 3.54
8563:- 3.50
8569:- 4.07
 
Yes. It does.

What i meant was - with numa OFF, what will numastat say for

Code:
local_node             266457007       215027747
other_node              36019057        97975207

If the BIOS values are reversed and OFF=ON this would be the fastest indication + numerical proof for it. If im not mistaken.
 
What could be reversed is memory<->CPU mapping, that is, Node 0 CPUs being associated
with Node 1 memory and vice versa.

That would be pretty gross error on BIOS part, though...

EDIT: Enable/Disable options are unlikely to be reversed, fahdiag (it uses /sys/devices/system/node/ data) reported two nodes when performance was lower

And an afterthought:
OR! Or...
Another BIOS error could be: BIOS generating SRAT but enabling node interleaving in the CPUs/IMCs, and, symmetrically, not generating SRAT but disabling node interleaving.
 
Last edited:
DL380p Gen8 (11/04/2013 bios)

2x IB-E (1 2.3, 1 2.7) all core turbo via i7z 2.6ghz

16x ddr3 1600 8gb dimms reg

Project: 8103 (Run 0, Clone 59, Gen 163)
Numa ON / node interleaving disabled.
TPF = 9:42,44,44,40,43,45,45,44

NUMA OFF / node interleaving enabled
TPF = 9:47,48,48,47
 
Last edited:
some more smp numbers:-

Project 8533

2p E5 2665 32 threads @2.4 3:12 for 102k PPD

1p E5v2 269x 24 threads @2.1-2.2, 3:54 for 75.9k PPD
 
Can a single 2680v2 make meaningful BA ?

You need 8 v1 cores at 3.0 GHz to make an 8101 deadline. A 2680 v2 is 10 cores at probably 3.2 GHz, so it would make current deadlines easily. The problem you are going to start running into is, is -smp 20 a valid option for the folding client? That, I don't know.
 
What would be the PPD for a dedicated 2680v2 ? How many GB RAM needed ? OS: Ubuntu
Plan is still to get a 2P system but start with one P first.
 
You would need four sticks of memory to get quad-channel performance, any size - 4 x 1GB is plenty for folding. I'm calculating ~210K ppd on the best units and 135K ppd on the worst based on my v1 frame times.
 
Update:- A 1p 12 core at 2.1-2.2 will do BA units on time, 162k on 8104 :)
 
What would be the PPD for a dedicated 2680v2 ? How many GB RAM needed ? OS: Ubuntu
Plan is still to get a 2P system but start with one P first.

I would recommend the E5-4650 CPU even if you're only doing 2P. This CPU can be acquired for ~$500 each, cheaper than most E5-2680's I've seen and it will run bigadv as well.
 
I would recommend the E5-4650 CPU even if you're only doing 2P. This CPU can be acquired for ~$500 each, cheaper than most E5-2680's I've seen and it will run bigadv as well.

Those would be with "reduced" support from manufacturer ? When I checked cErtain Sources I saw Quite Some differences in price. But I'm also concerned in power drawing/heat production as it suppose to fold 24/7; seems those pull 130watt each. Ivys run a bit cooler. My GPU rig keep my room already nice warm :p
 
Those would be with "reduced" support from manufacturer ? When I checked cErtain Sources I saw Quite Some differences in price. But I'm also concerned in power drawing/heat production as it suppose to fold 24/7; seems those pull 130watt each. Ivys run a bit cooler. My GPU rig keep my room already nice warm :p

They won't pull 130W unless you stress all features of the CPU at once. F@H does not appear to max out the TDP of Intel CPUs in my experience.

Also, if you're thinking about spending money on a 2680v2, more power to you. For the price of one of those extra spicy versions, you could get 2 or 3 E5-4650's. For the price of a retail one, you could get all 4 E5-4650's and a 4-socket motherboard and pull down 950K PPD easy (not going to happen on a single or dual 2680v2 to my knowledge).

A quad E5-4650 pulls between 600-700W at the wall depending on memory population, etc.
 
I was hoping to provide a better update but murphy just keeps biting me in the ass.

1p folding on E5 v2 12c/24t at 2.1ghz all core turbo

8104:- tpf 16:33 for 170k PPD
8105:- tpf 20:56 for 176k PPD

Smp folding

8533:- tpf 4:05 for 70k PPD
8556:- tpf 4:05 for 70k PPD

Now the good bit. power from the wall 134w folding on all 24 threads. That's on an asus Z9PA-D8 with corsair 1600C9 1.35v ram, 160gb hdd, 1 hsf and 2 case fans. Powered by a seasonic platinum 400w fanless PSU
 
Nathan,
To what "equivalent" E5-xxxx model # CPU are these numbers attributable?
It would be nice to know that for "browsing" sake :)
Thanks.
 
There isn't a retail 12c at 2.0/2.1, the closest is a 2692 at 2.2, which would be slightly faster
 
I was hoping to provide a better update but murphy just keeps biting me in the ass.

1p folding on E5 v2 12c/24t at 2.1ghz all core turbo

8104:- tpf 16:33 for 170k PPD
8105:- tpf 20:56 for 176k PPD

Smp folding

8533:- tpf 4:05 for 70k PPD
8556:- tpf 4:05 for 70k PPD

Now the good bit. power from the wall 134w folding on all 24 threads. That's on an asus Z9PA-D8 with corsair 1600C9 1.35v ram, 160gb hdd, 1 hsf and 2 case fans. Powered by a seasonic platinum 400w fanless PSU

Wow that's an impressively efficient folding System by all accounts.
 
One more question: I use on my i7 an All-In-One liquid cooler: Seidon 120M. Easy to install and keeps the i7 cool. Nice entry level for liquid cooling.

As I'm still dream/plan/think on a 2p xeon folding rig I would wonder if those cooler would be ok for CPU folding too. From cooling performance and reliability point of view. Any experience with those or similar ?

Could I stack the radiators; maybe with a fan in between ?
 
Not really a folder, but I can add some 2p 2695 results for y'all coming up soon.
 
One more question: I use on my i7 an All-In-One liquid cooler: Seidon 120M. Easy to install and keeps the i7 cool. Nice entry level for liquid cooling.

As I'm still dream/plan/think on a 2p xeon folding rig I would wonder if those cooler would be ok for CPU folding too. From cooling performance and reliability point of view. Any experience with those or similar ?

Could I stack the radiators; maybe with a fan in between ?

I run mine with CM 212's and they have no issues. These xeons really don't run that hot. The seidon 120 should be perfectly fine for cooling, especially since they can't be oc'd. Just make sure they get good contact with the cpu's.
I probably wouldn't stack the rads though, they need fresh air to work properly. You're going to be adding the heat from the first rad to the second and then trying to pull that out.
 
Back
Top