• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Tweaking a RP E5-4650 build

Joined
Apr 6, 2013
Messages
58
Just finished my 4P 4650 build and I'm looking for some advice to decrease the TPFs and of course increase PPD.

Here is my system:
1x SM X9QR7-TF+ mobo
4x E5-4650 ES C0
4 x SuperMicro SNK-P0050AP4 Heatsink
18x Hynix 4GB PC3-10600R ECC 2Rx4 (HMT151R7TFR4C-H9)
1x Corsair AX-1200 PSU

I have VT turned off and the turbo (custom power) on.

Prior to turning on the turbo I ran two 8104's:
(R0, C59, G65) @ 6:54 TPF and 610.2k PPD (w/o Kraken)
(R0, C23, G49) @ 5:18 TPF and 897.9k PPD (w/Kraken)

Right now I'm running an 8101 with Kraken & Turbo on:
(R13, C1, G247) @ 9:17 TPF and 614.1k PPD (@5%)

Are there any other settings I need to make and/or change in the BIOS. I thought I read somewhere there were a few other settings to change but I can't find them again.

I know that 1600MHz memory might help, but the 1333's will have to do for now.

Thanks in advance :)
 
Properly configured E5-4650 4P gets about 1M+ PPD on non-8101 BA units. Fact... end of story.
 
I am running the bios settings the boards have as-delivered - the only thing I changed was the time and boot order to install an OS. For OS tweaks, I just used the fahinstall script and nothing else.

Your ppd w/kraken look in-line with what you should see. I would expect an 8102/8103 to get you closer to 975K+ ppd.
 
This is what I've seen so far:
P8101 (R13, C1, G247) - TPF 9:16 - 599.1k PPD
P8103 (R0, C11, G85) - TPF 7:15 - 859.3k PPD

This is what I'm running right now:
P8104 (R0, C39, G60) - TPF 5:17 - 942.3k PPD (14% completion)

I'm thinking that 1600MHz ECC RDRAM is the next step in order to up the numbers.
 
Last edited:
Is the 18 sticks of ram correct? If so, is there a chance that those two extras are causing the memory to run below quad-channel? I know I saw a noticeable improvement when I went from dual- to quad-channel.
 
Is the 18 sticks of ram correct? If so, is there a chance that those two extras are causing the memory to run below quad-channel? I know I saw a noticeable improvement when I went from dual- to quad-channel.

Well, I went with the recommended SM memory configuration. I didn't use the "all blue" memory slots, like I did with my 4P SM AMD board, for quad channel.

Right now I'm using the following recommendation:
CPU-1 A1/2 B1/2 C1/2
CPU-2 E1/2 F1/2
CPU-3 J1/2 K1/2
CPU-4 N1/2 P1/2

I guess I'm basically using dual-channel and not quad. I can always run all the blue slots as well if someone thinks that will work even better.

I'm open to suggestions.
 
The system will use quad channel where it can and dual channel where necessary. The issue is, you basically can't control the location of the memory footprint of the application. If the memory allocation of the OS places it in the region of dual channel memory your perf suffers. I'd only use 16 memory sticks and leave the 2 out. They create more issues than benefit here. These asymetric memory structures are less impacting other workloads like databases, but by and large its always better to run the system in symmetric mode.
 
The system will use quad channel where it can and dual channel where necessary. The issue is, you basically can't control the location of the memory footprint of the application. If the memory allocation of the OS places it in the region of dual channel memory your perf suffers. I'd only use 16 memory sticks and leave the 2 out. They create more issues than benefit here. These asymetric memory structures are less impacting other workloads like databases, but by and large its always better to run the system in symmetric mode.

Point taken...and tyvm!

Edit:
Took out the two "offending" boards and then redistributed the ram to the "blue" slots to enable the quad-channel.

TPF dropped down momentarily to 5:02 for a ~980k PPD (down from the previous 5:17 earlier).

Still think that 1600MHz RAM will help some more...
 
Last edited:
The system will use quad channel where it can and dual channel where necessary. The issue is, you basically can't control the location of the memory footprint of the application. If the memory allocation of the OS places it in the region of dual channel memory your perf suffers. I'd only use 16 memory sticks and leave the 2 out. They create more issues than benefit here. These asymetric memory structures are less impacting other workloads like databases, but by and large its always better to run the system in symmetric mode.
For several years now, CPUs have been employing multiple independent channels
(controllers) without "ganging" (as you know it from socket 939 days -- expanding
effective channel width and working synchronously).

Four channels, with one DIMM each do _not_ work together (thus, there are no requirements
w.r.t. size, timings, or width/ranks/density when populating individual channels).

Instead, memory access logic interleaves memory accesses between multiple channels
(of single node) on page boundary (typically maps certain bits of physical address
to controller index but more elaborate schemes are available, too) -- that's how parallelism
is accomplished. That said, you don't really need to worry about memory population within
single node -- it's taken care of automatically (memory allocations are normally
contiguous so interleaving exploits all controllers populated with memory).

Controlling memory placement node-wise (as opposed to channel-wise) is something that
can be done and is done. NUMA-aware kernels (any stock Linux kernel today is
NUMA-aware) first attempt to allocate memory on the "home" node of the CPU that is
faulting a page.

If a thread is running on CPU16 (which is, say, on Node 4) and it faults a page, the kernel
will first try to find the memory on node 4 (to ensure memory locality).

Kraken binds FAH worker threads to respective CPUs... kernel does the rest (that is,
closest-node memory allocation).

The problem that OP had was that some channels were completely unpopulated due
to SM-recommended configuration (which is silly in the least). It's the general distribution
that was causing the degradation, not the "2 extra DIMMs".

One can determine configuration of each controller by running quick-hacked TPC version
that (only) supports -dram on SB-E: http://darkswarm.org/tpc/testing/tpc-svn69-sb-e.tar.gz
 
Instead, memory access logic interleaves memory accesses between multiple channels
(of single node) on page boundary (typically maps certain bits of physical address
to controller index but more elaborate schemes are available, too) -- that's how parallelism
is accomplished. That said, you don't really need to worry about memory population within
single node -- it's taken care of automatically (memory allocations are normally
contiguous so interleaving exploits all controllers populated with memory).

tear,
if an application is memory bound one should care about single node memory placement. BTW, there are also multiple interleaving schemes working concurrently. Address, page, rank and channel interleaving to name a few and depending on the available capabilities of the controllers and the memory sticks the performance implications with some apps can be significant.

If the application is CPU bound and has a memory footprint which can be predominantely served from the cache hierarchy, the impact of suboptimal placed single-node memory configurations is less visible.

If people are interested, some overview info:
Fujitsu Siemens E5-2600 memory whitepaper (applicable to E5-4600 as well)
Ulrich Drepper (CTO, RedHat):What every programmer should know about memory
John McCalpin's excellent 5-part series on memory performance with AMD CPU's (John McCalpin is the inventor of the STREAM benchmark)

Andy
 
And Fujitsu paper shows exactly what I said -- see table on page 14.

4-way has a reference performance of 1.00 while all others -- sub-one.

Point is, with _given_ set of sticks, uniform distribution of DIMMs across channels works better than
selective population (OP's problem).

The other point is that the initial 2 extra sticks weren't breaking anything and their removal (without
touching other sticks) would have gained nothing.
 
Thanks to all for your help, and your words of wisdom...I finally hit the one million mark on an 8104 (R0, C28, G62), TPF 5:02. Hit 1,012,109.7 at the 81% point.

Don't know as to the final numbers, but this is surely promising. Now all I need is some 1600MHz ECC RDRAM and maybe I can keep repeating this result.
 
While it may make some improvement - I wouldn't do it unless you can get the ram for free. I was BIOS limited on my 2P SB-E to 1333 for a little while - and once a new BIOS came out that supported 1600 I switched to it and saw no change (within the noise of WU variability).
 
Yeah, DDR3-1600 is only supported w/reg ECC memory on this platform (ECC UDIMMs
won't do). Even so, improvements are miniscule.

Money's better spent on another system...
 
Yeah, DDR3-1600 is only supported w/reg ECC memory on this platform (ECC UDIMMs
won't do). Even so, improvements are miniscule.

Money's better spent on another system...

Point taken.

In light of that, I have another 4P 2011 mobo, one that needs RMA due to bent pins. I still need to source another four CPUs, memory, PSU, and SSD, so it won't be up for a month or two.

BTW, where is the spreadsheet that lists the top folders by PPD? I found it once but I can't remember where it is.

Thanks again.

EDIT:

NVM, I found the chart. A wonderful thing this "search" function...duh!
 
Last edited:
Back
Top