• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

DL580 G7 with quad E7 4870

nwrtarget

Gawd
Joined
Aug 10, 2010
Messages
903
CPU
http://ark.intel.com/products/53579...7-4870-(30M-Cache-2_40-GHz-6_40-GTs-Intel-QPI)

So lets say I happen to be folding on one of these. I am pretty sure HT is disabled in Bios and it says on startup that it is mapping 40 to 32 threads. I started it with -smp 40 bigadv. It is running Linux. A little top action


top - 19:11:53 up 4 days, 4:56, 1 user, load average: 31.37, 24.78, 27.05
Tasks: 797 total, 1 running, 796 sleeping, 0 stopped, 0 zombie
Cpu(s): 0.0%us, 40.4%sy, 42.8%ni, 16.8%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Mem: 132142268k total, 5541424k used, 126600844k free, 101752k buffers
Swap: 4194296k total, 0k used, 4194296k free, 634524k cached

PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
29435 root 39 19 4743m 3.1g 3148 S 3198.8 2.5 125:53.01 FahCore_a5.exe


So it isn't fully loading the server (only 3200%). I could potentially enable HT and have 80 cores but would that help? I am pretty sure I saw a thread where some one asked a similar question about using more than 32 threads but I couldn't find it.

Currently folding a 6903 with 34 minutes ie 132K PPD or so which is a bit sad. I can't run wrappers or anything custom on this as it is destined for boring life as a production server after burn in but I could enable HT or tweak my config to get all of the performance out of this thing. 132K PPD is a little sad honestly for a server costing what this one does.

And yes it really does have 128 gigs of ram as that is what this particular model starts with.
 
Should it not be doing a whole bunch more?

That is the same as a SR-2 with l5640s at 3.2ghz
 
I agree is should be doing more. It isn't loading all cores in fact almost an entire CPU is idle in this thing right now. I could run two 20 thread bigadv instances but we already know that isn't the right way to go here.

If someone could point out the error in my ways it would be appreciated. I think it was some goofy flag that had to be set in the config file not through the command line.

Oh and this thing has 40 cores so could be 80 cores with HT enabled. 4 sockets 10 cores per socket.
 
hahahah Musky as a Nurse!


24% idle CPU is JUST WRONG!!!!

top - 19:34:12 up 4 days, 5:19, 1 user, load average: 32.00, 31.91, 30.81
Tasks: 797 total, 1 running, 796 sleeping, 0 stopped, 0 zombie
Cpu(s): 0.0%us, 13.6%sy, 62.5%ni, 24.0%id, 0.0%wa, 0.0%hi, 0.0%si, 0.0%st
Mem: 132142268k total, 7362476k used, 124779792k free, 102188k buffers
Swap: 4194296k total, 0k used, 4194296k free, 634560k cached

PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
29435 root 39 19 6726m 4.9g 3188 S 3198.1 3.9 839:40.29 FahCore_a5.exe
 
Ok, load average: 32.00, 31.91, 30.81 is looking better. What else is running on the machine?
 
Usual recommendations apply:

- enable HT (BIOS)
- enable NUMA / SRAT (BIOS)
- disable node interleaving (BIOS)
- make sure each CPU has memory
- use The Kraken

Other than that I'd be interested in output of this:
Code:
grep MemTotal /sys/devices/system/node/node[0-9]*/meminfo
 
grep MemTotal /sys/devices/system/node/node[0-9]*/meminfo
/sys/devices/system/node/node0/meminfo:Node 0 MemTotal: 33543856 kB
/sys/devices/system/node/node1/meminfo:Node 1 MemTotal: 33554432 kB
/sys/devices/system/node/node2/meminfo:Node 2 MemTotal: 33554432 kB
/sys/devices/system/node/node3/meminfo:Node 3 MemTotal: 33554428 kB

Oh and nothing else is running.
 
Ok, that tells us NUMA is most likely setup correctly.

That leaves HT and Kraken...
 
I know HT is off and I can't really run the kraken on this machine. It specifically says it is mapping down to 32 cores to go 8 x 4 x 1. If I go up to 80 via HT won't it just map it down to 64? It is just not using 8 cores for anything right now. I can hopefully turn on HT tomorrow and I know it will help out but I am just wondering if there is something else needed at a client level.



# Linux SMP Console Edition ###################################################
###############################################################################

Folding@Home Client Version 6.34

http://folding.stanford.edu

###############################################################################
###############################################################################

Launch directory: /root/folding
Executable: ./fah6
Arguments: -smp -bigadv

[21:05:19] - Ask before connecting: No
[21:05:19] - User name: nwrtarget (Team 33)
[21:05:19] - User ID: 50B0FC2755EDC8E0
[21:05:19] - Machine ID: 1
[21:05:19]
[21:05:19] Loaded queue successfully.
[21:05:19]
[21:05:19] + Processing work unit
[21:05:19] Core required: FahCore_a5.exe
[21:05:19] Core found.
[21:05:19] Working on queue slot 06 [November 19 21:05:19 UTC]
[21:05:19] + Working ...
[21:05:19]
[21:05:19] *------------------------------*
[21:05:19] Folding@Home Gromacs SMP Core
[21:05:19] Version 2.27 (Thu Feb 10 09:46:40 PST 2011)
[21:05:19]
[21:05:19] Preparing to commence simulation
[21:05:19] - Looking at optimizations...
[21:05:19] - Created dyn
[21:05:19] - Files status OK
[21:05:19] Error: Missing work file=<>
[21:05:19]
[21:05:19] Folding@home Core Shutdown: MISSING_WORK_FILES
[21:05:19] CoreStatus = 74 (116)
[21:05:19] The core could not find the work files specified. Removing from queue
[21:05:19] Deleting current work unit & continuing...
[21:05:19] - Preparing to get new work unit...
[21:05:19] Cleaning up work directory
[21:05:19] + Attempting to get work packet
[21:05:19] Passkey found
[21:05:19] - Connecting to assignment server
[21:05:20] - Successful: assigned to (130.237.232.237).
[21:05:20] + News From Folding@Home: Welcome to Folding@Home
[21:05:20] Loaded queue successfully.
[21:14:12] + Closed connections
[21:14:17]
[21:14:17] + Processing work unit
[21:14:17] Core required: FahCore_a5.exe
[21:14:17] Core found.
[21:14:17] Working on queue slot 07 [November 19 21:14:17 UTC]
[21:14:17] + Working ...
[21:14:17]
[21:14:17] *------------------------------*
[21:14:17] Folding@Home Gromacs SMP Core
[21:14:17] Version 2.27 (Thu Feb 10 09:46:40 PST 2011)
[21:14:17]
[21:14:17] Preparing to commence simulation
[21:14:17] - Looking at optimizations...
[21:14:17] - Created dyn
[21:14:17] - Files status OK
[21:14:30] - Expanded 57249576 -> 71846524 (decompressed 50.4 percent)
[21:14:30] Called DecompressByteArray: compressed_data_size=57249576 data_size=71846524, decompressed_data_size=71846524 diff=0
[21:14:31] - Digital signature verified
[21:14:31]
[21:14:31] Project: 6903 (Run 4, Clone 5, Gen 102)
[21:14:31]
[21:14:31] Assembly optimizations on if available.
[21:14:31] Entering M.D.
[21:14:42] Mapping NT from 40 to 32
[21:14:51] Completed 0 out of 250000 steps (0%)
[21:44:09] Completed 2500 out of 250000 steps (1%)
[22:18:26] Completed 5000 out of 250000 steps (2%)
[22:52:43] Completed 7500 out of 250000 steps (3%)
[23:27:00] Completed 10000 out of 250000 steps (4%)
[00:01:15] Completed 12500 out of 250000 steps (5%)
[00:35:31] Completed 15000 out of 250000 steps (6%)

Folding@Home Client Shutdown.


--- Opening Log file [November 20 01:06:47 UTC]


# Linux SMP Console Edition ###################################################
###############################################################################

Folding@Home Client Version 6.34

http://folding.stanford.edu

###############################################################################
###############################################################################

Launch directory: /root/folding
Executable: ./fah6
Arguments: -smp 40 -bigadv

[01:06:47] - Ask before connecting: No
[01:06:47] - User name: nwrtarget (Team 33)
[01:06:47] - User ID: 50B0FC2755EDC8E0
[01:06:47] - Machine ID: 1
[01:06:47]
[01:06:47] Loaded queue successfully.
[01:06:47]
[01:06:47] + Processing work unit
[01:06:47] Core required: FahCore_a5.exe
[01:06:47] Core found.
[01:06:47] Working on queue slot 07 [November 20 01:06:47 UTC]
[01:06:47] + Working ...
[01:06:47]
[01:06:47] *------------------------------*
[01:06:47] Folding@Home Gromacs SMP Core
[01:06:47] Version 2.27 (Thu Feb 10 09:46:40 PST 2011)
[01:06:47]
[01:06:47] Preparing to commence simulation
[01:06:47] - Looking at optimizations...
[01:06:47] - Files status OK
[01:07:00] - Expanded 57249576 -> 71846524 (decompressed 50.4 percent)
[01:07:00] Called DecompressByteArray: compressed_data_size=57249576 data_size=71846524, decompressed_data_size=71846524 diff=0
[01:07:01] - Digital signature verified
[01:07:01]
[01:07:01] Project: 6903 (Run 4, Clone 5, Gen 102)
[01:07:01]
[01:07:02] Assembly optimizations on if available.
[01:07:02] Entering M.D.
[01:07:07] Using Gromacs checkpoints
[01:07:19] Mapping NT from 40 to 32
[01:08:07] Resuming from checkpoint
[01:08:25] Verified work/wudata_07.log
[01:08:25] Verified work/wudata_07.trr
[01:08:25] Verified work/wudata_07.xtc
[01:08:25] Verified work/wudata_07.edr
[01:08:26] Completed 16775 out of 250000 steps (6%)
[01:17:28] Completed 17500 out of 250000 steps (7%)
[01:42:24] Completed 20000 out of 250000 steps (8%)
 
IIRC sfield did run w/HT on and hasn't mentioned any issues wrt thread number reduction (which doesn't
mean there weren't any).

In any way, you should get performance boost w/HT even if FahCore reduces the number of threads.

Why can't you run the Kraken? It's infinitely more open than FAH code :D (you build it from the source)
 
I know HT is off and I can't really run the kraken on this machine. It specifically says it is mapping down to 32 cores to go 8 x 4 x 1. If I go up to 80 via HT won't it just map it down to 64?

Not necessarily. I wouldn't expect it to use all 80 threads, but I'm not sure what it will drop down to. The issue you are seeing with only 32 of the 40 cores being used is due to Domain Decomposition (I asked the question you were trying to find). I'm not certain exactly what Domain Decomposition is. It is related to splitting up the processing among the threads requested. When you get to higher thread counts, it will drop down to a value that can be decomposed.

I just did a quick test on a P6903. Looks like it will work for all 80 threads:
[04:04:55] Mapping NT from 80 to 80
Starting 80 threads
Making 2D domain decomposition 8 x 1 x 10
 
IIRC sfield did run w/HT on and hasn't mentioned any issues wrt thread number reduction (which doesn't
mean there weren't any).

In any way, you should get performance boost w/HT even if FahCore reduces the number of threads.

Why can't you run the Kraken? It's infinitely more open than FAH code :D (you build it from the source)

It isn't about openess just about what the the guy who "owns" this server will think of me installing stuff. FAH is just a folder that can go away in the blink of an eye. He trusts the source of it and he is fine with it. IF I install anything else he will not be happy with me. He is letting me play with his toy while it isn't doing anything so I have to play by his rules.

I will play with HT tomorrow, if I can, and see where that goes.
 
Not necessarily. I wouldn't expect it to use all 80 threads, but I'm not sure what it will drop down to. The issue you are seeing with only 32 of the 40 cores being used is due to Domain Decomposition (I asked the question you were trying to find). I'm not certain exactly what Domain Decomposition is. It is related to splitting up the processing among the threads requested. When you get to higher thread counts, it will drop down to a value that can be decomposed.

I just did a quick test on a P6903. Looks like it will work for all 80 threads:

Ah very nice! So HT it is. Should I let it run out the unit it is on right now and try it on the next unit or just reboot and set it tomorrow morning?
 
Should I let it run out the unit it is on right now and try it on the next unit or just reboot and set it tomorrow morning?

Let it finish the current unit. Trying to change the thread count midway through a unit will cause issues.
 
It won't if checkpoint is sound (and that is a usual risk at every client restart == even if you don't change
the configuration).

nwrtarget, openness allows you to learn what the code does (so claim of openness does actually
pertain to trust) and, as a side note, Kraken is just another file in that folder which "goes away" ;)

Though your mind appears to be already set...
 
It won't if checkpoint is sound (and that is a usual risk at every client restart == even if you don't change
the configuration).

nwrtarget, openness allows you to learn what the code does (so claim of openness does actually
pertain to trust) and, as a side note, Kraken is just another file in that folder which "goes away" ;)

Though your mind appears to be already set...

Not my mind that is set here so let me explain.

If I "owned" the box I would give it a go but, since I don't, I don't have that lee way. I am just lucky that he is letting me on the thing. Technically I don't even work in that group anymore so I am more cautious than I would have been even last month.

Basically when a customer server gets assigned an engineer to set it up no one else is supposed to touch it. He is being nice and letting me play with his server. If he even thinks a little that I did something on there besides exactly what he expects then I won't be welcome next time. I install make and compile anything from source and this guy won't be friendly next time, that is how he is and I know that. I play by his rules on his server or I get out.

I may not even get to turn on HT... I will ask but may very well be told no. His box his rules.
 
Fair enough. Hope he agrees to enabling HT.
 
I'm sure tear could whip up a little shell script that has the same effect as theKraken but is run manually once per WU. Maybe that wouldn't be as scary.
 
Just use a compiled binary from another machine to install the kraken - two more files in the fah directory and nothing anywhere else. You don't need to compile it and add it to /usr/bin (make and make install) for it to work, which is your concern I believe.
 
Just use a compiled binary from another machine to install the kraken - two more files in the fah directory and nothing anywhere else. You don't need to compile it and add it to /usr/bin (make and make install) for it to work, which is your concern I believe.

How similar do the two machines have to be? Same kernel or just close? For example could I compile on CentOS 6.2 and then run it on Red Hat 6.X? If so then yes this is more possible. The only similarity between the hardware are that they are both Intel 64 bit. Is that close enough?
 
How similar do the two machines have to be? Same kernel or just close? For example could I compile on CentOS 6.2 and then run it on Red Hat 6.X? If so then yes this is more possible. The only similarity between the hardware are that they are both Intel 64 bit. Is that close enough?

Should definitely be close enough. I was going to give you one from a quad G34 machine to try. Yours would be much closer.
 
Last edited:
Found the instructions in your Linux guide so I guess I will give it a try tonight on something and then upload it once my current unit finishes.
 
Aye, there is. Don't remember the link though.
 
And yes it really does have 128 gigs of ram as that is what this particular model starts with.

Wow.

I imagine that in a few years we will all have 128 gigs of RAM in our desktops. But only if we are [H]ard.
 
Wow.

I imagine that in a few years we will all have 128 gigs of RAM in our desktops. But only if we are [H]ard.

The line between HDD and RAM are already blurring.... might end up you only need one in a few years.
 
The line between HDD and RAM are already blurring.... might end up you only need one in a few years.

It depends on what side of that line you like.

For long term storage I would prefer: non-volatile > volatile

oops, my ram drive battery died there goes all my stuff!
 
It depends on what side of that line you like.

For long term storage I would prefer: non-volatile > volatile

oops, my ram drive battery died there goes all my stuff!

I agree for hard storage that will be the case, but a good chunk of that will be in the cloud.

Flash is a middleground and prices are dropping.
 
I agree for hard storage that will be the case, but a good chunk of that will be in the cloud.

Flash is a middleground and prices are dropping.

As the manufacturing process gets smaller, which allows density to go up and prices to go down, flash becomes slower. Couple that with the wear issues and you will see that flash is a near term technology that barring any nice big breakthrough will be replaced by something else.
 
As the manufacturing process gets smaller, which allows density to go up and prices to go down, flash becomes slower. Couple that with the wear issues and you will see that flash is a near term technology that barring any nice big breakthrough will be replaced by something else.

I learned something today.

Would not have thought it got slower.

<the more you know rainbow>
 
With HT on and a 6901 I am looking at 7 minute frame times for 250K PPD roughly. See if it grabs a big one next.
 
7min on 6901 sounds similar to what I was getting on the bl680g7 with same chips....

I was rather disappointed though... somehow expect an intel box w/ 40/80 @2.4ghz to beat a 585g7 w/6172s but it doesn't...
 
7min on 6901 sounds similar to what I was getting on the bl680g7 with same chips....

I was rather disappointed though... somehow expect an intel box w/ 40/80 @2.4ghz to beat a 585g7 w/6172s but it doesn't...

It is curious, and since this hardware is so few and far between, we may never know why the performance is so low.

It seems as if the 8 core xeons performed much better than the 10 core ones...
 
It is curious, and since this hardware is so few and far between, we may never know why the performance is so low.

It seems as if the 8 core xeons performed much better than the 10 core ones...

That was an 8p...
but I get what you are saying...
Maybe when they increased L3 to 30mb they made it slower?

I will have to experiment....
 
Last edited:
That was an 8p...
but I get what you are saying...
Maybe when they increased L3 to 30mb they made it slower?

I will have to experiment....

Well, you would think 80 threads would have more performance than 64 threads. I have not seen what a full glorious 128 threads would do :D.

We need more DATA!!!
 
7min on 6901 sounds similar to what I was getting on the bl680g7 with same chips....

I was rather disappointed though... somehow expect an intel box w/ 40/80 @2.4ghz to beat a 585g7 w/6172s but it doesn't...

I have to agree as I thought a 40c/80t setup would do better. My 4P 6166HE @ 2.1 GHz is doing 6m34s on 6901s.
 
6903 TPF on first is 16' 40" for a listed PPD of 389K Completion in 1.16 days so that isn't to bad but really not much of an improvement over the previous CPU's in this same platform.

When I saw that they were putting the E7 CPU's in the G7 platform I was suspicious that the platform would hold the CPU's back. That may be what it happening here...

Update
The DL580 G7 still relies on the 7500 chipset but the DL380 Gen8 uses the C600 chipset. I wonder if the E5 in the Gen8 will perform on par with the DL580 G7 as configured. The DL380 G7 was so close that it almost wasn't worth buying the 580. That gap appears to have done nothing but get more narrow with the Gen8 although I haven't yet gotten my hands on a Gen8 server to verify this.
 
Last edited:
Back
Top