• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Build Help: ESXI Server

Gregyski

n00b
Joined
Jun 7, 2011
Messages
26
This build is complete. It is now in the storage benchmarking phase discussed below.

I'm planning to build a new ESXi server and would appreciate advice. This will be my first experience with ESXi, but I have a good friend who runs several. I'm particularly interested in anyone who has successfully used a LSI MegaRAID 926x-8i with the SuperMicro X9SCA MB as I've read that the LSI RAID cards can be extremely finicky about MB choice.
  1. ESXi server to run at least one Windows server, a couple of Windows workstations, and some Linux VMs - none with particularly heavy duty. It will also handle data storage. I want this to last me a lot of years - I won't be able to afford this expense again.
  2. Budget was approximately 2000 USD, but I've already gone over - I'm flexible but cost is obviously important
  3. FL, USA
  4. I will list my parts below
  5. No reused parts (aside from common hardware/cables)
  6. No OC
  7. No dedicated monitor
  8. I plan to purchase within a few days
  9. The features I need will be evident from my build or are discussed below
  10. I have all the OS licenses I need
Here is what I have come up with so far, all currently from Newegg:

CPU/MB combo: 510 USD
Intel Xeon E3-1270 Sandy Bridge 3.4GHz LGA 1155 80W Quad-Core Server Processor BX80623E31270
SUPERMICRO MBD-X9SCA-F-O LGA 1155 Intel C204 ATX Intel Xeon E3 Server Motherboard
Case: 150 USD
Antec Performance One Series P183 V3 Black Aluminum / Steel / Plastic ATX Mid Tower Computer Case
PS: 160 USD (I wanted as power efficient a PS as possible - I would think 550W will be sufficient even if I add some drives)
KINGWIN Lazer Platinum Series LZP-550 550W ATX 12V v2.2 / EPS 12V v2.91 / SSI EPS 12V v2.92 SLI Ready CrossFire Ready 80
RAID: 495 USD
LSI MegaRAID Internal Low-Power SATA/SAS 9261-8i 6Gb/s PCI-Express 2.0 w/ 512MB onboard memory RAID Controller Card, Single
Cables: 35 USD
3ware CBL-SAS8087OCF-06M 0.6m (SFF-8087) SAS/SATA breakout cable, forward, 4x (SFF8482) with power
HDs: 600 USD
4x Seagate Constellation ES ST31000424SS 1TB 7200 RPM SAS 6Gb/s 3.5" Internal Hard Drive -Bare Drive
RAM: 240 USD
2x Kingston 8GB (2 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333 (PC3 10600) ECC Unbuffered Server Memory Model KVR1333D3E9SK2/8G
Total: 2200 USD (including shipping)

I plan to run the drives, most likely, in a RAID6 configuration with the intention of live growing the array later. If (live) converting from 5 to 6 is an option with my RAID solution, I may do RAID5 initially until I add drives. I really wanted to get 2 sticks of 8GB for growing room later, but 8GB ECC UDIMMs aren't generally available yet. I'm hoping ESXi will support the temperature sensors for the MB since I like to monitor temps, but I don't know how to be sure of that. Nevertheless, it's a cool case in an almost always air-conditioned room.

Some alternate hardware worth considering:

RAID: 440 USD (I prefer the LSI passive cooling since the case should be adequately cooled;perhaps better management due to management port?;LSI has better documentation)
areca ARC-1222 PCIe x8 SATA / SAS (Serial Attached SCSI) RAID Card
HDs: 600 USD (Twice the storage for the same price but desktop SATA drives)
4x Western Digital Caviar Black WD2002FAEX 2TB 7200 RPM 64MB Cache SATA 6.0Gb/s 3.5" Internal Hard Drive -Bare Drive
HDs: 750 USD (RAID5 configuration - this option costs more and seems to have slightly more risk of data loss, but more future-proof and more space - will have to migrate to RAID6 if I add even one more drive)
3x Seagate ST32000445SS 2TB 7200 RPM 16MB Cache SAS 6Gb/s 3.5" Constellation ES Hard Drive with Secure Encryption
Case: 370 USD (+SAS backplane, +redundant PS, -less power efficient, -huge, -don't fully trust SM's PSes, -expensive)
Supermicro CSE-743TQ-865B-SQ Black Pedestal Server Case 865W 2 External 5.25" Drive Bays
Case: 250 USD (+SAS backplane, (big negative) -limited to 4 HDs, -less power efficient, -replacing PS could be problematic)
Supermicro Case CSE-733TQ-665B-DIST Black Workstation Mid Tower Dual 4 SATA/SAS Hot Swap HDD REV K with 665W Power Supply
Whole system: lotsa USD
Dell T410

Thank you for your sage advice. This is a really huge expense for me so I'm hoping to get it right.
 
Last edited:
I've moved this over to the Data Storage subforum due to your concerns over the LSI card and such..

Oh and don't get the Antec P183 as it's not worth its price at all, even for a quiet PC. I recommend any of these cases over the Antec:
$70 - Lian Li Lancool PC-K58W ATX Case
$90 - Lian Li Lancool PC-K56 ATX Case
$90 - Cooler Master CM690 II Advance ATX Case
$100 - Cooler Master HAF 922 RC-922M-KKN1-GP ATX Case
$90 - Lian Li Lancool PC-K7B ATX Case
$106 - Lian Li PC-7B Plus II ATX Case
$127 - Fractal Design Define R3 Black ATX Case
$110 - NZXT Whisper WHI - 001BK ATX Full Tower Case
$120 - Velocity Micro GX2-W Silver Classic Aluminum Case with Side Window
$140 - NZXT Phantom PHAN-001WT White Full Tower ATX Case
$140 - NZXT Phantom PHAN-001BK Black Full Tower ATX Case
$140 - Cooler Master HAF 932 RC-932-KKN1-GP ATX Case
$165 - Corsair Graphite Series 600T ATX Case
$160 - Silverstone RV02B-W ATX case
 
As an Amazon Associate, HardForum may earn from qualifying purchases.
Thanks for the input Danny. I had debated whether to post here instead of General Hardware since RAID discussion is thicker here, but I didn't want to chance it since this isn't going to be dedicated NAS. But I think you are right I'll likely get better input on this server here - particularly the thing that has me most concerned: the MB/RAID combo.

As for the case, I actually researched most of the cases you suggested, although I'll revisit them and check a few that I hadn't. Perhaps I'm not objective because I love my P180 (current server) and P182 (current gaming rig) cases, but it seemed to fit my preferences pretty well. Obviously a server pedestal chassis has a lot of advantages but I couldn't find one where the pros outweighed the cons for me.

Since I felt the P180 and P182 were worth the cost, it's possible I just have different preferences from you, but are there any particular issues you have with the P183 that I may not be considering or missed in reviews?
 
Since I felt the P180 and P182 were worth the cost, it's possible I just have different preferences from you, but are there any particular issues you have with the P183 that I may not be considering or missed in reviews?

Well just a FYI, the case for my main gaming rig is a P180 also. I still love it despite its many problems. Yes, the P183 is a good case overall but just not worth $150 in terms of quality and what you're actually getting.

Take a look at the Fractal Design R3 case, the Corsair 600T, and the Silverstone RV02 for example: The Fractal Design case is almost a copy of the Antec P183 design wise but at a lower price, with more cable management options, slightly better cooling, more room for hard drives, quieter and slightly better performing fans, and a helluva lot easier to work with than the P183. The Corsair 600T and Silverstone RV02 cases are almost in the same price range as the Antec P183 but have higher quality and significantly better cooling. In fact, the RV02 has better cooling than the Antec P183 while being just as quiet.

I just don't see how the P183 would be worth getting over just those three cases alone. If you're aiming for quiet, either the Fractal Design or Silverstone case would be a better choice for the money.
 
Thank you Danny, I understand your feelings towards the P183 much better now and will investigate those cases you mention more closely. Indeed both sufficient cooling and a low noise level (but not silent) are quite important to me.

Edit: I should mention one important factor regarding case selection: I want a capacity for at least 6 internal 3.5" HDs. 8 is preferable. I also prefer full towers for the extra space in which to work, but that's not critical.
 
Last edited:
Hi gregyski, nice to see you over here.

Thank you Danny, I understand your feelings towards the P183 much better now and will investigate those cases you mention more closely. Indeed both sufficient cooling and a low noise level (but not silent) are quite important to me.

I agree with Danny on the Fractal imo would be a better choice, even SPCR has it on their recommend list, Fractal Design Define R3 ATX Tower, a lot of the favoring of the P183 in SPCR is the usage of CP-850 and they like to supend drives, with wanting 6 to 8 hdd i doubt you will be able to. The Define R3 has 8 hdd slots and has sound dampening all over the case, with beautiful finish and a lot of holes/grommets to route cables, and its cheaper. i would probably swap the fans if you want a more quiet setup, i would probably go with Scythe Slip Stream 120mm x 25mm Fan - 800 RPM (SY1225SL12L), not so expensive and decent airflow, or if you are fan on spend high $$$ on fans, Noiseblocker NB-Multiframe M12-S1 120mmx25mm Ultra Silent Fan - 750 RPM - below 6 dBA. Or if you go PWM fans check Scythe Slip Stream 120mm x 25mm PWM Mid Speed Fan - (SY1225SL12LM-P).

I should mention one important factor regarding case selection: I want a capacity for at least 6 internal 3.5" HDs. 8 is preferable. I also prefer full towers for the extra space in which to work, but that's not critical.
I think the Fractal Design Define R3 Black ATX Mid Tower Silent PC Computer Case should be a good choice, but out of you liking bigger cases, another option is their full tower, Fractal Design Define XL Black ATX Full Tower Silent PC Computer Case with 10x hdd slots and 5x 5.25 slots, but this case uses 140mm fans and 1x 180mm fan, so would be harder to get quieter fans.

On the controllers/compatibility, hope someone else can help you here.

Good Luck,
 
Last edited:
Hi Abula, thanks for suggesting I post here. I had been considering it but your suggestion pushed me over the edge. :) Regarding the controller, I hope so too. As much as I greatly appreciate input on my other selections, hardware compatibility is my primary concern.

I've visited/revisited all the cases Danny suggested. Unfortunately, due to my drive bay requirements and the fact that I only like minimalist, clean, rectangular cuboid cases, the only alternative to the P183 that I liked was the Define R3. Because of this I've read and watched a slew of reviews. There are a few issues I have with the R3, however:

  1. This server will be in a living area and its having a power button so accessible on a server that itself will virtualize all my other servers is nerve-wracking in the extreme. I really wish it were inside the door.
  2. I'm a little concerned about heat as every reviewer said it was significantly warmer than most enthusiast cases with the default fans. While I wont have a hot video card like a gaming rig, the RAID card will put out a decent amount of heat and, being passive (if I go with the LSI), I want to make sure it stays nice and cool. Adding extra fans is an option, of course.
  3. I will only have easy access to the left side of the case once it's gone live, so I think I'll have to install the HDs with the ports towards the left instead of the right. I assume this will work though.
  4. I generally prefer full tower cases.
Despite those concerns, it has some definite advantages over the P183. I quite like it and currently consider it on par with the P183 for my own tastes. I'll just have to go with my gut come purchase time.

I didn't really like the Define XL at all, despite my preference for full tower. I just don't like the layout. I'd definitely go for the Define R3 over the XL.
 
What's the purpose of this? Production or just a home lab? Don't waste your money on Xeon CPU/Mobos and ECC-RAM. Just built an ESXi host from the following:

P8P67 Pro
Core i7 2600K
4x4GB GSKILL DDR3-1600 Ripjaws X
Intel Desktop Gigabit CT PCI-e NIC

Works just fine, OC'd to 4.5GHz too.

BTW, get yourself a Perc RAID card and flash it to LSI firmware. You'll save a *lot* of cash that way too.
 
I don't understand why you want 1TB/2TB sas drives. Normally you'd go with small sas drives for maximum iops/$ or large sata drives for maximum GB/$.
 
What's the purpose of this? Production or just a home lab?
Thanks for your suggestions. I'll certainly investigate them. For all intents and purposes it's a production server, but it'll be housed in a small home office so it's location in a living space is a bit out of the ordinary. I wouldn't call the scenario ideal, but it's tenable. Additionally, I'm usually comfortable doing exotic things with my systems, but I was trying to stay as much within-bounds on this build regarding things like hardware compatibility lists, OCing, etc. I'm willing to pay a little more for piece of mind with this system. Unfortunately, I can't play completely within the bounds as ESXi's HCL is limited in regard to new chipsets and LSI's HCL is fairly useless with regard to new hardware. I admit that I may be being a bit overly-conservative and I'll give your suggestions some thought.

Edit:
I don't understand why you want 1TB/2TB sas drives. Normally you'd go with small sas drives for maximum iops/$ or large sata drives for maximum GB/$.
Indeed with storage I'm more interested in GB/$ - I do not need highly performant storage. However, nearline SAS drives are about the same price as enterprise SATA drives and have a few advantages. And I've had too many problems using desktop SATA drives in a RAID configuration to feel confident about that.
 
Thanks for your suggestions. I'll certainly investigate them. For all intents and purposes it's a production server, but it'll be housed in a small home office so it's location in a living space is a bit out of the ordinary. I wouldn't call the scenario ideal, but it's tenable. Additionally, I'm usually comfortable doing exotic things with my systems, but I was trying to stay as much within-bounds on this build regarding things like hardware compatibility lists, OCing, etc. I'm willing to pay a little more for piece of mind with this system. Unfortunately, I can't play completely within the bounds as ESXi's HCL is limited in regard to new chipsets and LSI's HCL is fairly useless with regard to new hardware. I admit that I may be being a bit overly-conservative and I'll give your suggestions some thought.

Cool. Depending on your SLAs I would consider consumer hardware for use as production in terms of stability. Been hammering the living crap out of my new 1155 system as well as my overclocked i7-860 system for 6 months now without any issues.

The limiting factors of consumer hardware are the NICs and any onboard RAID. Any and all Realtek NICs will not work and while the P55/P67 chipset SATA ports work, you cannot utilize RAID.

To remedy this, buy a RAID card and some NICs on the HCL. Perc 5 or 6 is a very good choice in terms of performance and availability, you could grab 3-4 of them for the price of a decent LSI. Single-port gigabit Intel or Broadcom NICs are widely available for under $40, I snagged two new dual-port gigabit Intel NICs for $130 shipped not too long ago either.

Another option is to pick up used rackmount servers on the HCL. Dell 1950, Dell 2950 or HP DL360 G5 would work well but you will pay out the ++++++++ for ECC RAM so weigh your options accordingly.
 
Edit:
Indeed with storage I'm more interested in GB/$ - I do not need highly performant storage. However, nearline SAS drives are about the same price as enterprise SATA drives and have a few advantages. And I've had too many problems using desktop SATA drives in a RAID configuration to feel confident about that.

I would also recommend reconsidering the large SAS drives. At <140 I/O per 10K SAS drive it's not hard to be limited by your storage.

What's going to be running on the ESXi host?
 
Cool. Depending on your SLAs I would consider consumer hardware for use as production in terms of stability.
Well by production I mean it will be performing some production roles, but I have no SLAs to honor and don't expect any in the future. But yes, I understand. My initial plan was to buy all consumer hardware and I started my purchasing research with that as the plan. In fact, that's how I've always done my servers in the past - they were all consumer builds. But I've always found it a little limiting and sometimes regretted that choice. As long as it wasn't cost prohibitive, I felt like doing things "right" for this build and the price difference between consumer and enterprise hardware is not nearly as huge as it used to be. So this seemed like a good time to indulge a bit.
I would also recommend reconsidering the large SAS drives. At <140 I/O per 10K SAS drive it's not hard to be limited by your storage. What's going to be running on the ESXi host?
Initially: a Windows server PDC / software dev / light-duty services / data storage system, a few light-duty Windows workstations (probably only 1-2 running 24/7), and a few light-duty Linux VMs. However, I'm confident it really will not be using much simultaneous I/O. They will be largely idle - important to me, but idle. :)
 
It's possible I cast too wide a net. Would it objectionable for me to create a new post for just my single question regarding RAID/MB compatibility? I'm glad I made this post as I received a lot of good suggestions, but the primary issue is something I really need to try to work out if at all possible and as soon as possible. If not, no worries. Thanks.

Edit: Disregard the above as I'm about to pull the trigger. I decided to go with the Define R3 instead of the P183 (and it was close, I changed my mind around a dozen times) but otherwise the build stands as originally proposed. I will report back here with my experiences once it's up and running.

While I never got any replies regarding the MB/RAID combo, and more specific searches came up dry, I found a couple system integrators using unspecified Supermicro MBs w/ C202 and C204 chipsets that included an option for the LSI MegaRAID 9260-4i. So I suspect that my choice will be okay. Thanks again all for the input. I'll be in touch.
 
Last edited:
IWhile I never got any replies regarding the MB/RAID combo, and more specific searches came up dry, I found a couple system integrators using unspecified Supermicro MBs w/ C202 and C204 chipsets that included an option for the LSI MegaRAID 9260-4i. So I suspect that my choice will be okay. Thanks again all for the input. I'll be in touch.
Sorry you didnt get the answers you were looking, hopping the setup you chose will be compatible.

Im really interested into how your build turns out, a lot of the components you chose, im planing to use on my next server (aside from the psu and your hdd will be almost identical), so if have some time and feel like it, pls post some pics and your experience.

Btw was chekcing the mobo, and it has 4x 4pin fan connectors, so im assuming it would be capable of controlling the PWM fans rpm through the bios (not sure though), but might be worth considering if you end up changing the fans.
 
Last edited:
Good luck!
Thanks. :)
Im really interested into how your build turns out
Yep! I'll report back once she's running.
Btw was chekcing the mobo, and it has 4x 4pin fan connectors, so im assuming it would be capable of controlling the PWM fans rpm through the bios (not sure though), but might be worth considering if you end up changing the fans.
I'll be using a mix of stock and 3rd party fans. I haven't decided what will be necessary so I bought a few decently quiet and fairly cheap fans (three pin though) and will decide when I build. And, regarding controlling the fans, yes I hope so. Since I suspect I won't be able to use SpeedFan like I do with my current systems, I'm hoping I'll be able to have at least some control through the BIOS. Otherwise, I can use the manual control that comes with the case to control three of them. I'll be playing it by ear.
 
The build is complete. I'm going to post up my impressions of the hardware I chose soon. However, I'd like to post up some of my storage benchmark data as that's my primary concern right now.

I know my design is not highly optimized for throughput or IOPS, as was explained to me by forum members above, but I needed a decently inexpensive mix of speed, redundancy, and capacity. However, I do want to be sure, before I put this system into heavier use, that no one sees any significant under-performance given the system configuration I chose.

Obviously, benchmarking hardware RAID5 is known to be a challenge. I welcome input on my methodology. I tried many benchmarks but the only one that game me reasonably stable results from one run to another was HD Tune Pro (trial). Other benchmarks I tried were ATTO (Windows), iometer (from both Linux and Windows), and diskbench (Linux). diskbench might actually be the most useful but the flexibility in testing options and interpreting the results were over my head. I'll be happy to post up any results people feel would be beneficial.

My current setup is 4x1TB RAID5 with a ~750GB LUN/datastore for my VMs and a 2TB LUN upon which I setup a raw device mapping in ESXi for these tests. Ultimately, that LUN will also be an ESXi datastore used for data storage. All hardware used is as proposed in the original post with the exception of the chassis.

I ran each read and write test twice for each of two configurations:
1) read-ahead on, write-behind forced on
2) read-ahead off, write-behind off

Note: The image links to a gallery on imgur.com with all 8 benchmarks/images.



Note: The image links to a gallery on imgur.com with all 8 benchmarks/images.

(Disregard the read drop in the middle of the first test, that was my fault for accidentally launching MegaRAID Storage Manager during the test.)

Read-ahead had no significant effect on read performance for this particular benchmark. But, as you can see and is probably expected, the write-behind made for a massive difference. Indeed, I'm a little concerned about the performance without write-behind.

I'd appreciate input on how I can improve my testing methodology and/or regarding my results. Thank you.
 
Last edited:
Your write speeds look terrible. Although I do not trust HD Tune as a reliable benchmark. There was someone else posting in this forum recently who had poor write speeds with an LSI card.

Since you can run linux, the best benchmark to start with is simply sequential read and write speed with dd. For example:

# dd if=/dev/zero of=/mnt/raidvol/testfile bs=1M count=8000 oflag=direct

If the direct IO does not work, you can try

# dd if=/dev/zero of=/mnt/raidvol/testfile bs=1M count=8000 conv=fdatasync,notrunc

Adjust the count as necessary. For a sequential read test, you can try:

# dd if=/mnt/raidvol/testfile of=/dev/null bs=1M count=8000 iflag=direct

For more detailed benchmarking, I like iozone. Here are a couple tests you could try with iozone:

# iozone -a -y4K -q16M -s32M -I -e -f /mnt/raidvol/iotest -i0 -i1 -i2

This runs relatively quickly, but you probably should try a much larger file size after verifying that the above command works. Something like this perhaps:

# iozone -a -y64K -q16M -s8G -e -f /mnt/raidvol/iotest -i0 -i1 -i2

Or you can try this first, to get an idea of what file size to use:

# iozone -a -n1G -g32G -r512K -e -f /mnt/raidvol/iotest -i0 -i1
 
Last edited:
Hi John,

I ran your tests mostly as suggested. I did 16GB instead of 8GB. While dd w/ 'direct' worked, I ran a few with 'fdatasync,notrunc' as well. I haven't had the time to investigate the parameters for, and how to interpret the results of, iozone, but I intend to do so. So for the moment I just ran them with the parameters close to what you recommended and posted them up. These are all with read-ahead on and write-back forced on. Unlike my HD Tune tests, these are actually on a temporary 64GB VMDK I created on the VM datastore. Note that I had read the thread on the slow LSI card back before I even bought the parts, but I got the impression the person had just not forced write-back on.
Code:
dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=16000 oflag=direct
16777216000 bytes (17 GB) copied, 65.281 s, 257 MB/s
16777216000 bytes (17 GB) copied, 42.0595 s, 399 MB/s
16777216000 bytes (17 GB) copied, 40.1872 s, 417 MB/s
16777216000 bytes (17 GB) copied, 64.3609 s, 261 MB/s
16777216000 bytes (17 GB) copied, 45.5317 s, 368 MB/s
16777216000 bytes (17 GB) copied, 91.7425 s, 183 MB/s
16777216000 bytes (17 GB) copied, 39.4231 s, 426 MB/s
16777216000 bytes (17 GB) copied, 60.3335 s, 278 MB/s
16777216000 bytes (17 GB) copied, 53.066 s, 316 MB/s
16777216000 bytes (17 GB) copied, 54.8782 s, 306 MB/s
16777216000 bytes (17 GB) copied, 50.8727 s, 330 MB/s
Code:
dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=16000 conv=fdatasync,notrunc
16777216000 bytes (17 GB) copied, 74.0241 s, 227 MB/s
16777216000 bytes (17 GB) copied, 40.1471 s, 418 MB/s
16777216000 bytes (17 GB) copied, 41.5535 s, 404 MB/s
16777216000 bytes (17 GB) copied, 52.6267 s, 319 MB/s
16777216000 bytes (17 GB) copied, 47.1057 s, 356 MB/s
Code:
dd if=/media/testdrive/testfile of=/dev/null bs=1M count=16000 iflag=direct
16777216000 bytes (17 GB) copied, 39.7968 s, 422 MB/s
16777216000 bytes (17 GB) copied, 39.8331 s, 421 MB/s
16777216000 bytes (17 GB) copied, 40.1948 s, 417 MB/s
16777216000 bytes (17 GB) copied, 39.9849 s, 420 MB/s
16777216000 bytes (17 GB) copied, 40.1793 s, 418 MB/s
Code:
iozone -a -y4K -q16M -s32M -I -e -f /media/testdrive/testfile -i0 -i1 -i2

	Auto Mode
	Using Minimum Record Size 4 KB
	Using Maximum Record Size 16384 KB
	File size set to 32768 KB
	O_DIRECT feature enabled
	Include fsync in write timing
	Command line used: iozone -a -y4K -q16M -s32M -I -e -f /media/testdrive/testfile -i0 -i1 -i2
	Output is in Kbytes/sec
	Time Resolution = 0.000001 seconds.
	Processor cache size set to 1024 Kbytes.
	Processor cache line size set to 32 bytes.
	File stride size set to 17 * record size.
                                                            random  random    bkwd   record   stride                                   
              KB  reclen   write rewrite    read    reread    read   write    read  rewrite     read   fwrite frewrite   fread  freread
           32768       4   33486   36821    39474    39364   39808   36782                                                          
           32768       8   64280   66914    68254    68172   70182   67951                                                          
           32768      16  121467  121519   121307   128602  131264  127442                                                          
           32768      32  219216  205099   218322   228554  239779  227827                                                          
           32768      64  392177  341867   363579   365693  380695  373811                                                          
           32768     128  550650  470026   467840   541279  560712  573259                                                          
           32768     256  785200  689315   656619   701205  736129  783137                                                          
           32768     512  945222  799199   825891   883018  899576  948176                                                          
           32768    1024  960072  952863   833130  1106800 1103893 1145536                                                          
           32768    2048 1185566 1088601  1152868  1233324 1233634  935507                                                          
           32768    4096  948916  827826  1219091  1313611 1310929 1243894                                                          
           32768    8192 1335619 1100446   972084  1355790 1349202 1367499                                                          
           32768   16384 1228462 1123961  1199158  1364255 1360690 1223693
Code:
iozone -a -n1G -g32G -r512K -e -f /media/testdrive/testfile -i0 -i1 

	Auto Mode
	Using minimum file size of 1048576 kilobytes.
	Using maximum file size of 33554432 kilobytes.
	Record Size 512 KB
	Include fsync in write timing
	Command line used: iozone -a -n1G -g32G -r512K -e -f /media/testdrive/testfile -i0 -i1
	Output is in Kbytes/sec
	Time Resolution = 0.000001 seconds.
	Processor cache size set to 1024 Kbytes.
	Processor cache line size set to 32 bytes.
	File stride size set to 17 * record size.
                                                            random  random    bkwd   record   stride                                   
              KB  reclen   write rewrite    read    reread    read   write    read  rewrite     read   fwrite frewrite   fread  freread
         1048576     512  276018  354545   382922   379110                                                                          
         2097152     512  376264  370147   405147   371259                                                                          
         4194304     512  402496  403042   405474   390244                                                                          
         8388608     512  379867  174339   385288   396830                                                                          
        16777216     512  326169  401262   396298   394625                                                                          
        33554432     512  367951  127826   389086   388400
Code:
iozone -a -y64K -q16M -s16G -e -f /media/testdrive/testfile -i0 -i1 -i2

	Auto Mode
	Using Minimum Record Size 64 KB
	Using Maximum Record Size 16384 KB
	File size set to 16777216 KB
	Include fsync in write timing
	Command line used: iozone -a -y64K -q16M -s16G -e -f /media/testdrive/testfile -i0 -i1 -i2
	Output is in Kbytes/sec
	Time Resolution = 0.000001 seconds.
	Processor cache size set to 1024 Kbytes.
	Processor cache line size set to 32 bytes.
	File stride size set to 17 * record size.
                                                            random  random    bkwd   record   stride                                   
              KB  reclen   write rewrite    read    reread    read   write    read  rewrite     read   fwrite frewrite   fread  freread
        16777216      64  370788  221858   392284   391353    7120   19435                                                          
        16777216     128   92761  397263   403302   401328   11455   24932                                                          
        16777216     256  111255  375190   389480   405394   13311   33280                                                          
        16777216     512   83752  274814   371673   390371   15917   32062                                                          
        16777216    1024  178436  288805   407234   406706   30437   67594                                                          
        16777216    2048   76730  349773   405242   405454   53325   70056                                                          
        16777216    4096   78085  343188   407290   406065   81607   62108                                                          
        16777216    8192   83278  138820   395007   402984  131457   61081                                                          
        16777216   16384  137805   71652   403575   408249  187365   71182
Thank you.
 
Last edited:
I'm back with some more benchmark results. First off, for anyone following the thread, the post above this one would not have triggered this thread as having new posts since I edited an existing mostly-empty post a couple days back. So you may want to resume reading from there if you are interested. It contains the results of tests recommended by john4200.

In addition, I've now run some ATTO benchmarks on the same VM as I did the original HD Tune benchmarks a few posts above. The testing methodology is the same. The first three are with read-ahead ON and write-behind FORCED ON, the latter three have both OFF.

Note: The image links to a gallery on imgur.com with all 6 benchmarks/images.



Note: The image links to a gallery on imgur.com with all 6 benchmarks/images.

(I also accidentally ran some benchmarks using the much older ATTO 2.41 w/ the max total length of 256MB which are likely deprecated by those above which are with ATTO 2.47 w/ total length of 2GB.)

Also, I'm trying to run more dd and iozone benchmarks like those in the previous post except booting to Linux directly as opposed to running from an ESXi VM to see if there is any difference. However, I'm getting very low read speeds when I do so: around 100 MB/s. Can I assume that's because the RAID drivers included in the distribution are inferior to those provided by VMWare that I use on the VM?

I'd appreciate any input on improving my testing and in interpreting the results. I believe I have at least some reason for concern that my hardware is under-performing. Certainly while writes without write-back should be slower, should they be that much slower? I find it hard to believe they should. And if they really are that slow should I worry about my performance in real-world use, as opposed to benchmarking, even with write-back on? Thank you.
 
Last edited:
Hi John,

I ran your tests mostly as suggested. I did 16GB instead of 8GB. While dd w/ 'direct' worked, I ran a few with 'fdatasync,notrunc' as well. I haven't had the time to investigate the parameters for, and how to interpret the results of, iozone, but I intend to do so. So for the moment I just ran them with the parameters close to what you recommended and posted them up. These are all with read-ahead on and write-back forced on. Unlike my HD Tune tests, these are actually on a temporary 64GB VMDK I created on the VM datastore. Note that I had read the thread on the slow LSI card back before I even bought the parts, but I got the impression the person had just not forced write-back on.
Code:
dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=16000 oflag=direct
[/quote]

I looked over the results, and the read performance looks okay to me, both with dd and with iozone.

The dd results for sequential write speed seem to fluctuate wildly on successive trials. The maximum speeds look about right, but the results in the 100's and 200's look like something is wrong.

For the first iozone test with the 32MB file size, note that the purpose of that test was mostly to see if everything is working properly. Since the file size is only 32MB, most or all of the test was cached, and you ended up with enormous speeds (apparently the driver did not respect the direct I/O flag, which is not uncommon for RAID cards, I think).

For the iozone test with 1GB to 32GB file size (512KB record size), again the maximum write speeds look okay, but some of the speeds drop a lot, particularly a couple of the rewrite speeds, which are in the 100's.

For the iozone test with 16GB file size, record size from 64KB to 16MB, the write speeds are fluctuating wildly with record size, and some are even less than 100 MB/s. That is bad. 

I'm afraid I have no experience with either ESXi, or with configuring RAID on an LSI card, so I cannot help with that. But it certainly looks like something is wrong, judging from the write speeds in those tests.
 
Also, I'm trying to run more dd and iozone benchmarks like those in the previous post except booting to Linux directly as opposed to running from an ESXi VM to see if there is any difference. However, I'm getting very low read speeds when I do so: around 100 MB/s. Can I assume that's because the RAID drivers included in the distribution are inferior to those provided by VMWare that I use on the VM?

What RAID driver are you using, and what version? And what kernel?

If you don't know, you can try lsmod and look for something like megaraid. Once you know the module, you can try "modinfo [kernel module name]". Or you can look through the boot logs (or dmesg) to try to find it.
 
Thanks again for your continued and gracious assistance.

I've now set up 3 Linux environments (outside of ESXi so direct to the RAID controller) for testing this problem. All three are currently using the default drivers from the distributions. Below is the information you requested on each. I ran two dd write and two dd read tests on each environment all with read-ahead on and write-back forced on.

Debian 5.0.8
2.6.26-2-amd64
ext3
Write: 82 MB/s, 81 MB/s
Read: 219 MB/s, 181 MB/s
Code:
filename:       /lib/modules/2.6.26-2-amd64/kernel/drivers/scsi/megaraid/megaraid_sas.ko
description:    LSI MegaRAID SAS Driver
author:         megaraidlinux@lsi.com
version:        00.00.04.01
license:        GPL
srcversion:     49BEAEC53F4BE8F4646C64A
alias:          pci:v00001028d00000015sv*sd*bc*sc*i*
alias:          pci:v00001000d00000413sv*sd*bc*sc*i*
alias:          pci:v00001000d00000079sv*sd*bc*sc*i*
alias:          pci:v00001000d00000078sv*sd*bc*sc*i*
alias:          pci:v00001000d0000007Csv*sd*bc*sc*i*
alias:          pci:v00001000d00000060sv*sd*bc*sc*i*
alias:          pci:v00001000d00000411sv*sd*bc*sc*i*
depends:        scsi_mod
vermagic:       2.6.26-2-amd64 SMP mod_unload modversions
parm:           poll_mode_io:Complete cmds from IO path, (default=0) (int)
Ubuntu Server 10.04 LTS
2.6.32-32-generic
ext4
Write: 251 MB/s, 302 MB/s
Read: 112 MB/s, 82 MB/s
Code:
filename:       /lib/modules/2.6.32-32-generic/kernel/drivers/scsi/megaraid/megaraid_sas.ko
description:    LSI MegaRAID SAS Driver
author:         megaraidlinux@lsi.com
version:        00.00.04.01
license:        GPL
srcversion:     1AB2B4AC6534AB333DF4ACA
alias:          pci:v00001028d00000015sv*sd*bc*sc*i*
alias:          pci:v00001000d00000413sv*sd*bc*sc*i*
alias:          pci:v00001000d00000071sv*sd*bc*sc*i*
alias:          pci:v00001000d00000073sv*sd*bc*sc*i*
alias:          pci:v00001000d00000079sv*sd*bc*sc*i*
alias:          pci:v00001000d00000078sv*sd*bc*sc*i*
alias:          pci:v00001000d0000007Csv*sd*bc*sc*i*
alias:          pci:v00001000d00000060sv*sd*bc*sc*i*
alias:          pci:v00001000d00000411sv*sd*bc*sc*i*
depends:
vermagic:       2.6.32-32-generic SMP mod_unload modversions
parm:           poll_mode_io:Complete cmds from IO path, (default=0) (int)
Ubuntu Desktop 11.04
2.6.38-8-generic
ext4
Write: 236 MB/s, 319 MB/s
Read: 118 MB/s, 82 MB/s
Code:
filename:       /lib/modules/2.6.38-8-generic/kernel/drivers/scsi/megaraid/megaraid_sas.ko
description:    LSI MegaRAID SAS Driver
author:         megaraidlinux@lsi.com
version:        00.00.05.29-rc1
license:        GPL
srcversion:     650A1F5CFCD441F98FC6EBF
alias:          pci:v00001000d0000005Bsv*sd*bc*sc*i*
alias:          pci:v00001028d00000015sv*sd*bc*sc*i*
alias:          pci:v00001000d00000413sv*sd*bc*sc*i*
alias:          pci:v00001000d00000071sv*sd*bc*sc*i*
alias:          pci:v00001000d00000073sv*sd*bc*sc*i*
alias:          pci:v00001000d00000079sv*sd*bc*sc*i*
alias:          pci:v00001000d00000078sv*sd*bc*sc*i*
alias:          pci:v00001000d0000007Csv*sd*bc*sc*i*
alias:          pci:v00001000d00000060sv*sd*bc*sc*i*
alias:          pci:v00001000d00000411sv*sd*bc*sc*i*
depends:
vermagic:       2.6.38-8-generic SMP mod_unload modversions
parm:           poll_mode_io:Complete cmds from IO path, (default=0) (int)
parm:           max_sectors:Maximum number of sectors per IO command (int)
parm:           msix_disable:Disable MSI-X interrupt handling. Default: 0 (int)

Strangely, the reads are faster than the writes on Debian and the opposite is true on the two Ubuntu systems. And obviously none are performing as well as under ESXi. I picked the former two because they are listed as supported on the LSI driver releases, whereas the newer Ubuntu 11.04 is not (but which modinfo indicates has a newer driver).

I suspect my next step should be to see about compiling updated drivers for one or more of these environments and retesting? LSI's latest driver version for MegaRAID is 5.30. I've also opened a ticket with LSI but since I don't expect that to bear fruit I'm going to continue to try to diagnose this.
 
Those results are really bizarre. Ideally, you need to talk to someone at LSI who has experience using their RAID cards in linux.

By the way, the latest megaraid_sas dirver is 05.34-rc1. I am running it now (but only using my card in pass-thru mode). I did not need to compile it myself. I am running Archlinux, and the latest kernel, 2.6.39, came with it.

Code:
# uname -srv
Linux 2.6.39-ARCH #1 SMP PREEMPT Mon Jun 6 22:37:55 CEST 2011

# modinfo megaraid_sas
filename:       /lib/modules/2.6.39-ARCH/kernel/drivers/scsi/megaraid/megaraid_sas.ko.gz
description:    LSI MegaRAID SAS Driver
author:         megaraidlinux@lsi.com
version:        00.00.05.34-rc1
 
Ah! That'll be easier then. I'll throw the latest Archlinux on a partition and run the tests again. I'll report back with the results. Thanks again.

Edit: Or, probably even easier, I'll upgrade the Natty environment to 2.6.39.
 
I compiled and installed the 2.6.39.2 kernel for my Natty test system. The results show the same strange and distressing benchmarking patterns.
Code:
Linux 2.6.39.2 #2 SMP Sun Jun 26 02:29:41 EDT 2011

filename:       /lib/modules/2.6.39.2/kernel/drivers/scsi/megaraid/megaraid_sas.ko
description:    LSI MegaRAID SAS Driver
author:         megaraidlinux@lsi.com
version:        00.00.05.34-rc1
parm:           poll_mode_io:Complete cmds from IO path, (default=0) (int)
parm:           max_sectors:Maximum number of sectors per IO command (int)
parm:           msix_disable:Disable MSI-X interrupt handling. Default: 0 (int)

Code:
sudo dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=16000 oflag=direct
16777216000 bytes (17 GB) copied, 43.2148 s, 388 MB/s
16777216000 bytes (17 GB) copied, 55.7191 s, 301 MB/s
16777216000 bytes (17 GB) copied, 50.7508 s, 331 MB/s
16777216000 bytes (17 GB) copied, 42.9592 s, 391 MB/s
16777216000 bytes (17 GB) copied, 55.2481 s, 304 MB/s

sudo dd if=/media/testdrive/testfile of=/dev/null bs=1M count=16000 iflag=direct
16777216000 bytes (17 GB) copied, 131.599 s, 127 MB/s
16777216000 bytes (17 GB) copied, 197.016 s, 85.2 MB/s
16777216000 bytes (17 GB) copied, 131.65 s, 127 MB/s
16777216000 bytes (17 GB) copied, 109.959 s, 153 MB/s
16777216000 bytes (17 GB) copied, 150.353 s, 112 MB/s

Code:
sudo iozone -a -n1G -g32G -r512K -e -f /media/testdrive/testfile -i0 -i1
                                                            random  random    bkwd   record   stride
              KB  reclen   write rewrite    read    reread    read   write    read  rewrite     read   fwrite frewrite   fread  freread
         1048576     512  236692  180278 10690911 11106996
         2097152     512  309587  221234 10808616 10958107
         4194304     512  364403  131020 10713232 10797264
         8388608     512  358881  356811 10984410 11078063
        16777216     512  285312  182437   609378   625159
        33554432     512  318708  121747   379416   383181

Code:
sudo iozone -a -y64K -q16M -s16G -e -f /media/testdrive/testfile -i0 -i1 -i2
                                                            random  random    bkwd   record   stride
              KB  reclen   write rewrite    read    reread    read   write    read  rewrite     read   fwrite frewrite   fread  freread
        16777216      64  294510  161913   639922   655900   44810   17284
        16777216     128  304198  243102   663628   626907   62647   24969
        16777216     256  317265   71776   656590   649135   83424   38145
        16777216     512  131866  166327   603435   601439   92524   41362
        16777216    1024  221776  178563   334774   374351   71081   70161
        16777216    2048  218260  333600   249897   173169  179712   74613
        16777216    4096  152066  331113   307392   342866  173648   69582
        16777216    8192  174332  379355   494750   513194  573121   70571
        16777216   16384  269183  146856   616227   614224  738686   70669
So I've got the ticket in with LSI and will report back with what they say. Are there any bases I haven't covered in my investigation? Do we all feel pretty confident that there is something wrong?

Also, could these specific symptoms likely be caused by one or more of the HDs themselves? As I have no other SAS controllers and the 9261 cannot do JBOD, I wasn't able to test them individually. I suppose I could do single-drive RAID0 but it would be a significant project to redo everything I've done on the current RAID setup so I'd prefer not to wipe everything unless there was significant feeling that a drive could be a culprit. Thanks again.
 
So I've got the ticket in with LSI and will report back with what they say. Are there any bases I haven't covered in my investigation? Do we all feel pretty confident that there is something wrong?

Also, could these specific symptoms likely be caused by one or more of the HDs themselves? As I have no other SAS controllers and the 9261 cannot do JBOD, I wasn't able to test them individually. I suppose I could do single-drive RAID0 but it would be a significant project to redo everything I've done on the current RAID setup so I'd prefer not to wipe everything unless there was significant feeling that a drive could be a culprit. Thanks again.

It is interesting that your sequential writes with dd look reasonable now, but the dd sequential reads are low and variable. I wonder if the direct IO flag could be disagreeable to the LSI driver on reads?

You could try it without the direct flag. You would need to use a file size several times larger than your available RAM. Alternatively, you can clear the buffer-cache before the read test:

# dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=16000
# sync; echo 3 > /proc/sys/vm/drop_caches
# dd if=/media/testdrive/testfile of=/dev/null bs=1M count=16000

As for testing individual drives, that sounds like a good idea. I'm surprised the LSI 9261 cannot do pass-thru mode. I have three IBM ServeRAID M1015 cards that I have flashed with LSI 9240-8i firmware, and they automatically pass thru the drives to linux. I did not even have to configure the card to do that. Just plugged the drives in. Presumably, since I did not configure a RAID with the drives, the card just automatically passed them thru.

If the 9261 cannot do pass-thru, then what about plugging the drives into motherboard SATA. At least you could verify if any of the drives are faulty.

One other thought: have you tried swapping cables, and are you sure the drives are all getting enough power? When I get weird, intermittent results, my first troubleshooting steps are: 1) cables? and 2) power?
 
Last edited:
Hi John. It looks like your suspicion was correct regarding direct I/O and reads. Here are the results of your suggested test:
Code:
sudo dd if=/dev/zero of=/media/testdrive/testfile bs=1M count=96000
sudo sync; sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'
sudo dd if=/media/testdrive/testfile of=/dev/null bs=1M count=96000
sudo sync; sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

write: 100663296000 bytes (101 GB) copied, 333.265 s, 302 MB/s
read:  100663296000 bytes (101 GB) copied, 257.177 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 267.73 s, 376 MB/s
read:  100663296000 bytes (101 GB) copied, 257.191 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 285.936 s, 352 MB/s
read:  100663296000 bytes (101 GB) copied, 257.163 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 282.027 s, 357 MB/s
read:  100663296000 bytes (101 GB) copied, 257.448 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 259.887 s, 387 MB/s
read:  100663296000 bytes (101 GB) copied, 257.444 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 364.817 s, 276 MB/s
read:  100663296000 bytes (101 GB) copied, 257.432 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 485.989 s, 207 MB/s
read:  100663296000 bytes (101 GB) copied, 257.4 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 312.988 s, 322 MB/s
read:  100663296000 bytes (101 GB) copied, 257.555 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 370.459 s, 272 MB/s
read:  100663296000 bytes (101 GB) copied, 257.084 s, 392 MB/s

write: 100663296000 bytes (101 GB) copied, 445.236 s, 226 MB/s
read:  100663296000 bytes (101 GB) copied, 257.484 s, 391 MB/s

write: 100663296000 bytes (101 GB) copied, 278.636 s, 361 MB/s
read:  100663296000 bytes (101 GB) copied, 257.047 s, 392 MB/s
So reads could not be more stable now. Writes continue to fluctuate some.

I'll keep wiping the RAID and testing the drives as single-drive RAID0 as an option of last resort. I'll wait and see if LSI has anything to add before I go that far. With regard to plugging the HDs into MB SATA, it is my understanding that while SATA drives can plug into SAS controllers, the opposite is not true (and not just because of the connector "key"). I regret that the LSI has no pass-through, as I would have tested the drives as JBOD first thing after completing the build. But since it did not, I opted to go straight to the RAID5 testing.

Cabling and power are valid points. Unfortunately, since I dont have a SAS backplane in this chassis, I had to buy an expensive cable for connecting the drives, and I only currently have the one. I should _technically_ have enough power, using a well recommended 550W PS. I will try to find the time to monitor the voltages during a test but I don't have a convenient, good quality PS with which to swap without scavenging it from one of my active systems.

Thanks again.
 
Ah, I forgot that you had SAS disks.

One other thing you might want to try is writing directly to the block device, bypassing any filesystem. Something like:

# dd if=/dev/zero of=/dev/sd_ bs=1M count=16000 conv=fdatasync,notrunc

Alternatively, without the fdatasync, but larger write:

# dd if=/dev/zero of=/dev/sd_ bs=1M count=96000

And just for the heck of it, you could try:

# dd if=/dev/zero of=/dev/sd_ bs=1M count=16000 oflag=dsync

Obviously these will wipe out any partitioning and/or filesystem you might have on the device.

BTW, what filesystem did you have on your previous tests?
 
Last edited:
Hi John. Regarding your question, all Linux tests were performed on ext4, with the exception of the Debian Lenny tests which were on ext3.

I think the results of the latest batch of your suggested tests are interesting:
Code:
sudo sudo dd if=/dev/zero of=/dev/sdb bs=1M count=16000 conv=fdatasync,notrunc

16777216000 bytes (17 GB) copied, 41.472 s, 405 MB/s
16777216000 bytes (17 GB) copied, 43.299 s, 387 MB/s
16777216000 bytes (17 GB) copied, 41.7945 s, 401 MB/s
16777216000 bytes (17 GB) copied, 42.1298 s, 398 MB/s
16777216000 bytes (17 GB) copied, 42.3096 s, 397 MB/s

sudo dd if=/dev/zero of=/dev/sdb bs=1M count=96000

100663296000 bytes (101 GB) copied, 247.89 s, 406 MB/s
100663296000 bytes (101 GB) copied, 247.934 s, 406 MB/s
100663296000 bytes (101 GB) copied, 247.587 s, 407 MB/s
100663296000 bytes (101 GB) copied, 247.586 s, 407 MB/s
100663296000 bytes (101 GB) copied, 247.485 s, 407 MB/s

sudo dd if=/dev/zero of=/dev/sdb bs=1M count=16000 oflag=dsync

16777216000 bytes (17 GB) copied, 41.2639 s, 407 MB/s
16777216000 bytes (17 GB) copied, 41.212 s, 407 MB/s
16777216000 bytes (17 GB) copied, 41.3661 s, 406 MB/s
16777216000 bytes (17 GB) copied, 41.266 s, 407 MB/s
16777216000 bytes (17 GB) copied, 41.1869 s, 407 MB/s
I'm curious as to your interpretation of the significance of the speed fluctuations on the filesystem tests versus the splendid stability of these tests. (i.e. Should I be worried about something else?) Additionally, should I take this as a positive sign regarding my storage hardware? (I'm desperate to have a sigh of relief at some point. :p) And, finally, is there a way to do random read/write tests along these lines, directly to the block device? Thank you!
 
I'm curious as to your interpretation of the significance of the speed fluctuations on the filesystem tests versus the splendid stability of these tests. (i.e. Should I be worried about something else?) Additionally, should I take this as a positive sign regarding my storage hardware?

Yes, based on those results, it looks to me like you are getting reasonable performance out of your RAID card.

The earlier low speeds and fluctuations are probably due to either a poorly configured filesystem, or perhaps ext4 is just not well suited to RAID?

I do not use ext4 on striped RAID volumes. But I see that the mke2fs man page lists a couple options that should be set properly for a hardware RAID (for mdadm software RAID, I think they are automatically set):

Code:
-E extended-options
       Set extended options for the filesystem.  Extended options are
       comma separated, and may take an argument using the equals ('=')
       sign.  The -E option used to be -R in earlier versions of mke2fs.
       The -R option is still accepted for backwards compatibility.  The
       following extended options are supported:

       stride=stride-size
              Configure the filesystem for a RAID array with stride-size
              filesystem blocks. This is the number of blocks read or
              written to disk before moving to the next disk, which is
              sometimes referred to as the chunk size.  This mostly
              affects placement of filesystem metadata like bitmaps at
              mke2fs time to avoid placing them on a single disk, which
              can hurt performance.  It may also be used by the block
              allocator.

        stripe-width=stripe-width
               Configure the filesystem for a RAID array with
               stripe-width filesystem blocks per stripe.  This is
               typically stride-size * N, where N is the number of
               data-bearing disks in the RAID (e.g. for RAID 5 there
               is one parity disk, so N will be the number of disks in
               the array minus 1).  This allows the block allocator to
               prevent read-modify-write of the parity in a RAID stripe
               if possible when the data is written.

So, if you had a 4 disk RAID-0 with 512KB chunk size, I think you'd use a command like this assuming you want 4KiB filesystem blocks :

# mkfs.ext4 -b 4096 -E stride=128,stripe-width=512 /dev/sd_

If that still gives inconsistent write speeds, you might try mounting the filesystem without barriers.

Personally, I use XFS with striped RAID volumes. For a 512KB chunk 4-drive RAID-0, I would use:

# mkfs.xfs -d su=512k,sw=4 /dev/sd_

Again, you can try mounting without barriers for benchmarking, but for safety I generally avoid that option in production. Note that sw should be equal to the number of DATA drives. So it would be N for RAID-0, N-1 for RAID-5, and N-2 for RAID-6.
 
Last edited:
Very interesting. When it comes time to build the filesystems again, I'll put that information to good use!

However, I have some new data and some speculation to go with it. First the very distressing data. These are the same tests as yesterday but with write-behind disabled:
Code:
sudo dd if=/dev/zero of=/dev/sdb bs=1M count=16000 conv=fdatasync,notrunc

16777216000 bytes (17 GB) copied, 416.305 s, 40.3 MB/s
16777216000 bytes (17 GB) copied, 312.097 s, 53.8 MB/s
16777216000 bytes (17 GB) copied, 425.731 s, 39.4 MB/s

sudo dd if=/dev/zero of=/dev/sdb bs=1M count=96000

100663296000 bytes (101 GB) copied, 2240.52 s, 44.9 MB/s

sudo dd if=/dev/zero of=/dev/sdb bs=1M count=16000 oflag=dsync

16777216000 bytes (17 GB) copied, 605.773 s, 27.7 MB/s
16777216000 bytes (17 GB) copied, 584.235 s, 28.7 MB/s
16777216000 bytes (17 GB) copied, 590.544 s, 28.4 MB/s
So I'm seeing over a 10:1 difference between write-back and write-through. Please correct me if I'm wrong but that seems very wrong to me. Additionally, my speculation is this: could it be that the fluctuations we see with write-back on are caused by good performance when the cache isn't "full" but fluctuates more severely than is typical when it is because of this horrible write-through performance? It seems like this HD Tune benchmark result may be demonstrating my suspicion.
 
So I'm seeing over a 10:1 difference between write-back and write-through. Please correct me if I'm wrong but that seems very wrong to me. Additionally, my speculation is this: could it be that the fluctuations we see with write-back on are caused by good performance when the cache isn't "full" but fluctuates more severely than is typical when it is because of this horrible write-through performance? It seems like this HD Tune benchmark result may be demonstrating my suspicion.

I have tested Areca RAID cards in linux, and in write-through mode, they certainly do better than that! In fact, on large sequential writes, the speeds were generally proportional to the number of data drives in the RAID, minus a little overhead. This is what I would expect for any decent striped RAID controller. It is only for small, random writes where the cache should help (writes of less than a full stripe).

I vaguely recall that in my testing on the Areca, the individual drives needed to have NCQ enabled in order to get good performance when the Areca was in write-through. The Areca cards have a setting where you can change NCQ for the individual drives. I have no idea if NCQ is important for LSI RAID cards, and whether you can change it. And you have SAS drives, so I guess they have TCQ instead of NCQ?
 
I seem to recall that ext4 implements a write-barrier by default, which can screw write performance to the wall...
 
I seem to recall that ext4 implements a write-barrier by default, which can screw write performance to the wall...

Which possibility is why I suggested trying a benchmark without barriers in post #32.

For ext4, the mount option is: barrier=0

For XFS, the mount option is: nobarrier

Actually, I think nobarrier is also acceptable for ext4, but it is not the preferred method.

But people should keep in mind that barriers are generally helpful for data safety, so while it is okay to turn them off temporarily for benchmarking, barriers should generally be left on for production unless you are sure you know the consequences of turning them off. Generally, it is only safe to turn off barriers if your RAID card has a battery-backed cache AND you have disabled the write-cache on the individual HDDs.
 
Which possibility is why I suggested trying a benchmark without barriers in post #32.

For ext4, the mount option is: barrier=0

For XFS, the mount option is: nobarrier

Actually, I think nobarrier is also acceptable for ext4, but it is not the preferred method.

But people should keep in mind that barriers are generally helpful for data safety, so while it is okay to turn them off temporarily for benchmarking, barriers should generally be left on for production unless you are sure you know the consequences of turning them off. Generally, it is only safe to turn off barriers if your RAID card has a battery-backed cache AND you have disabled the write-cache on the individual HDDs.

I didn't see your post, john. I tend to prefer xfs. I've found that xfs seems to disable the write barrier when I am running it in a VM, claiming that the write barrier check failed (for whatever reason.) That said, the couple of times I tried playing with ext4 I gave up - the write performance was abysmal, and I wasn't doing anything customized - a straight up install.
 
The Areca cards have a setting where you can change NCQ for the individual drives. I have no idea if NCQ is important for LSI RAID cards, and whether you can change it. And you have SAS drives, so I guess they have TCQ instead of NCQ?
It should be something so easy to look up, but I'm having a heck of a time determining that. These are nearline SAS drives so they are SATA drives with a SAS interface. But, I can find no reliable source for what my Constellation ES SAS drives have. The manual mentions a "tag queue" in two places, but never actually mentions TCQ or tagged command queuing anywhere (or NCQ for that matter). And since there are SATA variants of the same line, that makes it more difficult to search.

NCQ is enabled on the LSI, for what that's worth. There are no NCQ controls for individual drives in its various configuration tools.
Code:
sudo /opt/MegaRAID/MegaCli64 -AdpGetProp NCQDsply -a0
Adapter 0: NCQ Status is Enabled

Edit: For the heck of it I disabled NCQ on the LSI and ran one of the dd tests again. It came out 31 MB/s so (unsurprisingly) that didn't help.
 
Last edited:
There may be a distinction between NCQ being enabled on the individual drives, and NCQ/TCQ being enabled on the "device" that the RAID card and hardware driver present to the OS.

For example, on an Areca 1680ix, I have options for (just listing the relevant ones):

Code:
Sys Config
  SATA NCQ support (enabled/disabled)
  HDD read-ahead cache (enabled/disabled)
  Volume Data Read ahead (norm/agg/cons/disabled)
  HDD queue depth (1,2,4,8,16,32)
  Disk write cache mode (auto/enabled/disabled)

Volume attributes
  Tagged Command Queuing (enabled/disabled)
  Volume cache mode (write back/write through)

I think the TCQ is for the overall RAID device, and the NCQ and HDD queue depth are for the individual drives (although the NCQ enabled/disabled seems superfluous, just set the queue depth to 1 to disable, I would think). I don't know if LSI has similar settings or not.
 
Last edited:
Back
Top