• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Using Desktop Drives for RAID -- My Experience

Quentin-

n00b
Joined
Jan 18, 2010
Messages
5
The thing about a Home RAID system is that you're not an enterprise with a multi-million dollar IT budget, so the cost of the drives DOES matter. You just want a good, safe place to put all your pictures, personal files, and music. And if you're like me you think it'd be just grand if you could rip your 1000+ DVD/Blu-Ray collection and play them all from your home theater screen using something like this.

The thing is, you can build a state of the art server/raid array combo for quite a reasonable price these days. The thing that really adds up are the drives. So you buy an enterprise class raid card capable of RAID-6 (and later on RAID-60) which means that two drives can catastrophically fail without suffering data loss. To me that's a great reason to use cheaper desktop class drives (remember that while we don't want to lose data, we're not a massive-budget enterprise either).

Of course the biggest issue we run into using these type of drives in combination with an enterprise class raid card is the time limited error recovery issue (TLER aka CCTL).

So we want to pick our drives carefully. Unfortunately there is nothing so nice as WDTLER available for any drives but Western Digital; and they are actually so greedy (for you to buy their 2x priced enterprise drives, even for a simple home raid) that not only did they change firmware to prevent the utility from permanently changing the settings--they actually went out of spec and disabled S.M.A.R.T. SCT Error Recovery functionality as well. Which is enough reason for me to never consider patronizing them again.

Which brought me to Seagate vs Hitachi vs Samsung. Seagate has had some ridiculous horror stories recently after their acquisition of Maxtor (is anyone really surprised?) and as much as I have been a huge fan in the past I just couldn't bring myself to trust that their latest generation of drives meet muster. Samsung on the other hand seems to have a very small market share, and generally little information available on the internet. Which leaves Hitachi (aka IBM). But not really, because IBM drives basically sucked. But apparently after Hitachi took over IBM's had drive division things really improved. You can find posts on the internet saying "...I will never buy from this manufacturer again because the drive I ordered was DOA..." for pretty much every single drive/manufacturer out there. But for Hitachi's recent lines the level of satisfaction seems to be quite high (compared especially to Seagate). Soo...I have decided to give them a try. Besides, they have tons of detailed documentation and utilities available on their website (see below).

Therefore, I have purchased 8 Hitachi 7K2000 drives for use in a RAID-6 array on an Adaptec 5085 adapter. This should give me just shy of 11 usable terabytes (2TB really means 1.8 and RAID-6 being N-2 so 6*1.8)--a great start. Without getting into system details (which for now doesn't matter much) I will add to the discussion with what I have learned.

First and foremost, the Adaptec card cannot even finish a build/verify on the RAID-6 without one of the drives dropping (having done nothing but hook up the raid box to the adapter and attempt to create a raid).

This was somewhat surprising, as I did check each drive's S.M.A.R.T. attributes (using smartctl 5.38, from GParted Live) prior to insertion in the RAID box and everything is fine. I knew I would have to address the CCTL (aka TLER) issue; I just didn't think it would affect me right out of the gate.

So, currently this is what I'm doing and why:
1) Use the Hitachi Feature Tool to:
1a) Disable the Write Cache on each drive. This is because the RAID card should be the one managing writes; and when it performs a write, it should do so with the expectation that each drive has actually written the data right then and there. This prevents a situation where a power loss to the RAID box could result in data loss that the adapter/OS would never be aware of. Aditionally, the ERC Write Timeout value should not have to be set (this is a good thing, follow the links below on SCT Write Timeout specification as to why). And yes, the Adaptec card has an option for supposedly disabling write cache on all connected drives, but to be safe I'll force it on the drive side.
1b) Enable Power Management with value 128 (0x80). As I am planning on using this array in a home environment, allowing the drives to reach idle states will save a substantial amount of electricity, furthermore the raid card *should* be capable of spin-down as well, adding further to potential savings. Dollar-wise the savings pale in comparison to the cost of the drives, but since there is no actual need (in my application) to have them constantly running at full power, it is still a good idea.
2) Use the Hitachi Drive Fitness Test to perform an Advanced Test on each drive (this takes HOURS) to verify that each drive is truly in prime condition.

The above mentioned tools are available at the following URL: http://www.hitachigst.com/hdd/support/download.htm

Now, after having done all of this there should (theoretically) be one final step, and that is to set a Read ERC timeout. I won't finish performing advanced tests on all drives until a few days hence, so for now I will just add what information I have learned, and later update with results.

So the first tool I came across for changing SCT Error Recovery values is the famed HDAT2. It works, you can plug the drive directly into an SATA port and access it, change the value for read timeout to 70 (7000 ms/7s) and it sticks through reboot. The problem is, I haven't figured out any way to use it while the drive is plugged into the RAID box and that box is connected to the Adaptec 5085. If anyone knows how please let me know and I'll try it just to let people know.

In the meantime, since values changed via S.M.A.R.T. SCT do not keep after power has been removed (e.g. unplug the drive from the direct connection to the computer and plug it back into the RAID box) it is not a very useful tool for my purposes. After all, if you have the drives directly connected to your computer (not using a hardware RAID card) you will likely want to use ZFS, and then read/write timeout values will be a non-issue anyway.

The other issue with the HDAT2 approach is that inevitably your raid box will lose power at some point (even if you have a UPS, which I do). Reasons include moving equipment, long term power failure, etc. Lets face it, we're building the RAID because we want a highly reliable, low administration, home storage network. Having drives revert after power loss and having to remember to re-run an application doesn't meet that criteria.

Having a service on server boot check (and if necessary set) those values however does. Which is why I have extremely high hopes for this. I haven't gotten around to porting his patch to smartmontools 5.39 yet, but if It works as advertised I'll be able to have a simple script run on server boot that checks the read recovery timeout value, and if necessary sets it to 70. I will post back later with my experience once I have all my drives tested and have tried this.

For information on Error Recovery Control, and specifically the SCT commands to set timeouts see http://www.t13.org/Documents/UploadedDocuments/docs2008/D1699r6a-ATA8-ACS.pdf, section 8.3.4.

For specification information on the 7K2000 series drives (including additional specific information for SCT commands) see http://www.hitachigst.com/tech/techlib.nsf/techdocs/5F2DC3B35EA0311386257634000284AD/$file/7K2000_A7K2000_Spec_r1.0.pdf.
 
Damn. Great start for a first post.

I, too, just picked up 4 of the 2TB Hitachi drives for use in a NAS device in RAID 5. I don't have movies, etc ripped, but I'm backing up 4 computers to the device. For me, this is an extra level of safety. I already have an external HD on my desktop that stores my automated backups (documents, etc). I don't save important files on the OS drive of my desktop, either (which saved my bacon the last two weeks due to some emergency reformats, damn bad RAM).

I'm wondering if I should have done some of this testing and the like prior to building my RAID array...
 
After all, if you have the drives directly connected to your computer (not using a hardware RAID card) you will likely want to use ZFS, and then read/write timeout values will be a non-issue anyway.

Actually, TLER/ERC/CCTL can be an issue for ZFS

Having a service on server boot check (and if necessary set) those values however does. Which is why I have extremely high hopes for this. I haven't gotten around to porting his patch to smartmontools 5.39 yet, but if It works as advertised I'll be able to have a simple script run on server boot that checks the read recovery timeout value, and if necessary sets it to 70. I will post back later with my experience once I have all my drives tested and have tried this.

Thank you for this! I've been looking for a way to set TLER/ERC/CCTL through linux and not through hdat2.
 
Thank you for this! I've been looking for a way to set TLER/ERC/CCTL through linux and not through hdat2.

+1! I wasn't aware of this either.

If anyone needs it, I have built a patched version of 5.39 for Debian/amd64, seems to work fine. Probably would install fine on Ubuntu or similar distros too.

http://www.gotroot.ca/~ktims/dump/smartmontools_5.39-2_amd64.deb

Edit: Oh, and APT on my machine wants to upgrade it to the 5.39-2 from the repos even though the version number is the same, so I just put smartmontools on hold. Also interesting to note that in my array that consists of 5xWD5000AAKS, the most recently added one (couple months ago) has had the command disabled, while the others work fine.
 
Last edited:
Actually, TLER/ERC/CCTL can be an issue for ZFS

To digress I'll address that. The only problem ZFS would have with a desktop class drive is that by nature a desktop class drive may try extremely hard to read data (whereas a server class drive would give up--expecting that its in a RAID and the RAID can still find the data). This behavior is fully controllable within the ZFS configuration by the way. You should never have the "ERC timeout drop from raid" if using RAID-Z/RAID-Z2/RAID-Z3 though. That is not to say that tuning the ERC timeout values would not be beneficial, its just that the symptom that people using a hardware raid controller having a drive drop won't happen in ZFS software raid.

Anyway...I finished porting the patch for smartmontools to 5.39. I have uploaded the patch file here. Simply download the source code for the 5.39 tools, put this patch in the directory that contains the smartmontools-5.39 folder, and execute: patch -p0 < smartmontools-ercv2.patch then compile as usual. On my Adaptec 5085 card I can query the values by: smartctl -d sat -l scterc /dev/sg*. Setting the values is similar: smartctl -d sat -l scterc,70,100 /dev/sg*. I can confirm that the tools work, and can both read AND set the ERC timeout values for both read and write. This is obviously great news because it can easily be tested and if necessary set in a boot time script.

Meanwhile, I have now put 4 of the 8 drives I purchased through the nearly 6 hour fitness test with full success so far. In the meantime it occurred to me that I really shouldn't be having a TLER/CCTL related issue to a failed build/verify of a RAID-6 array with brand new drives, so I took some time to narrow down the issue. It turns out I have a bad lane on port 0 of the raid card. I won't bother going through every step of the testing I did to arrive at that realization, needless to say a faulty ~$750 raid card was the last thing I expected. Hurray for Adaptec though, a 10 minute call to their support and they are shipping me another card via Fedex and in the meantime I can still use the faulty card (only one lane out of 8 remember) for testing and setup purposes.

@Wipeout: It really depends on your need, but full testing of drives prior to RAID creation, and stress testing/degraded state testing of the RAID after creation is a really good idea. After all, you built the raid because you want a safe place to store files. Spending a bit of time early on causing degraded states (removing a drive, etc.), and stress testing is a great way to make sure you actually have what you intended to build.

Furthermore, when using a hardware raid card like many in these forums with Desktop drives, its a good idea to be aware of differences in these cheaper desktop drives and their enterprise equivalents (at least insofar as we can change settings). You may find (unfortunately) that your NAS box (QNAP, etc) has a list of acceptable drives and rarely will you find desktop drives on the list for this very reason. Which is why I am attempting to (through pre-setup/boot scripts) to overcome the issues which can be overcome. I'll purposefully stray away from the question of whether a given vendors' enterprise drives are actually physically superior in any way or if the difference is mostly firmware.

Anyway, that's it for now, I'll update further as I get further along in my build.
 
I've had a lot of luck with LSI RAID cards under linux. This is the first I've heard of ERC (etc), I'll at least check how well the MegaRAID cards support it.
 
Adaptec does list both the 7K2000 and A7K2000 (with JKAOA20N firmware) as compatible with their 5-series SAS RAID cards, so once you get your replacement, you would hope it would function without issues.

Since the 7K2000 drives do pull a lot of power under load, you may want to make sure drives aren't dropping out because of voltage drops (or lack of amperage). If you have a multi-rail PSU <850W, it wouldn't be a bad idea to split the load between different rails since all the drives combined will put a peak load of 192W (16A) on the 12v rail at startup but more like 120W (10A) avg under heavy usage which could be too much for a single rail, especially if other things are on it as well.

For example, in order for me to get 14 drives running on my 650W PSU (12v @ 42A avg / 45A peak) with four 12v rails (8-18A but no more than 10.5A avg), I needed to balance the rails almost perfectly or I would have issues of drives at the end of the daisy-chain dropping-out (power-cycling) when all the drives on a particular rail were under heavy load (e.g. initializing/rebuilding RAID array, defragmenting).

Of course if you have a single rail PSU with decent wattage, you shouldn't be having any power issues like this, assuming it has decent voltage regulation.
 
Last edited:
Adaptec does list both the 7K2000 and A7K2000 (with JKAOA20N firmware) as compatible with their 5-series SAS RAID cards, so once you get your replacement, you would hope it would function without issues.

That would be a natural assumption, however in the case of vendor compatibility lists (such as the Adaptec one you mention) I believe they only mean you can use them, not that they are recommended. In this FAQ they are specifically recommending enterprise drives for their timeouts over desktop drives, which while they work can exhibit the TLER/CCTL problem.

Since the 7K2000 drives do pull a lot of power under load, you may want to make sure drives aren't dropping out because of voltage drops (or lack of amperage). If you have a multi-rail PSU <850W, it wouldn't be a bad idea to split the load between different rails since all the drives combined will put a peak load of 192W (16A) on the 12v rail at startup but more like 120W (10A) avg under heavy usage which could be too much for a single rail, especially if other things are on it as well.

I'm quoting this mainly for others reading, this is a very good point and people building their own enclosures should definitely be aware of this. Don't forget that a lot of PSU's have horrible efficiency, so to be on the safe side multiply whatever your PSU is by 70% and use that figure for your power calculations. A dedicated box with 15 drives and a SAS Expander should be on ~450W or so (which means that at a minimum you should expect to get 315W efficiency). If you're running a system board in the box with CPU, RAM, etc you'll have to have a larger supply to cover that as well.

Even though I have a dedicated 15A circuit for my "network" closet, and the RAID enclosure I'm using was professionally purpose built, this was one of the first things I checked into.

Additionally, for those with cards that support it and large numbers of drives, enable staggered spin-up on your drives if you are using power saving low-spin or spin-down settings.
 
...~450W or so (which means that at a minimum you should expect to get 315W efficiency).

I believe you've got this mixed up. The efficiency problem is on the input side rather then the output side.

For example, a 450W PSU with 75% efficiency when fully loaded should be able to output 450W, but in doing so it would be pulling 600W from the wall socket.
 
Thanks for this post, been looking to upgrade my current raid for a while now and I'm now going to highly consider the Hitachi drives.

I've been running a pretty low end raid on WD RE2's and a RocketRaid 2320 for almost 2 years now with 0 issues, but I think I've been lucky. I will probably invest in a better card like the one you're using this time around.
 
Its been a little bit since I first started my storage project, but it is now nearing completion. The new RAID card arrived overnight from Singapore (talk about good service). Its probably time to talk a little about hardware as well, so that you will have a bit of perspective on my results as well.

For my system I chose a Super Micro SuperServer 5015A-H. I purchased the optional PCI Riser Card (CSE-RR1U-E8) to support the Adapted 5085 RAID (2249100-R), and also the 2x 2.5" HDD Bracket (MCP-220-00044-0N) into which I have installed two Seagate Momentus 7200.4 (ST9250410AS) drives which as I will explain later are using software RAID-1 (not the Intel chipset fakeRAID).

I am using an iStarUSA dAGE840TL-ML2 case for the drives which as I already mentioned are 8 Hitachi Deskstar 7K2000 (HDS722020ALA330) running firmware JKAOA20N. The case is connected to the raid card via 2x Multilane Infiniband SFF-8470 to miniSAS SFF-8088 (1m) cables.

My first issue immediately after closing the Super Micro 5015A-H case and installing it in my rack was the fact that it is passively cooled, and with the case closed and installed the RAID card jumped after a couple of hours to a blistering 106C (ouch!). This was quite well above the maximum (100 C) operating temperature of the card and so needed to be solved ASAP. The case has plenty of room for just about any style fan(s) you could wish, and the motherboard has connectors for up to 3 fans--either 3 or 4 pin (however one of the connectors is actually underneath the RAID card and therefore mostly inaccessible). I opted for installing a couple Evercool <21db 5.6cfm el-cheapo fans--they are 3 pin which means they report RPM, but do not control RPM based on MB temperatures. This is fine as they added virtually no noise to the box, but the RAID card now runs at 74C which for a home environment is perfectly acceptable. Plus the Adaptec Manger software emails me not only about issues with the RAID itself, but also card status (i.e. temperatures, etc.) as well, so if one the fans fail and the temperature of the card raises it will be easy enough to replace them.

My next issue is that one of the fans in in the backplane of the iStarUSA drive enclosure was already in the process of failing (buzzing, off tilt). Did I offend someone or something--why am I having so many issues right off the bat? :p Anyway, I contacted iStarUSA about a replacement and they had me fill out (no BS) a two page form to authorize my credit card be charged $69.00 to cross-ship me a replacement unit, which required no less than every piece of information known to man, including copies of my credit card front AND back, and my drivers license. Then their automated system for converting faxes to PDFs was apparently on the fritz and it took them two days to finally say they could ship me the product. That's when they finally decided to mention that UPS ground (5 days) was the only way they could ship it to me. Frustrated, I gave them my Fedex number to overnight it to me. Which by the way cost me about $90 because they selected early morning (vs $40 for next day 3pm). And now I still have to return the other unit or I won't get refunded the $69.00. So basically I could have purchased an entire replacement at an online e-tailer for $63.00, selected 2 day shipping, paid less and had it sooner, and had spare backplane parts to boot. Lesson learned. I can't recommend this company or their products because of this fiasco. Especially when Adaptec internationally overnighted me a product that MSRPs for $950 without even asking for a credit card. The stark contrast in customer service left a bad taste in my mouth.

Anyway, all the hardware is up and working now. I have chosen to use a dual boot operating system because IMHO there is no comparison to the capabilities that OpenSolaris/ZFS offer me, however the one thing that does not work in in Solaris is smartmontools with the TLER/CCTL query/settings. Luckily the Adaptec Manager software monitors the S.M.A.R.T. status of all drives in the array and can automatically email me information about that from Solaris, but I have installed CentOS 5.4 to use the patch I mentioned in posts above to set the timeouts using smartctl on the drives. This isn't the totally optimal setup because if I hotswap a drive I still have to reboot into CentOS to set the timeouts again (because that removes power to the drives). However, the other benefits of Solaris (IPMP and native kernel CIFS/SMB for instance) and ZFS (checksumming, deduplication, L2ARC, ZIL offloading, easy iSCSI, NFS, CIFS/SMB etc.) are so desirable that this is the situation I will deal with until such time as I can set the TLER/CCTL timeouts from Solaris. I won't bother expounding on all the reasons ZFS is the best file system (especially for data integrity) out there hands down, a simple Google search with some of the terms I mentioned will give you more information than you could possibly want.

So my initial testing reveals that the raw RAID-6 array speed is much faster than I expected it to be. My initial tests for sequential read and write of 4GB (this is so that I surpass any possible System RAM/RAID RAM) reveal the following:
Code:
dd raw write command:
dd if=/dev/zero of=/dev/rdsk/c4t0d0 bs=128k count=32000

dd raw read command:
dd if=/dev/rdsk/c4t0d0 of=/dev/null bs=128k count=32000

dd raw write test: 334 MB/s (Avg 5 tests)
dd raw read test: 550 MB/s (Avg 5 tests)

Wow! Now keep in mind that's just the raw throughput. You'll not see this when using files on top of an actual filesystem unless using enterprise class processors and cache levels. However, imagine if I used a pair of SSD drives in my Super Micro server instead of those spindle drives. I could then use partitions on them for L2ARC cache AND a mirrored ZIL device. Then the entire system could probably come close to achieving those speeds internally as the processors (dual core with hyper threading remember, so 4 simultaneous threads--these atom 330 based systems are actually pretty impressive for what they are).

Anyway, my purpose is a NAS based system which means gigabit line speed. That's 125 MB/s minus overhead (so realistically 100-110 possibly). I haven't gotten around to testing that yet, I'm currently running multiple iozone/bonnie++ tests in different configurations using partitions on my two Seagate 250GB drives in the server to provide ZIL mirror log and L2ARC cache devices to see if (even using crappy spindle drives) I gain (or lose) any benefit (IOPS or Read/Write speed). After I've decided on a particular configuration that way I'll get around to starting network tests/tuning. Keep in mind that with Solaris and IPMP I'll get the benefit of being able to pump out upwards of 2Gbps into my network at any given time, so my ultimate goal will be to use the configuration that best allows my expected workloads to achieve 220+ MB/s read speeds.

I won't post my final results of tests with filesystems yet, however I do want to post a really interesting finding--ZFS is already substantially faster at certain operations than linux?!? See the following snippets of iozone output:
Code:
CentOS 5.4
                                                            random  random
              KB  reclen   write rewrite    read    reread    read   write
         4194304     128  123771  134205   252528   272590   17820   18228
         4194304     256  117605  160794   250512   270268   20989   25092
         4194304     512  120464  101472   249588   265197   27938   37846
         4194304    1024  120710  154903   245667   264532   34498   52728
         4194304    2048  123287  155914   247618   269044   50405   64229
         4194304    4096  124749  118991   250668   267004   78982   78399

Code:
OpenSolaris snv_131 w/ZIL offload w/o L2ARC
                                                            random  random
              KB  reclen   write rewrite    read    reread    read   write
         4194304     128  185339  174198   334740   329705   18111  190387
         4194304     256  181021  169919   318968   323905   30918  169092
         4194304     512  162369  158631   312491   317606   38205  162445
         4194304    1024  149131  174788   308877   319532   47994  174654
         4194304    2048  180793  173083   315660   320117   56720  167129
         4194304    4096  178612  166717   309651   311945   71683  176542

How odd is that? I will leave detailed information from a myriad of tests for my next post, but as we can see from these numbers I'm going to be very happy with my choice for a build. I'm using cheap desktop drives in a very secure RAID array (2 drives can completely fail before I have to worry about possible data loss) and achieving speeds that are easily going to surpass line speed (even dual line speed with IPMP output, i.e. reads). If you're looking at those random reads thinking that they kinda suck, you'd be wrong. If a 4GB+ file actually existed that was so heavily fragmented those numbers would be correct, but thats actually a bit unrealistic. And just so you can see if it were a 32MB file heavily fragmented (where random read/write speed might be a little more realistic) I leave you with this (can you say oh my god):
Code:
OpenSolaris snv_131 w/ZIL offload w/o L2ARC
                                                            random  random
              KB  reclen   write rewrite    read    reread    read   write
           32768     128 1116992 1111050  1394500  1434303 1308732 1103485
           32768     256 1063721 1129030  1077286  1099486 1056622 1149801
           32768     512  984880 1050356  1041442  1109946 1070782  993667
           32768    1024  938779  952665  1121915  1148820 1143934  975029
           32768    2048  946315  973213  1157382  1139363 1165954  972455
           32768    4096  945788  988713  1168332  1191435 1182231  958104
 
I don't know anything about Solaris, but if it is anything like linux, your last test is just measuring the speed of the OS buffer cache. In linux, I give iozone the -I option to use direct I/O, which bypasses the buffer cache.
 
Back
Top