The thing about a Home RAID system is that you're not an enterprise with a multi-million dollar IT budget, so the cost of the drives DOES matter. You just want a good, safe place to put all your pictures, personal files, and music. And if you're like me you think it'd be just grand if you could rip your 1000+ DVD/Blu-Ray collection and play them all from your home theater screen using something like this.
The thing is, you can build a state of the art server/raid array combo for quite a reasonable price these days. The thing that really adds up are the drives. So you buy an enterprise class raid card capable of RAID-6 (and later on RAID-60) which means that two drives can catastrophically fail without suffering data loss. To me that's a great reason to use cheaper desktop class drives (remember that while we don't want to lose data, we're not a massive-budget enterprise either).
Of course the biggest issue we run into using these type of drives in combination with an enterprise class raid card is the time limited error recovery issue (TLER aka CCTL).
So we want to pick our drives carefully. Unfortunately there is nothing so nice as WDTLER available for any drives but Western Digital; and they are actually so greedy (for you to buy their 2x priced enterprise drives, even for a simple home raid) that not only did they change firmware to prevent the utility from permanently changing the settings--they actually went out of spec and disabled S.M.A.R.T. SCT Error Recovery functionality as well. Which is enough reason for me to never consider patronizing them again.
Which brought me to Seagate vs Hitachi vs Samsung. Seagate has had some ridiculous horror stories recently after their acquisition of Maxtor (is anyone really surprised?) and as much as I have been a huge fan in the past I just couldn't bring myself to trust that their latest generation of drives meet muster. Samsung on the other hand seems to have a very small market share, and generally little information available on the internet. Which leaves Hitachi (aka IBM). But not really, because IBM drives basically sucked. But apparently after Hitachi took over IBM's had drive division things really improved. You can find posts on the internet saying "...I will never buy from this manufacturer again because the drive I ordered was DOA..." for pretty much every single drive/manufacturer out there. But for Hitachi's recent lines the level of satisfaction seems to be quite high (compared especially to Seagate). Soo...I have decided to give them a try. Besides, they have tons of detailed documentation and utilities available on their website (see below).
Therefore, I have purchased 8 Hitachi 7K2000 drives for use in a RAID-6 array on an Adaptec 5085 adapter. This should give me just shy of 11 usable terabytes (2TB really means 1.8 and RAID-6 being N-2 so 6*1.8)--a great start. Without getting into system details (which for now doesn't matter much) I will add to the discussion with what I have learned.
First and foremost, the Adaptec card cannot even finish a build/verify on the RAID-6 without one of the drives dropping (having done nothing but hook up the raid box to the adapter and attempt to create a raid).
This was somewhat surprising, as I did check each drive's S.M.A.R.T. attributes (using smartctl 5.38, from GParted Live) prior to insertion in the RAID box and everything is fine. I knew I would have to address the CCTL (aka TLER) issue; I just didn't think it would affect me right out of the gate.
So, currently this is what I'm doing and why:
1) Use the Hitachi Feature Tool to:
1a) Disable the Write Cache on each drive. This is because the RAID card should be the one managing writes; and when it performs a write, it should do so with the expectation that each drive has actually written the data right then and there. This prevents a situation where a power loss to the RAID box could result in data loss that the adapter/OS would never be aware of. Aditionally, the ERC Write Timeout value should not have to be set (this is a good thing, follow the links below on SCT Write Timeout specification as to why). And yes, the Adaptec card has an option for supposedly disabling write cache on all connected drives, but to be safe I'll force it on the drive side.
1b) Enable Power Management with value 128 (0x80). As I am planning on using this array in a home environment, allowing the drives to reach idle states will save a substantial amount of electricity, furthermore the raid card *should* be capable of spin-down as well, adding further to potential savings. Dollar-wise the savings pale in comparison to the cost of the drives, but since there is no actual need (in my application) to have them constantly running at full power, it is still a good idea.
2) Use the Hitachi Drive Fitness Test to perform an Advanced Test on each drive (this takes HOURS) to verify that each drive is truly in prime condition.
The above mentioned tools are available at the following URL: http://www.hitachigst.com/hdd/support/download.htm
Now, after having done all of this there should (theoretically) be one final step, and that is to set a Read ERC timeout. I won't finish performing advanced tests on all drives until a few days hence, so for now I will just add what information I have learned, and later update with results.
So the first tool I came across for changing SCT Error Recovery values is the famed HDAT2. It works, you can plug the drive directly into an SATA port and access it, change the value for read timeout to 70 (7000 ms/7s) and it sticks through reboot. The problem is, I haven't figured out any way to use it while the drive is plugged into the RAID box and that box is connected to the Adaptec 5085. If anyone knows how please let me know and I'll try it just to let people know.
In the meantime, since values changed via S.M.A.R.T. SCT do not keep after power has been removed (e.g. unplug the drive from the direct connection to the computer and plug it back into the RAID box) it is not a very useful tool for my purposes. After all, if you have the drives directly connected to your computer (not using a hardware RAID card) you will likely want to use ZFS, and then read/write timeout values will be a non-issue anyway.
The other issue with the HDAT2 approach is that inevitably your raid box will lose power at some point (even if you have a UPS, which I do). Reasons include moving equipment, long term power failure, etc. Lets face it, we're building the RAID because we want a highly reliable, low administration, home storage network. Having drives revert after power loss and having to remember to re-run an application doesn't meet that criteria.
Having a service on server boot check (and if necessary set) those values however does. Which is why I have extremely high hopes for this. I haven't gotten around to porting his patch to smartmontools 5.39 yet, but if It works as advertised I'll be able to have a simple script run on server boot that checks the read recovery timeout value, and if necessary sets it to 70. I will post back later with my experience once I have all my drives tested and have tried this.
For information on Error Recovery Control, and specifically the SCT commands to set timeouts see http://www.t13.org/Documents/UploadedDocuments/docs2008/D1699r6a-ATA8-ACS.pdf, section 8.3.4.
For specification information on the 7K2000 series drives (including additional specific information for SCT commands) see http://www.hitachigst.com/tech/techlib.nsf/techdocs/5F2DC3B35EA0311386257634000284AD/$file/7K2000_A7K2000_Spec_r1.0.pdf.
The thing is, you can build a state of the art server/raid array combo for quite a reasonable price these days. The thing that really adds up are the drives. So you buy an enterprise class raid card capable of RAID-6 (and later on RAID-60) which means that two drives can catastrophically fail without suffering data loss. To me that's a great reason to use cheaper desktop class drives (remember that while we don't want to lose data, we're not a massive-budget enterprise either).
Of course the biggest issue we run into using these type of drives in combination with an enterprise class raid card is the time limited error recovery issue (TLER aka CCTL).
So we want to pick our drives carefully. Unfortunately there is nothing so nice as WDTLER available for any drives but Western Digital; and they are actually so greedy (for you to buy their 2x priced enterprise drives, even for a simple home raid) that not only did they change firmware to prevent the utility from permanently changing the settings--they actually went out of spec and disabled S.M.A.R.T. SCT Error Recovery functionality as well. Which is enough reason for me to never consider patronizing them again.
Which brought me to Seagate vs Hitachi vs Samsung. Seagate has had some ridiculous horror stories recently after their acquisition of Maxtor (is anyone really surprised?) and as much as I have been a huge fan in the past I just couldn't bring myself to trust that their latest generation of drives meet muster. Samsung on the other hand seems to have a very small market share, and generally little information available on the internet. Which leaves Hitachi (aka IBM). But not really, because IBM drives basically sucked. But apparently after Hitachi took over IBM's had drive division things really improved. You can find posts on the internet saying "...I will never buy from this manufacturer again because the drive I ordered was DOA..." for pretty much every single drive/manufacturer out there. But for Hitachi's recent lines the level of satisfaction seems to be quite high (compared especially to Seagate). Soo...I have decided to give them a try. Besides, they have tons of detailed documentation and utilities available on their website (see below).
Therefore, I have purchased 8 Hitachi 7K2000 drives for use in a RAID-6 array on an Adaptec 5085 adapter. This should give me just shy of 11 usable terabytes (2TB really means 1.8 and RAID-6 being N-2 so 6*1.8)--a great start. Without getting into system details (which for now doesn't matter much) I will add to the discussion with what I have learned.
First and foremost, the Adaptec card cannot even finish a build/verify on the RAID-6 without one of the drives dropping (having done nothing but hook up the raid box to the adapter and attempt to create a raid).
This was somewhat surprising, as I did check each drive's S.M.A.R.T. attributes (using smartctl 5.38, from GParted Live) prior to insertion in the RAID box and everything is fine. I knew I would have to address the CCTL (aka TLER) issue; I just didn't think it would affect me right out of the gate.
So, currently this is what I'm doing and why:
1) Use the Hitachi Feature Tool to:
1a) Disable the Write Cache on each drive. This is because the RAID card should be the one managing writes; and when it performs a write, it should do so with the expectation that each drive has actually written the data right then and there. This prevents a situation where a power loss to the RAID box could result in data loss that the adapter/OS would never be aware of. Aditionally, the ERC Write Timeout value should not have to be set (this is a good thing, follow the links below on SCT Write Timeout specification as to why). And yes, the Adaptec card has an option for supposedly disabling write cache on all connected drives, but to be safe I'll force it on the drive side.
1b) Enable Power Management with value 128 (0x80). As I am planning on using this array in a home environment, allowing the drives to reach idle states will save a substantial amount of electricity, furthermore the raid card *should* be capable of spin-down as well, adding further to potential savings. Dollar-wise the savings pale in comparison to the cost of the drives, but since there is no actual need (in my application) to have them constantly running at full power, it is still a good idea.
2) Use the Hitachi Drive Fitness Test to perform an Advanced Test on each drive (this takes HOURS) to verify that each drive is truly in prime condition.
The above mentioned tools are available at the following URL: http://www.hitachigst.com/hdd/support/download.htm
Now, after having done all of this there should (theoretically) be one final step, and that is to set a Read ERC timeout. I won't finish performing advanced tests on all drives until a few days hence, so for now I will just add what information I have learned, and later update with results.
So the first tool I came across for changing SCT Error Recovery values is the famed HDAT2. It works, you can plug the drive directly into an SATA port and access it, change the value for read timeout to 70 (7000 ms/7s) and it sticks through reboot. The problem is, I haven't figured out any way to use it while the drive is plugged into the RAID box and that box is connected to the Adaptec 5085. If anyone knows how please let me know and I'll try it just to let people know.
In the meantime, since values changed via S.M.A.R.T. SCT do not keep after power has been removed (e.g. unplug the drive from the direct connection to the computer and plug it back into the RAID box) it is not a very useful tool for my purposes. After all, if you have the drives directly connected to your computer (not using a hardware RAID card) you will likely want to use ZFS, and then read/write timeout values will be a non-issue anyway.
The other issue with the HDAT2 approach is that inevitably your raid box will lose power at some point (even if you have a UPS, which I do). Reasons include moving equipment, long term power failure, etc. Lets face it, we're building the RAID because we want a highly reliable, low administration, home storage network. Having drives revert after power loss and having to remember to re-run an application doesn't meet that criteria.
Having a service on server boot check (and if necessary set) those values however does. Which is why I have extremely high hopes for this. I haven't gotten around to porting his patch to smartmontools 5.39 yet, but if It works as advertised I'll be able to have a simple script run on server boot that checks the read recovery timeout value, and if necessary sets it to 70. I will post back later with my experience once I have all my drives tested and have tried this.
For information on Error Recovery Control, and specifically the SCT commands to set timeouts see http://www.t13.org/Documents/UploadedDocuments/docs2008/D1699r6a-ATA8-ACS.pdf, section 8.3.4.
For specification information on the 7K2000 series drives (including additional specific information for SCT commands) see http://www.hitachigst.com/tech/techlib.nsf/techdocs/5F2DC3B35EA0311386257634000284AD/$file/7K2000_A7K2000_Spec_r1.0.pdf.