• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

RAID0, Newbies, and How To Handle Your Data

Am I an idiot?

  • Yes

    Votes: 7 21.2%
  • No

    Votes: 10 30.3%
  • Magical Brownies?

    Votes: 9 27.3%
  • Your post was too long, so I didn't read it.

    Votes: 7 21.2%

  • Total voters
    33

Stiletto One

2[H]4U
Joined
May 7, 2002
Messages
3,906
So. There have been ALOT of threads lately with newcomers asking for suggestions on their spec lists (pre-purchase, thankfully), which have included the likes of ThermalTake cooling gear and RAID0 arrays.

On ThermalTake cooling gear: I'm not going to bother defending myself, just don't buy it. Their power supplies aren't TOO bad, but I'll just say to get a Fortron or Antec or Enermax.

Now for RAID0: I'm seeing people who just don't know anything about it except for "makes drives into one big fast drive".
1.) Yes, this is true. RAID0 is called "striping" and splits data across every drive in the array. Theoretically, this doubles read and write speeds, because the RAID controller can split data traffic over the two drives. One drive can handle X amount of data per second; two drives can handle 2X amount of data per second, and so on. So that's what RAID0 does, and why it exists.

2.) Theory != practice. The kind of RAID controller that ships strapped onto consumer-level motherboards and controller cards are "host based". This has the same meaning as with printers: the host CPU does pretty much all of the processing work. So a RAID controller can easily suck up over 20% of your CPU resources when there's heavy hard drive traffic. If you're doing a home file/web server, this might be OK; in gaming, it's bad for you.

2.b) You will almost never get anywhere near full theoretical speed; the increase in access latency (the time it takes for your storage system to retrieve whatever file you're requesting) will offset any perceived speedup from increasing raw bandwidth. The car analogy works very well for this: would you rather have a car with a top speed of 200mph and a 0-60 of 12s, or one with top speed of 140 and a 0-60 of 8s?

3.) Probably the worst for those of us with no money, data redundancy. Comma. Lack of. RAID stands for "Redundant Array of Inexpensive Disks"—RAID0 is redundant. Hence, the zero moniker: RAID0 isn't actually RAID. (Zero times X does not give a multiple of X, it gives zero. Algebra, foo'.) This means that if any one drive goes sour, then you lose EVERYTHING. No data recovery, no "fdisk /mbr", just GONE. Take everything I just wrote, and drop out every other character. How readable is it? NOT VERY. And it only gets worse as you add drives.

My conclusion: if you want lots of storage on more than one, the smart way is to just use those drives and partition them up. For example, instead of having one huge 160GB volume striped over two 80GB drives, I would do this:

Drive 1: Three partitions
#1: C drive. 10GB, for Windows and swap file. Nothing else.
#2: D drive. 30GB for program files. Literally, put Program Files there. Go to regedit, and navigate to HKEY_LOCAL_MACHINE\SOFTWARE\MICROSOFT\WINDOWS\CURRENTVERSION and set the ProgramFilesDir directory to whatever you want to call it.
#3: E drive. 40GB for infrequently used files which you want to hold onto, like patches and installers. I guess music is OK, since the files aren't huge and can be quickly loaded into RAM.
Drive 2: One partition
#1: F drive. 80GB. Done. I would put My Documents on F drive (right-click on My Documents on the desktop, and it lets you map it), and just work from there.

This will still be faster than just one huge 160GB drive, on account of separating data accesses from everything else. And if anything corrupts a partition (most likely to be your OS partition), your data is still intact and all of the patches, drivers, and major service packs are already there (on E drive) for you to quickly get yourself to a semblance of up-to-date-ness.

Backups, of course, are still the way to go for data security. Archive any data you want to hold onto once you've decided that you won't be working with it anymore, and then go ahead and delete it from hard disk. Most of us home users don't have any need for multi-locality of data—after all, does it REALLY matter if you can't find that backup disk with your high school junior year English term paper, four years down the road?

OK, rant done. Breakfast time. :)
 
Having your data partition on a separate physical drive from your os partition used to be a big deal, but if you have enough memory (I have 1 gig) Windows will barely even access the OS files after bootup.

I would simplify your setup one further by having just one partition per drive and using separate folders in place of your specialized partitions. That way they all share the 80 gigs of that drive. You don't waste any extra space on your OS and you don't have to do any partition magicing if you happen to need more space for "program files" or whatnot.

Still, I agree with you as to RAID 0 being dangerous. It is mathematically twice as likely to fail as a single drive.
 
Yeah, works for some people. I've run into enough cases where people hosed one partition on a single drive to prefer keeping stuff highly compartmented.

10GB for the OS partition works pretty well; it allows for a 4GB swap file (if this happens, GET MORE RAM!) as well as some temp folder bloat. *shrug

And you can always start sticking programs on another partition.
 
DocSavage said:
It is mathematically twice as likely to fail as a single drive.

That's not really true... Suppose the probability of a single drive failing on any given day is 10% (1/10). By your logic, if you have 10 drives, one will CERTAINLY fail on any given day, because the probability of a 10-drive aray failing is 10x the probability of 1 drive failing.

The way to calculate the probability of a 2-drive array failing on any given day is like this:

Probability of a 2-drive array failing = probability of drive A failing and drive B not failing + probability of drive A not failing and drive B failing + probability of drive A and B both failing.

So, using the example of 10%, Probability of a 2-drive array failing = 0.1 * 0.9 + 0.9 * 0.1 + 0.1 * 0.1 = 19%. Correct me if I'm wrong please.
 
I'm pretty sure that's incorrect, but I'm not sure, either. If you have two drives that are both equally likely to fail, and one drive failing renders the other unreadable as well, I'd think it would double your chances of data loss, wouldn't it?
 
Ugly_Jim said:
I'm pretty sure that's incorrect, but I'm not sure, either. If you have two drives that are both equally likely to fail, and one drive failing renders the other unreadable as well, I'd think it would double your chances of data loss, wouldn't it?

The problem with this is that if you keep adding drives, you can eventually have a probability of failure which is > 100%, which is meaningless.


Anyway, I think a decent alternative to Stiletto One's RAID-0 alternative would be a setup like this:

Keep using RAID-0 with two big, fast drives, but add a 3rd, smaller drive for backup. Have some automated tool make nightly backups of the important stuff (or daily backups if you turn your comp off at night... leave it on when you go to work/school or something :) ). Most of us probably have an older HD we can use for backup purposes.
 
TheSpoon said:
The problem with this is that if you keep adding drives, you can eventually have a probability of failure which is > 100%, which is meaningless.

very well said. i think your math stands, but you should also realize that RAID 0 is usually used (in homes) for two drives. so, the difference between 20% and 19% (or whatever the two numbers would be) is negligible. in fact, as you decrease the probability of an individual failure, the difference between the estimated and actual failure probabilities becomes smaller.

in any case, a point well made, i believe.

EDIT: bah HTML...
 
if u flip a coin randomly u have a 50% chance of heads or tails on any given throw, flipping two coins also nets a 50% of heads or tails. The chance of heads or tails is not cummulative on consecutive tosses. So assuming .1 failure rate for each drive the chance of any 1 drive failing is still .1. Correct me if im wrong. im just drunk and need to learn statistics again...
 
The probability of 1 drive failing is still .1, but we're discussing the probability of the array failing. The array will fail if either of the drives fails, or if both fail.
 
MTTF(RAID 0) = MTTF / # of Disks

MTTF(RAID 1) = (MTTFDisk)² / # of Disks

and thats directly from 3ware

MTTF = Mean Time To Failure

so the odds of data loss does double with a dual channel RAID 0 array
and yes it only gets worse with the number of drives you have, where your getting tripped up is in how MFFT and MTBF is figured
and its worse if you figure in the controllers MTTF, and even worse if the controler is on the motherboard

I wrote off a RAID 0 array rather than replacing the mobo based on what RAID controller it had :p

http://tech-report.com/reviews/2001q2/realraid/index.x?pg=2

so wheres my brownie?

a more useful way to figure is say your drive has a warranty of 3 years
or 1095 days, a dual channel RAID 0 would halve that, and a 4 channel quarter it
as rough odds, these rapidly increase to a 100% chance of loosing a 12 channel RAID 0 array built with drives of various ages and or condition, always run a drive for a few months under heavy access after you get it, giving it a chance to develop "shipping" damage
then trust it to a RAID 0 array, and monitor its S.M.A.R.T. status, they are of limited usefulness for predicting slow degredation of a drive, but completely useless against catastrophic failures

a 12 channel made of three year warranty drives would see an average failure around the 3rd month if you employed that formula,
which would probably be accurate for someone just picking 12 vendors at random and buying 12 drives,
it might even be optimistic :p
 
I just wish people would read the storagereview FAQ entry on RAID 0 before they start thinking sequential transfer results are going to correlate at all with their real world usage.

Seeks people, seeks!
 
The debate over raid reliability formulas got me curious, and I eventually stumbled across this .pdf from sun: http://www.sun.com/blueprints/0602/816-5132-10.pdf "Network Storage Evaluations Using Reliability Calculations"

Some very interesting reading in there. They use a drive's MTBF to come up with a percentage failure rate for the year. They then use that percentage in their raid formulas.

I would try to retype what they are doing in here, but it is confusing my mind this early in the morning...
 
there use to be a RAID Java calculator
I had a link to that included parity levels
but unfortunately it went dead last year
and it would spit out these amazing figures

but such formulas just address hardware failure at the HDD and controller level
if everything else goes well
completely ignoring the leading causes of data loss
pilot error
corruption (memory, power, filesystem, ect)
>Hardware Failure
malware

(a personal list)
 
Pilot error?

I think the numerical setup we use in math is the counter probability.

Instead of saying that in, say, a two-year period, we have a fail rate of 10%, we have a pass rate of 90%.

So with one drive, our pass probability is 90%.

With two drives, it becomes 90% chance of drive 1 passing, then times 90% chance of drive 2 passing. 81% overall pass rate.

With three, it becomes (90%)^3, for 72.9%. And so on.

...

I think Spoon was trying to get at this, but I'm not quite sure.
 
TheSpoon said:
The problem with this is that if you keep adding drives, you can eventually have a probability of failure which is > 100%, which is meaningless.
.

No, as the number of drives approaches 'infinity' (for arguement's sake) the limit goes to 100% with a RAID 0 array.

Clearly there are practical limits, such as 15 slots on a SCSI channel or a 12 channel SATA RAID card...
 
Stiletto One i dont know why i couldnt say that (i was thinking it atleast lol). Anyways i just ordered a 76gig raptor from newegg. Maybe ill raid it later this year but for now ill spin at 10k,
 
Stiletto One said:
Pilot error?.

Oh yes pilot error
from a mistake in a commandline to an error in a manual
to a n00b reformatting the wrong partition

I'll relate my story for those that havent already heard it
migrating my 6 Channel RAID 5 array for the second time (new case)
I disconnected all the cables without marling which channel is which
cause the manual said I could, only when I hooked them back up 2 drives appeared
get into the knowledge base and buried deep is,
never disconnect all the drives at the same time :rolleyes:

Dual Channel RAID
AB BA

Tri Channel RAID
ABC ACB BAC
BCA CBA CAB

Quad Channel RAID
ABCD ABDC ACBD ACDB ADCB ADBC
BACD BADC BCDA BCAD BDCA BDAC
CABD CADB CBDA CBAD CDAB CDBA
DACB DABC DBCA DBAC DCBA DCAB

and a six channel RAID :eek:

needless to say I lost the array
a later BIOS update fixed the problem, but too late for me.
Luckily it was a little over 90% backed up to Hard Media
 
jen4950 said:
No, as the number of drives approaches 'infinity' (for arguement's sake) the limit goes to 100% with a RAID 0 array.

Clearly there are practical limits, such as 15 slots on a SCSI channel or a 12 channel SATA RAID card...
Yes, if you follow the rule of thumb that every time you double the amount of drives, the probability of failure doubles, the probability will eventually exceed 100%. Which is why the rule does not work.
 
Ice Czar said:
MTTF(RAID 0) = MTTF / # of Disks

MTTF(RAID 1) = (MTTFDisk)² / # of Disks

and thats directly from 3ware

MTTF = Mean Time To Failure

so the odds of data loss does double with a dual channel RAID 0 array
and yes it only gets worse with the number of drives you have, where your getting tripped up is in how MFFT and MTBF is figured
and its worse if you figure in the controllers MTTF, and even worse if the controler is on the motherboard

I wrote off a RAID 0 array rather than replacing the mobo based on what RAID controller it had :p

http://tech-report.com/reviews/2001q2/realraid/index.x?pg=2

so wheres my brownie?

a more useful way to figure is say your drive has a warranty of 3 years
or 1095 days, a dual channel RAID 0 would halve that, and a 4 channel quarter it
as rough odds, these rapidly increase to a 100% chance of loosing a 12 channel RAID 0 array built with drives of various ages and or condition, always run a drive for a few months under heavy access after you get it, giving it a chance to develop "shipping" damage
then trust it to a RAID 0 array, and monitor its S.M.A.R.T. status, they are of limited usefulness for predicting slow degredation of a drive, but completely useless against catastrophic failures

a 12 channel made of three year warranty drives would see an average failure around the 3rd month if you employed that formula,
which would probably be accurate for someone just picking 12 vendors at random and buying 12 drives,
it might even be optimistic :p

I agree that the MTTF is halved when you have two drives as opposed to one drive, since between them they accumulate twice as many working hours as a single drive. Of course, this is assuming that HD failure rates are exponentially distributed (otherwise you have to take age into account, etc etc), which seems to be the case from the reading I've done.

However, this does not mean that the probability of failure on any given day is doubled. Again, this can't make sense because if you have two drives whose probability of failure is 90% (for argument's sake), I can assure you that the probability of failure of a RAID-0 array comprised of these two drives is not 180%.

Stiletto One: Exactly. According to your calcs, with 2 drives there's an 81% probability of non-failure, and according to mine there's a 19% probability of failure. Same thing. Your way is quicker, though :)
 
TheSpoon said:
Yes, if you follow the rule of thumb that every time you double the amount of drives, the probability of failure doubles, the probability will eventually exceed 100%. Which is why the rule does not work.


NO. That's not how it works (statistics)-

Hypothetically, say a drive has a probability of success of 0.9995 (1 error in (~5.5years))

And for a 2 drive RAID0 array the probability of success is:

(0.9995^2)*(0.0005^0) = 0.9990

For a 200 drive RAID0 array the probability of success is:

(0.9995^200)*(0.0005^0) = 0.9048

For an infinite number of drive in a RAID0 array, probability of success:

limit as x -> infinity: (0.9995^x)*(0.0005^0) = 0.


Likewise, for a 200 drive RAID1 array, arranged in 100 2 disk arrays:

[(0.9995^1)*(0.0005^1)]^100*{1-[(0.9995^1)*(0.0005^1)]}^0 = 0.999975



You cannot have a probability over 100%. EVER.
 
It seems you haven't noticed, but I am refuting the idea that 2x the drives = 2x the probability of failure. How is it that you do not see that we are on the same side?
 
TheSpoon said:
It seems you haven't noticed, but I am refuting the idea that 2x the drives = 2x the probability of failure. How is it that you do not see that we are on the same side?

nm. - I don't know how to read.
 
jen4950 said:
When you say that the probability will eventually exceed 100%.

I'm a stickler for details.

I said that the probability eventually exceeds 100% if you believe that 2x the drives = 2x the probability of failure, which is the idea I am refuting.
 
TheSpoon said:
I agree that the MTTF is halved when you have two drives as opposed to one drive, since between them they accumulate twice as many working hours as a single drive. Of course, this is assuming that HD failure rates are exponentially distributed (otherwise you have to take age into account, etc etc), which seems to be the case from the reading I've done.

yup MTBF \ MTTF is definately convoluted means of rating drives or any components
but as it sums up in that link (Mean Time Between Failures (MTBF)
Overall, MTBF is what I consider a "reasonably interesting" reliability statistic--not something totally useless, but definitely something to be taken with a grain of salt. I personally view the drive's warranty length and stated service life to be more indicative of what the manufacturer really thinks of the drive. I personally would rather buy a hard disk with a stated service life of five years and a warranty of three years, than one with a service life of three years and warranty of two years, even if the former has an MTBF of 300,000 hours and the latter one of 500,000 hours.

In the real world, the actual amount of time between failures will depend on many factors, including the operating conditions of the drive and how it is used; this section discusses component life. Ultimately, however, luck is also a factor

I did however want to point out the difference between drive failure and data loss again, as you increase the stripe width, you are adding more than just drives that can fail, the increased number of cables, and often caddies with the additional interface, the increased power load, and thermal aspect. I think that if all the variables where to be addressed and tested with a large sample, youd find that as time passes from the initialization of a large RAID array the potential for data loss actually decreases for the first several months, then increases again when the end of the useful service life of the drives approaches.

the "shakeout" period I mentioned earlier ;)
 
Ice Czar said:
I did however want to point out the difference between drive failure and data loss again, as you increase the stripe width, you are adding more than just drives that can fail, the increased number of cables, and often caddies with the additional interface, the increased power load, and thermal aspect. I think that if all the variables where to be addressed and tested with a large sample, youd find that as time passes from the initialization of a large RAID array the potential for data loss actually decreases for the first several months, then increases again when the end of the useful service life of the drives approaches.

the "shakeout" period I mentioned earlier ;)

Yep, RAID-0 certainly introduces several other factors (breaking points) besides the added drive(s) alone. I can't comment on the curve you're describing. Would be nice to see some real (empirical) data, plotting drive failure rates (especially in a RAID-0 environment) vs. time. The exponential distribution seems like somewhat of an oversimplification, but that's just intuition talking...I could be totally wrong.
 
I base that on what Ive seen in here regarding a certain class of drive failure
the same as described in the SR FAQ
Which Brand of Hard Drive is Most Reliable?

if such damage doesnt manifest itself within the first month or so, I think the odds of it actually reaching its rated service life increase as time passes
 
Having read the above, I think that's what one has larger capacity drives vs faster RPM drives.

I use two WD Raptor drives for system and programs. I have three other hard drives (2 200 gig Maxtor 8meg cacher's and 1 100gig WD SE 8meg cacher). I partitioned these into a small section on each for a static (unchanging) paging file on each drive.

I tried going without a paging file, but Windows XP seems smart enough not to page very often anyway and a lot of programs that I use needed the page files. So I thought by spreading the page files across multiple drives (each one attached to its own controller except the 100giger (which is on the same channel as the CD-RW I use rarely) I'd be giving the computer the ability to access each page file simultaenously incurring less of a penalty than by having the page file on the same drive as the Windows or program drive.

The rest of the hard drives (except the page files) are for data. I used to have the two 200giger's in RAID 0, but I decided that was just too dangerous for data and took them out of it. I've noticed they seem just a hair slower in responsiveness, but the better certainty is nice as far as losing data. As far as Windows or my programs going with the RAID'ed Raptors partition, well, I really wouldn't be bothered. I'd just reinstall.

So I see RAID 0 as nice for programs and Windows because it does add to performance in swifty loading and file transfer (files that usually start out first in C). Data drives and page file drives do not need to be as fast just so long as they do what they do occassionally. Seems to me a page file on a RAID'ed drive is just asking for the computer to crash, though.

Which is why I put those page files on other directories and watch Windows share across various channels.

I mean, if I'm going to be using Hyperthreading to make my computer more responsive, then why not have a lot of hard drives accessed via several channels rather than have just one drive access by one channel?
 
jen4950 said:
NO. That's not how it works (statistics)-
Hypothetically, say a drive has a probability of success of 0.9995 (1 error in (~5.5years))
And for a 2 drive RAID0 array the probability of success is:
(0.9995^2)*(0.0005^0) = 0.9990
For a 200 drive RAID0 array the probability of success is:
(0.9995^200)*(0.0005^0) = 0.9048
For an infinite number of drive in a RAID0 array, probability of success:
limit as x -> infinity: (0.9995^x)*(0.0005^0) = 0.

It seems to me that this 100% way of calculating array reliability is not working...

Say your drive has a failure rate of 1 in 5.5 years. That translates to a MTBF of 5.5yrs/failure * 8766hrs/yr = 48213hrs/failure. To figure out annual failure rate (AFR) use 8766hrs/yr / 48213hrs/failure = 0.1818181818failure/yr.

Now suppose I have a 2 drive raid 0 array with the stats above. My odds of having either of the drives fail are 2 * 0.1818181818failure/yr = 0.3636363636failure/yr.

Suppose I have a 200 drive raid 0 array... 200 * 0.181818181818failure/yr = 36.36363636failure/yr.

Heheehe refute this!!!
 
DocSavage said:
It seems to me that this 100% way of calculating array reliability is not working...

Say your drive has a failure rate of 1 in 5.5 years. That translates to a MTBF of 5.5yrs/failure * 8766hrs/yr = 48213hrs/failure. To figure out annual failure rate (AFR) use 8766hrs/yr / 48213hrs/failure = 0.1818181818failure/yr.

Now suppose I have a 2 drive raid 0 array with the stats above. My odds of having either of the drives fail are 2 * 0.1818181818failure/yr = 0.3636363636failure/yr.

Suppose I have a 200 drive raid 0 array... 200 * 0.181818181818failure/yr = 36.36363636failure/yr.

Heheehe refute this!!!

Here's a good place for you to start: http://mathforum.org/probstat/probstat.lessons.html

Lots of good help for you there-

You're probability of having either one of the drives fail is 1 out of 2. (refuted.)

And I mentioned out of thin air a probability of success of %99.95 - which translates into having a failure on one day out of 2000.

Therefore, having one array depend on two drives with a probability of success of %99.95 each results in a probability of success of the array of %99.9. Two hundred drives, %90.5 - reflecting the greater risk of relying on more drives that all must function for the array to be operational.

at a %10 failure rate per drive, 15 drives will knock your chance of having a reliable array down to 20 percent.
 
jen4950 said:
You're probability of having either one of the drives fail is 1 out of 2. (refuted.)
^^^^
I don't get that statement at all, but whatever...

Now that I understand what your units are, it makes more sense. You are stating that with the 200 drive array, that there is a .095 (1.00-.905) chance of a drive failing on any given day.

My figures of 36.3636failure/yr / 365.25day/yr = 0.0995failure/day.

Our results are so close, that I am not sure that it is a fluke or not... Maybe it is the same information in two formats? Percent Probability of Daily Failure vs. Average Failures Per Day?

ps. Thanks for the math link... I've been looking for something like that :)
 
DocSavage said:
^^^^
I don't get that statement at all, but whatever...

Well, you have 2 drives and you want to know the probability of either one of the drives failing.

Drive one can fail
Drive two can fail.
There are two drives.

The probability of either drive failing is 1 (one of the drives) out of 2 (2 total drives). :confused:

You can't have a fraction of a drive fail.
 
jen4950 said:
Well, you have 2 drives and you want to know the probability of either one of the drives failing.

Drive one can fail
Drive two can fail.
There are two drives.

The probability of either drive failing is 1 (one of the drives) out of 2 (2 total drives). :confused:

You can't have a fraction of a drive fail.

Those are just outcomes. I don't see how the desired outcome divided by number of possible outcomes has anything to do with the chances of one of the outcomes happening. It's not like there are equal chances of either result...
 
DocSavage said:
Oh good lord... do you have any mtbf figures for your dog f'ing up the computer? :D

actually thats the MTBC mean time before chew :p

and that would be Schrödinger's dog, it has both chewed the cables and not chewed the cables, and until you try to initialize the array and colapse the probability field you dont know :p
 
DocSavage said:
Now that I understand what your units are, it makes more sense. You are stating that with the 200 drive array, that there is a .095 (1.00-.905) chance of a drive failing on any given day.

My figures of 36.3636failure/yr / 365.25day/yr = 0.0995failure/day.

Our results are so close, that I am not sure that it is a fluke or not... Maybe it is the same information in two formats? Percent Probability of Daily Failure vs. Average Failures Per Day?
I went ahead and plugged in 2, 50, & 200 drives in both of our formulas and each time the results are very close. When I am bored I may try to analyze and compare our methods.

ps. LOL Ice Czar :)
 
Back
Top