• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

URGENT! RAID-5 Problem, need help fast!

sykotic

Limp Gawd
Joined
Jan 2, 2003
Messages
154
One of my customers has their main server which houses something like 300gb of title images(for land, houses, etc). There are 3 150gb drives in a RAID-5 array, alotting something around 279gb of space. Today the server reboot itself and when I got there after their call it has an error that says:
"Following Arrays have missing required members and cannot be configured:
Array-0-RAID-5"

I open the scsi utility and manage the array and the scsi ID's 0 and 2 are a dark grey and 1 is white. These drives are in a 5 bay hot swap enclose and I tried switching drives 0 and 1 and upon rebooting the scsi utility still shows 0 and 2 as grey and shows 1(which is usually 0) as white. How in the world am I supposed to figure out which drive is bad or if a drive is even bad?! I am stumped and I have to get this thing fixed, they are going to have to close tomorrow and turn away all their title closings because this thing is down, I HAVE to fix this thing this weekend. The only thing I could think to do is to just verify media on all 3 drives, I have checked all cables and connections, everything is ok it seems.

Oh yea, the controller card is an Adaptec 2120S.

There is another array which is a RAID-0 with 2 300gb drives. These drives were not in use. I don't know if this is related or just coincidence but today, they decided to use these drives, well I originally had the 300gb partition mounted in a folder on one of the partitions from the RAID-5 array. Today I unmounted it and made the 300gb partition it's own drive letter. The server just so happened to restart and hour or so later while they were copying some new images to the new drive letter. Is this at all related???
 
First things first...DO NOT MIX UP THE DRIVE NUMBERS. If you don't know what you are doing step away from the box as you seem to be messing with a companies core data.

And my guess is you can't just verify each drive, delete the array, recreate it, and restore from the backup....because there is no current back up???

I'm getting to some things to try, because I know how it feels to be under a shitstorm, but unless you are coming in behind someone else (which you don't appear to be doing)...you really need to think about how you are handling peoples livelihoods.

THAT SAID...(enough ball busting)

It is highly unlikely (although not impossible) that both drives went bad at the exact same time. It is more likely that either bad power, or some other kind of connection to the drives is having issues.

I would first take the drives out of the hotswap bays, and go get some new cables (just incase) and plug the drives up directly to the card. IT IS VITAL that the hard drives stay connected to the same port on the controller that they were originally connected to.

Start up the box and see what happens...hopefully all will be well.

If it isn't then the next step would be to fun a WD Diag on each hard drive to see whether or not they are infact good (as my next step isn't worth a piss if a drive(s) is dead)

If they all pass, and for some unforsaken reason they will not recongnize the array, then the final step is to delete the array. Yes, I know it sounds scary, but all you are doing is deleting the array from the Adaptec RAID BIOS. Once it is gone you are going to create a new array, using the EXACT same drives, on the EXACT same ports. ALL CONFIGURATION POINTS NEED TO BE EXACT. What I mean by this is that stripe size, name, everything has to be configured identically on the new array. Now this next step is absolutely crucial. Once you have configured the array again, it will ask you how you want the controller to implement the array. Last time I had to do this on an adaptec BIOS there was an option to just write the meta data, and to leave the disks intact. That is the option you want...because if the drives are good, and the controller is good, and the cables are good, then the raid arrays metadata was corrupted. And all this does is recreate the metadata, that is why it has to be setup exactly as the old raid array, and that is why you do the nondestructive raid creation.

You could have a bad controller I guess...possible, in which case you would want to do the same thing (although I have seen Intel controllers pickup an existing array - moving an array from one dead 975badaxe to an identical board).

You didn't by chance enable write caching...and if so, is there a battery backup module on your RAID controller? If there is no battery on your controller then in the future disable write caching. It will give you a performance hit, but it can prevent loss of data/array.
 
It may be a good idea to make drive images before deleting and recreating the array?

I mean you could always go someplace and get a 500-750GB IDE HDD and find a utility that makes "raw" images?
 
You could, I mean it can't hurt at this point. In my mind though the more you do with these drives, the more you run the risk of further screwing something up. Deleting the array and recreating it (using the non-destructive/meta-data rebuild) will not do anything whatsoever to the data on the drives.............if it is still there to begin with.


My first question to pretty much every business client is "how much downtime are you willing to work with...a day, half a day, an hour, none?"

Then develop a solution to do that. But I have seen it said before, and I have seen it said on these forums...RAID is not a backup.
 
Thanks guys, I did just as you said and got the drives back up by recreating the array with a quick init status. right now I am copying all the crucial data over to the 300gb mirror array using robocopy, it is going ULTRA slow and will only copy in sperts. I have no idea what caused this but I ordered 2 new drives identical to these model drives, a new hot swap enclose(these drives are ultra320, whoever built this server put these drives in an enclosure that only supports ultra160), and I also ordered a point-to-point non-terminated 68pin cable because whoever built this server used a 5 device 68pin cable that is terminated(the hotswap board has its own termination). So for now they are using the server as it backs up(they are only using it if they absolutely have to). I overnighted all those parts so they will be here tomorrow, hopefully I can get things together tomorrow. Next week i'll be ordering a tape backup system, can anyone recommend a drive/software/controller card? I am thinking a 400/800gb system.
 
I'd also look at an online storage system, like a sata raid6 array. It only takes a terabyte to have 6 backups of 150 gigs, and that's 4 500gig disks plus controller. That's probably a lot cheaper than losing all the data. Tape backup is a must, yes, but for quick restoral disk-to-disk is much faster.
 
Back
Top