• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Areca 1880i RAID Set/Volume Set Missing - Help Needed

pirivan

Limp Gawd
Joined
Feb 22, 2009
Messages
346
Hi All,

I am running an Areca 1880i (FW version 1.48B, BIOS 1.22D) with an HP SAS expander (FW v2.02) and 11 Hitachi HDS722020ALA drives. The 1880i is connected to the HP SAS expander with two cables. The HP SAS Expander is connected to a Norco 4020 backplane (with SAS to 4x SATA cables). The drives were all 1 large RAID6 array with no hotspare (128K stripe size, 16K clusters). In the Areca boot time BIOS unfortunately it does not appear that I can pull the member disk sequence; all I have is a screenshot of how it looks currently below.

This evening while I was out, I came back to hear the server beeping loudly, two drive slots blinking and the server hard locked. I hard shut it down and booted it back up. Much to my alarm the RAID volume was not accessible. I fully shut down and rebooted several times but the RAID Set/Volume appears to be missing. Besides shutting down/rebooting a few times I have done nothing else except look at the controller BIOS or web UI. Below is what a picture of the web GUI looks like now (there should be one large RAID6 volume with ALL the drives in it):

http://i.imgur.com/2ngAyl.jpg - Link to larger version of picture below
2ngAyl.jpg


Here is what the RAID controller event log looks like now (the text as well as a couple of pictures):

011-10-30 22:25:36 192.168.001.002 HTTP Log In
2011-10-30 22:02:38 RS232 Terminal VT100 Log In
2011-10-30 22:01:53 H/W Monitor Raid Powered On
2011-10-30 21:56:13 192.168.001.002 HTTP Log In
2011-10-30 21:55:06 H/W Monitor Raid Powered On
2011-10-30 21:53:52 192.168.001.002 HTTP Log In
2011-10-30 21:50:29 H/W Monitor Raid Powered On
2011-10-30 21:50:19 192.168.001.002 HTTP Log In
2011-10-30 21:47:03 192.168.001.002 HTTP Log In
2011-10-30 21:45:47 H/W Monitor Raid Powered On
2011-10-30 21:32:15 RAID6 2TB DRIVES RaidSet Degraded
2011-10-30 21:32:15 ARC-1880-VOL#001 Volume Degraded
2011-10-30 21:31:56 Enc#2 SES2Device Time Out Error
2011-10-30 21:31:52 Enc#2 Slot#13 Time Out Error
2011-10-30 21:31:39 Enc#2 SES2Device Time Out Error
2011-10-30 21:31:09 Enc#2 Slot#11 Device Removed
2011-10-30 21:31:09 RAID6 2TB DRIVES RaidSet Degraded
2011-10-30 21:31:09 ARC-1880-VOL#001 Volume Degraded
2011-10-30 21:31:05 Enc#2 Slot#11 Time Out Error

http://i.imgur.com/a1K5gl.jpg - Link to larger version of picture below
a1K5gl.jpg

http://i.imgur.com/wVpHpl.jpg - Link to larger version of picture below
wVpHpl.jpg


What should my next move be? Should I try to rescue the RAID set? I want to be as careful as possible of course so I wanted to ask here before I started blindly doing anything and make my situation worse. Any help would be HIGHLY appreciated. I would be happy to provide any information you require; thanks in advance for any feedback/suggestions.
 
Last edited:
Well first of as a matter of procedure I would email Areca.

Secondly, dont do anything yet... let me digest this. I recently went through a massive failure on my R6 volume so this is still fairly fresh.
 
Hi Nitrobass24,

Thank you for the quick follow up; I appreciate it. Indeed, I just got finished essentially copying and pasting this post and sending it to Areca support.

I will absolutely hold off until on doing anything until I hear back from some experts here; I absolutely do not want to make my situation worse/more complicated. I have literally not done anything else besides what I mentioned above. If it means anything to anyone I did find this forum post that pretty much is my exact same situation but I wasn't sure if I should follow his end recommendation right away or not:
http://forums.storagereview.com/ind...rt-beeper->rebooted->everything-is-gon/

Basically he suggests re-creating the array and there is a lot of concern about 'drive order'. Hopefully I wouldn't have to worry about that because I haven't even physically removed any of the drives from their slots.

Anyhow, thanks again for the fast response and any feedback; as you might imagine I am a bit stressed!
 
Yea i will read that tomorrow when I am at work, but i wouldnt do anything in there just yet. The key to data recovery is to take everything one step at a time & PATIENCE!

One thing that we need to know and have written down are the following:
1. Raidset/volumeset info - raid level, stripe size, # of disk.
2. What Firmware revision do you have?
3. Original Drive order - go into the Areca boot-time BIOS, "Raid Set Function" -> "Raid Set Information", and write down the "Member Disk Channels sequence" if it is available. This will aid greatly in the recovery process.
4. Do you have a backup of this data? Or a way to dump the disks sector-by-sector to somewhere else?
 
Hopefully some of the other Areca experts around here will pop in and offer their advice. They helped me w/ my recovery
 
Yea i will read that tomorrow when I am at work, but i wouldnt do anything in there just yet. The key to data recovery is to take everything one step at a time & PATIENCE!

One thing that we need to know and have written down are the following:
1. Raidset/volumeset info - raid level, stripe size, # of disk.
2. What Firmware revision do you have?
3. Original Drive order - go into the Areca boot-time BIOS, "Raid Set Function" -> "Raid Set Information", and write down the "Member Disk Channels sequence" if it is available. This will aid greatly in the recovery process.
4. Do you have a backup of this data? Or a way to dump the disks sector-by-sector to somewhere else?

Hi Nitrobass24,

Let me answer each individually (I also updated my first post a bit with some of this info):

1. RAID6, all 11x Hitachi 2TB drives were in one big RAID6 array (no hotspare). 128K stripe size, 16K clusters
2. Areca 1880i (FW version 1.48B, BIOS 1.22D) with an HP SAS expander (FW v2.02)
3. It says 'no RAID' set available unfortunately; argh. I am not sure where else I might be able to get this
4. Yes, I do have an offsite backup essentially as recent as Saturday morning. I do not at the moment have a way to dump the disks sector by sector but I do have said offsite backup.

Hopefully some of the other Areca experts around here will pop in and offer their advice. They helped me w/ my recovery
Interesting. It does look like your case was slightly different than mine. The only good news in your thread I see so far for my case is that someone said :
The "Rescue RAIDset" function is only for raidsets that don't show up, but yours shows up as failed, so that doesn't apply.

In my case that DOES apply. My RAID set/RAID volume does NOT show up now. Basically, it's like it had trouble communicating with a drive in slot 11, then because of this it thought the RAIDSet and Volume was degraded and the device was ejected, then it had an issue with another drive in slot 13, to which it again responded that the RAIDSet and Volume were degraded. At that point I had hard shut it down, and it somehow decided that the RAIDSet/Volume did not exist.
 
Last edited:
Privan-

The guys you want to talk to are odditory or jus, you can find them here or at forums.2cpu.com in the storage forums.
 
Privan-

The guys you want to talk to are odditory or jus, you can find them here or at forums.2cpu.com in the storage forums.

Hi Mwroobel,

Thank you for the heads up; I will send them a quick PM and beg nicely for them to take a look at this thread.

There is perhaps in some small way a real irony to the situation. I had previously ready nitrobass24's recover procedure and had even posted in his thread saying how much the recovery process he had to go through scared me given my knowledge of said advanced recovery tactics and I thought at the time "boy I hope I don't have to deal with that." And now, here I am! Luckily (haha), after reading his thread my situation may be less complicated (and I do have some quite recent backups) and in my case the rescue RAIDset might even work (though I am not quite clear the exact syntax and casing of the command based on the back in forth in that thread).
 
Last edited:
@pirivan

You do not happen to have battery backup unit attached to the controller, do you?

Whatever you do in order to recover your data, BBU must not be attached.
 
Last edited:
#1: Report back on what Kevin @ Areca says. My *guess* is they'll have you try a rescue raidset command or two, to write the meta/signnature back to the dropped drives.

#2: How hot is the HP SAS Expander heatsink when you touch it? If its too hot to touch then maybe airflow is a problem in that case and its overheating which I've seen cause those kind of drops/timeouts out of the blue - especially when its two drives at the same time like that. Nothing wrong with the drives. Best remedy is a 40x40x10mm fan mounted to heatsink.

#3: Resist the urge to panic, your data hasn't gone anywhere.

#4: Power down and remove BBU if one exists - as I think was going to be Jus' point.
 
Last edited:
As Odditory says, if it is not a problem with your expander, then recovering the missing RAIDset might be as simple as applying RESCUE and SIGNAT commands.

If I were you, I'd first update F/W to the latest as well as all other file sets. I've seen cases where it was sufficient to get RAIDset(s) back.
 
Last edited:
Thanks for all the replies Jus and Odditory! Let me address each point individually.

@pirivan

You do not happen to have battery backup unit attached to the controller, do you?
Whatever you do in order to recover your data, BBU must not be attached.

I indeed do not have a BBU at all, so it will not be attached.

#1: Report back on what Kevin @ Areca says. My *guess* is they'll have you try a rescue raidset command or two, to write the meta/signnature back to the dropped drives.

#2: How hot is the HP SAS Expander heatsink when you touch it? If its too hot to touch then maybe airflow is a problem in that case and its overheating which I've seen cause those kind of drops/timeouts out of the blue - especially when its two drives at the same time like that. Nothing wrong with the drives. Best remedy is a 40x40x10mm fan mounted to heatsink.

#3: Resist the urge to panic, your data hasn't gone anywhere.

#4: Power down and remove BBU if one exists - as I think was going to be Jus' point.

1. Excellent, I will let you know as soon as I hear from Areca. I think it is likely that they will have me run that command, it is my hope that the RAID card just dropped the configuration for some reason when it had an issue talking to the two drives
2. Unfortunately I can't precisely answer that. I turned off the server last night and when I was taking a look at it, I did have the cover off the Norco as well. However, even given all of that the server had been on for quite a few hours prior (longer than usual, a day or two so at least) so that is quite likely part of the issue. I will investigate getting some kind of a small fan ordered. Should I just zip tie it or something to the HP SAS Expander? Any fan suggestions?
3. I am trying! I actually have a Microsoft exam to study for/take tomorrow morning so I will have to resist working on this issue until tomorrow afternoon
4. Indeed, power is down, no BBU

@pirivan
As Odditory says, if it is not a problem with your expander, then recovering the missing RAIDset might be as simple as applying RESCUE and SIGNAT commands.

If I were you, I'd first update F/W to the latest as well as all other file sets. I've seen cases where it was sufficient to get RAIDset(s) back

Boy I hope that is indeed the case that a couple of commands might help me. How are those commands supposed to be entered? All caps? My guess is that the expander is still technically fine given that on reboots it saw all of the drives connected (they are connected to the expander after all) they just weren't part of any RAID set/volume. However, I think it is a good idea to get a fan or some kind of slot cooler installed for the SAS expander.

By update firmware do you just mean the 1.49 update here: http://www.areca.us/support/main.htm? I don't see any other update files there. I'll admit I would feel a little nervous about applying an FW update before running the recovery.
 
No reason to be nervous about apply a FW update to the card, it cant hurt your data. But if it really makes you uncomfortable you can always pull the drive trays out an inch or two so they are disconnected before you start up.
 
(in addition to what nitro said)
Furthermore, if you pull the drives out, update the F/W, and re-insert one drive at a time while the system is powered ON, the controller may recognize your lost RAIDset without entering RESCUE and SIGNAT commands.

Here is where you could get all the file sets for F/W/ upgrade.
ftp://ftp.areca.com.tw/RaidCards/BIOS_Firmware/ARC1880/149-20101210/

Flash them in the following order (NO reboots in between):
FIRM
BIOS
BOOT
MBR

Once applied, reboot the system.

F/W 1.49 has HP SAS expander card 2.02 support added, so I'd highly recommend getting the F/W flashed before you do anything else.
 
(in addition to what nitro said)
Furthermore, if you pull the drives out, update the F/W, and re-insert one drive at a time while the system is powered ON, the controller may recognize your lost RAIDset without entering RESCUE and SIGNAT commands.

Here is where you could get all the file sets for F/W/ upgrade.
ftp://ftp.areca.com.tw/RaidCards/BIOS_Firmware/ARC1880/149-20101210/

Flash them in the following order (NO reboots in between):
FIRM
BIOS
BOOT
MBR

Once applied, reboot the system.

F/W 1.49 has HP SAS expander card 2.02 support added, so I'd highly recommend getting the F/W flashed before you do anything else.

Hi Jus,

Thanks for the link! I think I will give that a shot. As far as the flashing goes, I assume that I can flash all of the 4 items above from the Web GUI (the web GUI built on to the card, not the ARCHttp environment) the 'Fimware Update' PDF isn't terribly specific? I assume that I can just re-insert the drives in the order that they are physically in the server or do they need to be re-inserted in some sort of specific order? The only issue is I can't get the "Member Disk Channels sequence" from the RAID controller (it doesn't show that in the RAID controller BIOS because it doesn't see a RAID set/volume anymore) if I need to put them in that order.
 
Last edited:
Yes you can do it from the WebGui.

As for the drive order, dont change it from what it is currently....for now.
 
Yes, use Web GUI. Flash all 4 file sets in the order I specified. Once done, reboot the system and start re-inserting the drives in their sequential order of ports on the expander. Frankly, I have found it to be irrelevant, as long as you give the controller enough time to detect each drive - wait for each drive to blink several times and after the LED finally stops blinking, proceed to next drive.
 
Please run "LeVeL2ReScUe" first and check the data if the raidset comes back after a power cycle. Send me an entire event log and hierarchy via email then. Do not issue "SIGNAT" until we can confirm the data is back.

You have a couple of timeout errors on slot #11 and #13. It could be caused by the controller, SAS expander, backplane and cabling. You'll need to isolate the problem later. Also, HDS722020ALA330 is a desktop version hard disk drive which is not recommended for raid array.

There is no order for updating all four BIN files of the firmware. You may do it after the raid set is back.

So, above is the reply I received from 'Benjamin' at Areca. I haven't done anything yet (updated firmware, ejected drives etc, I was sort of waiting to hear what Areca might say as well before I did anything). Should I try updating firmware first with drives ejected OR running the LeVel2ReScUe command as he suggested? Also, if I run a LeVel2ReScUe and then do a power cycle, won't I lose the results of that command? I thought it was sort of a 'read only' command that would be lost if you didn't 'save' it essentially with 'SIGNAT' (I could be completely wrong).
 
SIGNAT only regenerates the RAIDset signature. So, leave it for later.

Most likely Benjamin assumed you had latest F/W on your controller. So, get to it first. Next, try what Benjamin suggests.
 
Oh boy, well things got interesting. So, I removed all the drives, booted up, updated the firmware, rebooted and then 1 by 1 installed the drives. I did not change which drive went into which drive bay, I simply took them out a little bit and then pushed them back in. It started beeping after I started inserting drives but I kept going. Now it appears to have re detected the array but it looks very bizarre and appears to think there are multiple versions of my RAID6 array (not feeling so good now), what do I do now? It shows each of the RAID6 sets as 'incomplete' and it thinks that they each should have 11 2TB drives. So basically it is 3 versions of my RAID config but with the disks all spread out between them. The server is still on, I have not powered it off or rebooted since doing this.

93byR.jpg

7MDtc.jpg
 
Last edited:
Okay I ran LeVeL2ReScUe and then rebooted. It told me during the controller BIOS boot process that some RAID volumes had failed but I ignored that and continued to let it boot.

However, once it booted things are looking up! Only 1 RAID6 volume and it thinks it is OK

uHDHU.jpg
 
Now, try accessing it in OS and find a relatively big file, a movie or something. See that it is in one piece, I mean not corrupted... Check things randomly. The more you check the better.
 
Now, try accessing it in OS and find a relatively big file, a movie or something. See that it is in one piece, I mean not corrupted... Check things randomly. The more you check the better.

Hmm it does not currently appear in the OS for access (not in disk management, no errors in the event log). The Areca controller is listed under storage controllers but there is only my OS hard drive listed under disk drives. Should I reboot again or will that clear out the results of the "LeVeL2ReScUe" command?
 
Go to RAID Set functions and choose "Activate incomplete RAID set". Then go back to Windows and re-scan the drives.
 
Go to RAID Set functions and choose "Activate incomplete RAID set". Then go back to Windows and re-scan the drives.

Just to clarify, still do this even though the RAID6 set no longer shows as incomplete (it shows as normal with all 11 drives in 1 RAID6 array)?
 
Yes, even though the controller shows it as Normal, it may still be inactive.
 
OK, now reboot it one more time and observe the state of the RAID set. Once booted into OS you should be able to access the data.
 
Hmm, okay after rebooting the RAID6 set is not back in the OS and the web UI is back to where I was to start with; it show no RAID6 volume or array in the RaidSet Hierarchy and it shows all of the drives as 'free'. Is this because I have now rebooted after running the "LeVeL2ReScUe" command?
 
Run LeVeL2ReScUe again and reboot. When you get it back and it is reported in Normal state, use SIGNAT, then reboot again.
 
Run LeVeL2ReScUe again and reboot. When you get it back and it is reported in Normal state, use SIGNAT, then reboot again.

Okay I ran LeVeL2ReScUe and then rebooted. It told me during the controller BIOS boot process that some RAID volumes had failed but I ignored that and continued to let it boot.

After the reboot process finished it display 1 RAID6 volume and it thinks it is OK.

I then ran the SIGNAT command and rebooted again (no errors reported during the controller BIOS boot process).

Now the drive shows up in Windows and is accessible. I tested copying an 18GB file to the C: drive and it copied and then played without any issues. Should I run a volume check on the Areca controller or a chkdsk? Or should I just copy everything I had backed up over from my backups?

I will keep copying/testing files as well.
 
OK, a good news for a change! :)

Now, first run the volume set check, but for the first run UNcheck the scrubbing and parity recalculation.
Once completed, do the same, but with both options checked. If it completes with no errors then consider the data as completely recovered. Still, additional checks (such as chkdsk) would not hurt.
Of course, running the check twice takes time, I know, but if the first round starts spitting out errors (check event log periodically), then would be no need to run the second round with scrubbing.

Edit: OR, if you do have a good backup, run the check with both scrubbing and parity re-calculation and in case if you see errors, restore your data from your last good backup.
 
Last edited:
Restoring from backup is unnecessary. By the way were the drives spun down when the timeouts occurred? If so then even more reason for there to be nothing to worry about. But as Jus recommended run a scrub.
 
OK, a good news for a change! :)

Now, first run the volume set check, but for the first run UNcheck the scrubbing and parity recalculation.
Once completed, do the same, but with both options checked. If it completes with no errors then consider the data as completely recovered. Still, additional checks (such as chkdsk) would not hurt.
Of course, running the check twice takes time, I know, but if the first round starts spitting out errors (check event log periodically), then would be no need to run the second round with scrubbing.

Edit: OR, if you do have a good backup, run the check with both scrubbing and parity re-calculation and in case if you see errors, restore your data from your last good backup.

Restoring from backup is unnecessary. By the way were the drives spun down when the timeouts occurred? If so then even more reason for there to be nothing to worry about. But as Jus recommended run a scrub.

Yeah; I am glad things are starting to go in the right direction. I really can't thank you enough (both Jus and Odditory, extremely helpful as per usual); I genuinely appreciate it.

So far out of the 70 or so files I have opened/played nothing error'ed or wouldn't open etc.

I decided to go the 'safest' route and scheduled a volume consistency check with both options unchecked. I do have a good backup but it doesn't have absolutely EVERYTHING in it (but almost). I will run one with both options if it completes without any errors and then probably run a chkdsk. So, it will be a day and a half or so until that all finishes I expect (the volume checks aren't exactly quick. it is at 1% now).

As an aside, while I wait for that, in order to put myself in a better situation in case this sort of thing occurs in the future, what should I document about my RAID controller setup? Just take screenshots of all the controller configuration in the UI? Write down/take pictures of the Areca, BIOS member disk configuration?

Also, I did notice that the HP SAS expander heatsink is indeed fairly hot to the touch. Anyone have suggestions for a 40x40x10mm to secure to the heatsink? Also what would be a good (read safe) way to secure it on there? Or would it be better to try to purchase a slot based cooler? I have an open PCI slot two slots away (there is only a very tiny PCI-E E-SATA card in between the PCI slot and the SAS expander) that I could put a slot based cooler into. Any recommendations there?

For now maybe I will just blow an external fan or something onto it with the case open.

EDIT: @Odditory.

The drives were not spun down I don't believe. I went out for the evening for 45 minutes or so, and when I left I was updating a database for "Moving Pictures (a plugin) for an application called "MediaPortal". It is possible that that finished before the timeouts occurred and the drives spun down but it is also likely that the timeouts occurred while that was going on and thus the drives were not spun down. Either way, I am running a check right now with both options unchecked, the a run with both checked and possibly a chkdsk on the volume

-------------------EDIT/UPDATE:---------------------

Hmm well during the volume consistency check it is now at 7.3% and it has found two errors. I am not sure how concerning that is. Should I continue ahead with the plan and run it with the scrubbing and parity recalculation options checked once this run finishes? Or should I prepare to restore data from backups?
 
Last edited:
The consistency check is now at 10.7% and it has found 6 errors. For clarity's sake as far as the volume check goes I will just update this post as it goes along and post again separately when it is done. I am not sure if 6 errors for 10% is concerning or not. I assume that is what the "scrubbing and parity recalculation" options are for, to fix said items? I was hoping it wouldn't find any given that it never has before but I guess you can't always get that lucky.
 
Back
Top