• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

ARC-1880 Raid Failure - looking for advice?

Joined
Mar 16, 2014
Messages
10
Hi All,

I have a WHS running the SAS version of the ARC-1880, been running well for 3+ years. Last week, I started having some issues with it, and 2 drives failed (RAID 5 - 7 drive array).

I read that 2 drives can fail due to the raid card overheating - so I shut it off overnight.

Started it up, and only 1 of the drives failed, so I replaced it with a new drive. Now it says that the RAID Array Failed (not the volume), and the ARC GUI just says "Rebuilding" with no percentage Status - AND - I don't see the lights blinking on the enclosure.

So - do I just be patient and give this thing 2-3 days to just run and see if it rebuilds? Shouldn't I see some status? Shouldn't I see the lights blinking some?

Thanks in advance
 
Please post a complete log from the card, and what firmware is your card at?
 
Thank you in advance.


firmware is:
Controller Name ARC-1880X
Firmware Version V1.52 2014-02-07
BOOT ROM Version V1.49 2010-12-10 (I updated this one, and it replied that it is at 1.52 - but I have not rebooted yet... SHOULD I?)
PL Firmware Version 18.0.0.0
Serial Number Y117CACPAR200129
Unit Serial #
Main Processor 800MHz PPC440
CPU ICache Size 32KBytes
CPU DCache Size 32KBytes/Write Back
System Memory 512MB/800MHz/ECC
PCI-E Link Status 8X/5G


Snap-Shot of Hierarchy:

RAID Set Devices Volume Set(Ch/Id/Lun) Volume State Capacity

Raid Set # 000

E#1Slot#1 ARC-1880-VOL#000(0/0/0) Failed 12002.4GB
E#1Slot#2
E#1Slot#8
E#1Slot#4
E#1Slot#5
E#1Slot#6
E#1Slot#7




Enclosure#1 : ARECA SAS RAID AdapterV1.0
Device Usage Capacity Model
Slot#1(E) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#2(C) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#3(B) Hot Spare 2000.4GB WDC WD20EARS-00MVWB0 [Raid Set # 000 ] (NOTE: This is a temporary replacement)
Slot#4(D) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#5(10) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#6(A) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#7(F) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0
Slot#8(9) Raid Set # 000 2000.4GB WDC WD2002FYPS-02W3B0



LOGS AS FAR AS THEY GO:

Time Device Event Type Elapse Time Errors
2014-03-16 19:15:39 Proxy Or Inband HTTP Log In
2014-03-16 18:15:00 Proxy Or Inband HTTP Log In
2014-03-16 14:24:51 Proxy Or Inband HTTP Log In
2014-03-16 14:24:12 Enc#1 Slot#5 Device Inserted
2014-03-16 14:23:53 Enc#1 Slot#7 Device Inserted
2014-03-16 14:23:50 Enc#1 Slot#1 Device Inserted
2014-03-16 14:23:49 Enc#1 Slot#4 Device Inserted
2014-03-16 14:23:47 Enc#1 Slot#2 Device Inserted
2014-03-16 14:23:44 Enc#1 Slot#3 Device Inserted
2014-03-16 14:23:25 Enc#1 Slot#6 Device Inserted
2014-03-16 14:23:19 Enc#1 Slot#8 Device Inserted
2014-03-16 14:20:18 Proxy Or Inband HTTP Log In
2014-03-16 14:16:03 H/W Monitor Raid Powered On
2014-03-16 14:09:35 Proxy Or Inband HTTP Log In
2014-03-16 14:04:03 Proxy Or Inband HTTP Log In
2014-03-16 11:24:33 H/W Monitor Raid Powered On
2014-03-16 11:23:17 RS232 Terminal VT100 Log In
2014-03-16 09:47:30 RS232 Terminal VT100 Log In
2014-03-15 23:43:23 RS232 Terminal VT100 Log In
2014-03-15 22:34:21 RS232 Terminal VT100 Log In
2014-03-15 21:25:03 RS232 Terminal VT100 Log In
2014-03-15 20:36:31 RS232 Terminal VT100 Log In
2014-03-15 19:51:17 RS232 Terminal VT100 Log In
2014-03-15 19:21:01 RS232 Terminal VT100 Log In
2014-03-15 18:08:02 RS232 Terminal VT100 Log In
2014-03-15 17:57:10 RS232 Terminal VT100 Log In
2014-03-15 17:56:48 H/W Monitor Raid Powered On
2014-03-15 17:45:36 Proxy Or Inband HTTP Log In
2014-03-15 17:20:35 Proxy Or Inband HTTP Log In
2014-03-15 17:13:33 H/W Monitor Raid Powered On
2014-03-15 16:43:50 Enc#1 Slot#3 Device Inserted
2014-03-15 16:41:55 Enc#1 Slot#3 Device Removed
*** NOTE: I installed a new drive in place of #3, and it is now the hot spare for this array ***

2014-03-15 16:35:16 Proxy Or Inband HTTP Log In
2014-03-15 16:24:56 Enc#1 Slot#3 Device Failed(SMART)
2014-03-15 16:24:56 H/W Monitor Raid Powered On
2014-03-15 16:21:26 H/W Monitor Raid Powered On
2014-03-15 16:19:39 RS232 Terminal VT100 Log In
2014-03-15 16:18:55 H/W Monitor Raid Powered On
2014-03-15 16:12:28 Enc#1 Slot#3 Device Inserted
2014-03-15 16:09:37 Raid Set # 000 Rebuild RaidSet
2014-03-15 16:09:37 Enc#1 Slot#2 Device Inserted
2014-03-15 16:09:26 Enc#1 Slot#3 Device Removed
2014-03-15 16:09:08 Enc#1 Slot#2 Device Removed
2014-03-15 16:09:08 Raid Set # 000 RaidSet Degraded
2014-03-15 16:09:08 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:06:11 Enc#1 Slot#6 Device Inserted
2014-03-15 16:06:11 Enc#1 Slot#7 Device Inserted
2014-03-15 16:06:09 Raid Set # 000 Rebuild RaidSet
2014-03-15 16:06:09 Enc#1 Slot#1 Device Inserted
2014-03-15 16:06:09 Enc#1 Slot#2 Device Inserted
2014-03-15 16:06:08 Enc#1 Slot#3 Device Failed(SMART)
2014-03-15 16:06:08 Enc#1 Slot#3 Device Inserted
2014-03-15 16:06:07 Enc#1 Slot#4 Device Inserted
2014-03-15 16:06:05 Enc#1 Slot#5 Device Inserted
2014-03-15 16:06:05 Enc#1 Slot#8 Device Inserted
2014-03-15 16:05:35 Enc#1 Slot#5 Device Removed
2014-03-15 16:05:35 Enc#1 Slot#6 Device Removed
2014-03-15 16:05:35 Enc#1 Slot#7 Device Removed
2014-03-15 16:05:35 Enc#1 Slot#4 Device Removed
2014-03-15 16:05:35 Enc#1 Slot#8 Device Removed
2014-03-15 16:05:34 Raid Set # 000 RaidSet Degraded
2014-03-15 16:05:34 Raid Set # 000 RaidSet Degraded
2014-03-15 16:05:34 Raid Set # 000 RaidSet Degraded
2014-03-15 16:05:34 Raid Set # 000 RaidSet Degraded
2014-03-15 16:05:34 Enc#1 Slot#1 Device Removed
2014-03-15 16:05:34 Enc#1 Slot#3 Device Removed
2014-03-15 16:05:33 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:05:32 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:05:32 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:05:31 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:05:31 Raid Set # 000 RaidSet Degraded
2014-03-15 16:05:31 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:04:23 Enc#1 Slot#3 Device Failed(SMART)
2014-03-15 16:04:23 Enc#1 Slot#3 Device Inserted
2014-03-15 16:04:13 Enc#1 Slot#2 Device Removed
2014-03-15 16:04:13 ARC-1880-VOL#000 Stop Rebuilding 000:03:07
2014-03-15 16:04:13 Raid Set # 000 RaidSet Degraded
2014-03-15 16:04:13 ARC-1880-VOL#000 Volume Failed
2014-03-15 16:04:01 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:56 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:51 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:46 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:42 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:37 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:32 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:28 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:23 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:18 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:16 Enc#1 Slot#3 Device Removed
2014-03-15 16:03:11 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:07 Enc#1 Slot#2 Reading Error
2014-03-15 16:03:02 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:57 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:53 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:47 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:39 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:34 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:30 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:25 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:20 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:15 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:11 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:06 Enc#1 Slot#2 Reading Error
2014-03-15 16:02:01 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:57 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:52 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:47 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:42 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:38 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:33 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:28 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:23 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:18 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:14 RS232 Terminal VT100 Log In
2014-03-15 16:01:12 Enc#1 Slot#2 Reading Error
2014-03-15 16:01:05 ARC-1880-VOL#000 Start Rebuilding
2014-03-15 16:00:53 Enc#1 Slot#3 Device Failed(SMART)
2014-03-15 16:00:53 Raid Set # 000 Rebuild RaidSet
2014-03-15 16:00:53 ARC-1880-VOL#000 Failed Volume Revived
2014-03-15 16:00:53 H/W Monitor Raid Powered On
2014-03-14 20:56:50 Proxy Or Inband HTTP Log In
2014-03-14 18:47:41 Enc#1 Slot#2 Device Failed
2014-03-14 18:47:41 ARC-1880-VOL#000 Stop Rebuilding 000:05:02
2014-03-14 18:47:41 Raid Set # 000 RaidSet Degraded
2014-03-14 18:47:41 ARC-1880-VOL#000 Volume Failed
2014-03-14 18:47:36 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:31 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:27 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:22 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:17 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:12 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:08 Enc#1 Slot#2 Reading Error
2014-03-14 18:47:03 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:58 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:53 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:49 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:44 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:39 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:34 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:29 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:25 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:20 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:15 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:10 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:05 Enc#1 Slot#2 Reading Error
2014-03-14 18:46:01 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:56 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:54 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:46 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:42 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:36 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:30 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:25 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:16 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:11 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:06 Enc#1 Slot#2 Reading Error
2014-03-14 18:45:01 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:56 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:52 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:47 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:42 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:37 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:31 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:26 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:21 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:16 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:12 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:06 Enc#1 Slot#2 Reading Error
2014-03-14 18:44:01 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:56 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:51 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:47 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:42 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:37 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:32 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:28 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:23 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:18 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:14 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:09 Enc#1 Slot#2 Reading Error
2014-03-14 18:43:04 Enc#1 Slot#2 Reading Error
2014-03-14 18:42:59 Enc#1 Slot#2 Reading Error
2014-03-14 18:42:55 Enc#1 Slot#2 Reading Error
2014-03-14 18:42:44 Enc#1 Slot#2 Reading Error
2014-03-14 18:42:38 ARC-1880-VOL#000 Start Rebuilding
2014-03-14 18:41:41 Enc#1 Slot#3 Device Failed(SMART)
2014-03-14 18:41:41 Raid Set # 000 Rebuild RaidSet
2014-03-14 18:41:41 ARC-1880-VOL#000 Failed Volume Revived
2014-03-14 18:41:41 H/W Monitor Raid Powered On
2014-03-14 18:35:54 Enc#1 Slot#2 Device Failed
2014-03-14 18:35:53 ARC-1880-VOL#000 Stop Rebuilding 000:06:14
2014-03-14 18:35:53 Raid Set # 000 RaidSet Degraded
2014-03-14 18:35:53 ARC-1880-VOL#000 Volume Failed
2014-03-14 18:35:49 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:44 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:39 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:34 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:30 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:25 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:20 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:15 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:11 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:06 Enc#1 Slot#2 Reading Error
2014-03-14 18:35:01 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:56 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:52 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:47 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:42 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:37 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:33 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:28 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:23 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:18 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:14 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:09 Enc#1 Slot#2 Reading Error
2014-03-14 18:34:04 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:59 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:55 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:50 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:45 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:40 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:35 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:31 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:26 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:21 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:16 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:12 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:12 Enc#1 Slot#2 Reading Error
2014-03-14 18:33:02 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:57 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:53 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:48 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:42 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:36 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:31 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:22 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:18 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:13 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:08 Enc#1 Slot#2 Reading Error
2014-03-14 18:32:03 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:59 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:54 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:49 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:45 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:40 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:35 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:30 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:26 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:21 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:16 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:11 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:05 Enc#1 Slot#2 Reading Error
2014-03-14 18:31:00 Enc#1 Slot#2 Reading Error
2014-03-14 18:30:55 Enc#1 Slot#2 Reading Error
2014-03-14 18:30:50 Enc#1 Slot#2 Reading Error
2014-03-14 18:30:46 Enc#1 Slot#2 Reading Error
 
Did you have this setup as a RAID5 with a hotspare initially? Do you have all the specs of your array (stripe size, drive order, etc?)
 
Hello mwroobel,

Actually - I have never moved the drives around. I RESEATED them in their same slot tho. (That is why you see 'inserted drive' a few times)

I do remember tho - that just before this happened - I was having some issues with disks #2 and #3.

So - I seem to recall that Disk #3 failed (SMART FAILED) some time ago, but the array was just degraded. So Disk #2 was a hot spare, and had not kicked-in. So when I logged-on, I made Disk #2 the spare, and the array came back online, and I thought completely re-built itself. HOWEVER - I am wondering if I was wrong there.

Question: Is there a way to REMOVE Disk #2 from the Array, and put disk #3 back in where it was before?

Since I have updated to the 1.52BIOS - I am now able to 'reactivate disk" - and I am able to reactivate Disk #3. There are smart errors, but it is working. Could I ACTIVATE #3 long enough to bring the array back online, and copy the data off??

Thoughts?
 
Unfortunately, from the log it seems you have lost 2 drives in an array that can only handle 1 loss. In addition, with a degraded array did you (while the machine was on and running) remove and reinsert each of the drives? Do you have a backup of the array?
 
Actually - I really only lost 1 drive. The other drive was due to the card overheating. It is actually back now.

I removed the drives while it was powered off, and then powered it all on, and the inserted all the drives. I read online that this was a way to get the controller to recognize the ARRAY, and make sure I did not have to re-build from scratch.

Ok - so let's say that only 1 drive failed. What would I do next at this point?
 
With all the slot #2 read errors that does not look like a card over-heat problem at all. I have never seen a problem cause a bunch of read errors on a single drive like that.

Unfortunately you don't have enough of an event log to see the full history due to all the read errors. The oldest entries I see are:

Code:
2014-03-14 18:41:41 Enc#1 Slot#3 Device Failed(SMART)
2014-03-14 18:41:41 Raid Set # 000 Rebuild RaidSet
2014-03-14 18:41:41 ARC-1880-VOL#000 Failed Volume Revived
2014-03-14 18:41:41 H/W Monitor Raid Powered On
2014-03-14 18:35:54 Enc#1 Slot#2 Device Failed
2014-03-14 18:35:53 ARC-1880-VOL#000 Stop Rebuilding 000:06:14
2014-03-14 18:35:53 Raid Set # 000 RaidSet Degraded
2014-03-14 18:35:53 ARC-1880-VOL#000 Volume Failed

So It looks like slot 3 was probably already failed (its being failed for SMART), slot 2 then failed after a bunch of read errors.

Slot 2 is the one disk you absolutely can't lose.

I would take that disk out and try to dd_rescue it to a new disk if you haven't already.

If you can get this machine booted on riplinux or something I would check smart status of the drives behind the controller to see the situation but its looking quite a bit like slot 2 is really failing to me, I have seen this happen on tons of arrays.

Just because the drive comes back online after reboots and stuff does not mean the drive is not having issues and won't fail again under any kind of load.

This right here is why I don't run raid5. I have recovered arrays on 30+ areca machines and looking at your event history I would not be optimistic.

Again what I would do in this situation is either 1) use linux and check smart behind the controller or pull slot 2 and put it in another computer and look at smart stats to see the status of the disk. I really don't trust that the drive is 'fine'

2) There must be another drive that was in the fail state at some point if what you pasted is accurate. A screenshot of the actual raid hierarchy page would be better just to see whats going on.

I suspect what happened is disk 8 was your hot spare and slot 3 failed (smart) and then disk 8 took its spot and now became logical disk 3 (going by your paste info this follows) and it was still rebuilding when slot 2 failed causing the entire array to be in the fail state. This is pretty much always when I see arrays fail because they have no redundancy during the rebuild and the disk its rebuilding from starts failing under the heavy load of the rebuild.

I would recover this by first notating all the raid volume set sizes (exactly), stripes, w etc....

1) do a ddrescue on slot 2 to a new drive on a seperate computer. I would not trust that drive.

2) leave disk order alone....

3) delete entire raid set.

4) re-create raid set using disks slot 1-7.

5) re-create volume sets using previously noted settings with the no init (rescue option).

6). Fail (through web-interface or CLI) slot 3 which should be a new blank disk or simply pull it from the array degrading it putting it back in the same state it was when it failed.

7) Boot linux and try mounting the file-system in read-only to see if data is accessible. I know of no way to do a read-only mount on windows.... You don't want to write anything to the array if its possibly inconsistent.

8) If data looks ok, remove/re-insert slot 3 and let it rebuild.

9) After its a normal state again add slot 8 (your original hot spare) to the raid set and change it from raid5 to raid6. Raid 6 is way more fault tolerant than raid5 + hot spare. Honestly don't know why anyone would ever do raid5 + hotspare over raid6.

I wouldn't really recommend doing any of the above unless you know what you are doing. If you want you can boot the machine into riplinux:

ISO is here:
http://www.tux.org/pub/people/kent-robotti/looplinux/rip/RIPLinuX-13.7.iso

run:

dhcpcd eth0
ifconfig eth0

(to see the IP it pulled from DHCP)

then:

passwd
/usr/sbin/sshd


(passwd prompts for a password to set a password, and second command starts ssh).

If you forward any port to that machines IP on port 22 and PM me with your IP, the port, and the password you set I am willing to take a look at the array and help you out in reviving it.

I literally did raid recoveries all the time on 3ware, lsi, and areca at my old job and between the three brands have recovered 50+ arrays so I have done this a lot. The biggest reason for having to do so many as instead of raid6 they were raid10 with the same problem of a single point of redundancy and sadly most of the disks were seagate non-enterprise disks and had CRAP reliability. Seriously those are the shittiest disks on the planet. This caused arrays to fail pretty much all the time.

EDIT:

Also you need to notate whether the raid set was created with 128 volume support or not. It should show this on the raid set information page from the web interface if it says:

Supported Volumes 128
 
Last edited:
Hi All - just wanted to give an update here. I finally received a response from the ARECA folks. They noted that drives 2,3,8 were all an issue - and since disk 3 failed initially - then disk 2 built itself back from a degraded raid - I had 2 disks that were 'capable' of bringing it back online.

So - WHILE the server was running. They had me pull drives 2,3 and 8. Then - reinsert them. I did a couple different comninations of these 3 drives and bingo! I was able to get the volume back on-line in a degraded state. I have been busy copying the data off.

UPDATE: Was working the first part of this week (I have about 85% of the data - and the MOST CRITICAL stuff I needed) and now I am running into disk failures, etc...

THANK YOU all that responded to help me - truly appreciate it!
 
Back
Top