• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Problem with Areca ARC-1880

jbsmith22

n00b
Joined
Jul 28, 2012
Messages
8
Hi folks,

I have two questions and I'll keep 'em short :)

1) Has anyone experienced an ARC-1880 dropping an entire enclosure? System (running in a Norco 4220 chassis) has been running steady state for a year with zero problems but the first failure occurred when I upped the total disk count from 13 to 15 and was expanding my primary array from 9 to 10 disks (RAID-6). Now it just happened again, not sure what the system was doing which leads to my second question...

2) After the failure during the expansion, my primary array is in a "Failed Migration" state. After some jimmying around (I didn't do much, pulled the new drives, swapped them around, etc), the last time I rebooted the server last night I got an entry in the event log that said "Rebuild RaidSet" and nothing else for around 18 hours. During that time, the Raid Set Hierarchy page continued to say "Failed Migration" for the state (in other words, it did NOT change to "Rebuilding (xx%)" so I don't know if it was actually rebuilding. Is there any way to tell if it was actually rebuilding the RaidSet or was I leaving it along and was it just sitting there? And the corollary question: is there any way to force a rebuild when the state is semi-hung like that? I mean it only got like 10% into rebuilding and since it's a RAID-6 array, failure during expansion should be immediately reversible shouldn't it? But yet no magic command that I could find to do so.

I also sent an email to Areca support, but since it's the weekend, thought I might get some helpful ideas here somewhat faster.

Annotated event log posted below for reference. Any more data that could be helpful, I can provide as well.

Thanks everyone!
Jason


Raid Set # 000/ARC-1880-VOL#000 - primary array, RAID 6, was 9 matching Samsung drives, added 1 to the array at the start of this trouble, kept an 11th in reserve as a hot spare, Enclosure #2 (not sure what enclosure 1 is), Originally Slots 05, 09-16, added drives in Slots 06-07, designated one for expansion, one for hot spare. Have rearranged those two drives ONLY since.


TVRecordings/TVRecordinVolSet - secondary array, RAID 0+1, four matching Toshiba drives, Enclosure #2 (not sure what enclosure 1 is), Slots 01-04.

2012-07-29 14:37:58 192.168.135.053 HTTP Log In
2012-07-29 14:10:16 TVRecordinVolSet Start Rebuilding
2012-07-29 14:10:14 TVRecordings Rebuild RaidSet

**** It doesn't show up here but I declared the drive in Slot 1 as a hot spare so the secondary array grabbed it and started rebuilding. Note that it didn't try to start rebuilding the primary array this time.

2012-07-29 14:08:43 Proxy Or Inband HTTP Log In
2012-07-29 14:04:33 H/W Monitor Raid Powered On
2012-07-29 14:00:41 192.168.135.053 HTTP Log In

**** Rebooted server after enclosure disappeared

2012-07-29 14:00:15 Enc#2 SLOT 03 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 02 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 15 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 14 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 13 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 12 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 11 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 10 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 09 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 07 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 05 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 04 Device Removed
2012-07-29 14:00:15 TVRecordings RaidSet Degraded
2012-07-29 14:00:15 TVRecordinVolSet Volume Failed
2012-07-29 14:00:15 TVRecordings RaidSet Degraded
2012-07-29 14:00:15 TVRecordinVolSet Volume Failed
2012-07-29 14:00:15 Enc#2 SLOT 01 Device Removed
2012-07-29 14:00:15 Enc#2 SLOT 16 Device Removed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Failed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Degraded
2012-07-29 14:00:15 Enc#2 SLOT 06 Device Removed
2012-07-29 14:00:15 Raid Set # 000 RaidSet Degraded
2012-07-29 14:00:15 ARC-1880-VOL#000 Volume Degraded
2012-07-29 14:00:14 Enclosure#2 Removed

**** Entire enclosure disappeared and the drive located on the previously functioning secondary array vanished.

2012-07-29 14:00:02 Enc#2 SES2Device Time Out Error
2012-07-29 13:59:48 Enc#2 SES2Device Time Out Error
2012-07-29 13:59:17 Enc#2 SLOT 01 Device Failed
2012-07-29 13:59:17 TVRecordings RaidSet Degraded
2012-07-29 13:59:17 TVRecordinVolSet Volume Degraded
2012-07-29 13:54:42 Enc#2 SLOT 01 Reading Error
2012-07-29 13:54:24 Enc#2 SLOT 01 Reading Error
2012-07-29 13:53:58 Enc#2 SLOT 01 Time Out Error
2012-07-29 13:53:03 Enc#2 SLOT 01 Time Out Error
2012-07-29 13:50:54 Enc#2 SLOT 01 Reading Error
2012-07-29 13:44:08 Enc#2 SLOT 01 Reading Error

**** Was streaming video off the secondary array (watching a TV show) which started stuttering badly until it finally failed

2012-07-28 19:59:09 Raid Set # 000 Rebuild RaidSet
2012-07-28 19:59:09 Enc#2 SLOT 07 Device Inserted

**** Slid the drive in Slot 07 back into the chassis and the primary array picked it up and claimed to start rebuilding. However, no sign of activity and no change to the summary webpage which continued to just say "Failed Migration" so no way to know if it was really rebuilding the array.

2012-07-28 19:16:41 192.168.135.053 HTTP Log In
2012-07-28 18:57:20 H/W Monitor Raid Powered On
2012-07-28 18:53:38 192.168.135.053 HTTP Log In
2012-07-28 18:37:59 RS232 Terminal VT100 Log In
2012-07-28 18:37:25 H/W Monitor Raid Powered On
2012-07-28 18:32:12 Enc#2 SLOT 08 Device Removed
2012-07-28 18:32:11 Raid Set # 000 RaidSet Degraded
2012-07-28 18:32:11 ARC-1880-VOL#000 Volume Degraded
2012-07-28 18:29:49 Enc#2 SLOT 08 Time Out Error
2012-07-28 18:27:14 Enc#2 SLOT 06 PassThrough Disk Deleted
2012-07-28 18:21:23 Proxy Or Inband HTTP Log In
2012-07-28 18:19:28 H/W Monitor Raid Powered On
2012-07-28 18:18:16 Enc#2 SLOT 06 PassThrough Disk Created
2012-07-28 18:16:28 RS232 Terminal VT100 Log In
2012-07-28 18:16:00 H/W Monitor Raid Powered On

**** I decided I would try to reformat the drive that was originally a hot spare to clear any RAID information off it in the hope that it would use it to rebuild the primary array. Alas, it didn't seem to make any difference, other than I stopped getting Time Out Errors on it.

2012-07-28 16:14:32 Enc#2 SLOT 06 Time Out Error
2012-07-28 16:14:12 Enc#2 SLOT 08 Time Out Error
2012-07-28 16:13:59 192.168.135.053 HTTP Log In
2012-07-28 14:13:12 Raid Set # 000 Rebuild RaidSet
2012-07-28 14:13:12 Enc#2 SLOT 08 Device Inserted
2012-07-28 14:02:26 H/W Monitor Raid Powered On
2012-07-28 14:00:18 RS232 Terminal VT100 Log In
2012-07-28 13:59:34 H/W Monitor Raid Powered On
2012-07-28 13:54:09 192.168.135.053 HTTP Log In
2012-07-28 13:53:36 RS232 Terminal VT100 Log In
2012-07-28 13:53:18 H/W Monitor Raid Powered On
2012-07-28 13:49:25 192.168.135.053 HTTP Log In
2012-07-28 13:47:43 Enc#2 SLOT 08 Device Removed
2012-07-28 13:47:43 Raid Set # 000 RaidSet Degraded
2012-07-28 13:47:43 ARC-1880-VOL#000 Volume Degraded
2012-07-28 13:46:39 RS232 Terminal VT100 Log In
2012-07-28 13:45:46 H/W Monitor Raid Powered On
2012-07-28 13:29:11 Enc#2 SLOT 08 Time Out Error
2012-07-28 13:29:01 Enc#2 SLOT 06 Time Out Error
2012-07-28 12:42:54 H/W Monitor Raid Powered On

**** Started to do some experiments with pulling out and pushing in the two new drives. Also experimented with the command line and the BIOS setup to see if it had any additional capabilities.

2012-07-28 10:48:12 Enc#2 SLOT 08 Time Out Error
2012-07-28 10:48:02 Enc#2 SLOT 06 Time Out Error
2012-07-28 10:46:18 SW API Interface API Log In
2012-07-28 10:11:00 192.168.135.053 HTTP Log In
2012-07-28 09:55:29 H/W Monitor Raid Powered On
2012-07-27 22:59:21 192.168.135.053 HTTP Log In
2012-07-27 22:57:59 H/W Monitor Raid Powered On
2012-07-27 22:42:19 192.168.135.053 HTTP Log In
2012-07-27 22:27:11 H/W Monitor Raid Powered On
2012-07-27 22:24:47 RS232 Terminal VT100 Log In
2012-07-27 22:24:32 H/W Monitor Raid Powered On
2012-07-27 22:19:48 192.168.135.053 HTTP Log In
2012-07-27 22:17:38 H/W Monitor Raid Powered On
2012-07-27 22:14:48 TVRecordinVolSet Modify Volume

**** Renamed my secondary array to better be able to tell it apart from the primary

2012-07-27 22:10:13 RS232 Terminal VT100 Log In
2012-07-27 22:09:27 H/W Monitor Raid Powered On
2012-07-27 22:06:45 192.168.135.053 HTTP Log In
2012-07-27 21:31:36 ARC-1880-VOL#001 Complete Rebuild 003:32:34
2012-07-27 21:06:15 Raid Set # 000 Offlined
2012-07-27 20:50:43 192.168.135.141 HTTP Log In
2012-07-27 18:19:55 Enc#2 SLOT 06 Device Inserted
2012-07-27 18:19:43 Raid Set # 000 Rebuild RaidSet
2012-07-27 18:19:42 Enc#2 SLOT 08 Device Inserted
2012-07-27 18:19:42 Enc#2 SLOT 07 Device Removed
2012-07-27 18:19:42 Raid Set # 000 RaidSet Degraded
2012-07-27 18:19:42 ARC-1880-VOL#000 Volume Degraded
2012-07-27 18:12:13 Enc#2 SLOT 07 Time Out Error
2012-07-27 18:00:40 Raid Set # 000 Rebuild RaidSet

*** Notice that it claims to be building the primary raidset but it is getting timeouts from the new disks. So I decide to try moving them to a different slot to see if it will help.

2012-07-27 18:00:40 Enc#2 SLOT 07 Device Inserted
2012-07-27 17:59:01 ARC-1880-VOL#001 Start Rebuilding
2012-07-27 17:58:59 Enc#2 SLOT 07 Device Removed
2012-07-27 17:58:59 TVRecordings Rebuild RaidSet
2012-07-27 17:58:59 TVRecordings RaidSet Degraded
2012-07-27 17:58:59 ARC-1880-VOL#001 Volume Degraded
2012-07-27 17:58:52 Enc#2 SLOT 06 Device Removed
2012-07-27 17:58:52 Raid Set # 000 RaidSet Degraded
2012-07-27 17:58:52 ARC-1880-VOL#000 Volume Degraded
2012-07-27 17:56:36 Enc#2 SLOT 07 Time Out Error
2012-07-27 17:56:23 192.168.135.053 HTTP Log In
2012-07-27 17:46:13 H/W Monitor Raid Powered On
2012-07-27 17:17:45 ARC-1880-VOL#001 Complete Rebuild 003:29:38
2012-07-27 16:48:47 Proxy Or Inband HTTP Log In
2012-07-27 16:31:01 Enc#2 SLOT 06 Time Out Error
2012-07-27 13:48:07 ARC-1880-VOL#001 Start Rebuilding
2012-07-27 13:48:05 TVRecordings Rebuild RaidSet

**** Since I didn't want the samsung used in the secondary array, I purposefully fail it by pulling the drive and setup the original hitachi drive as a hot spare so that it rebuilds itself successfully. Now I'm back to a normally functioning secondary array, and a broken primary array but at least I have my expansion and hot spare disks back in place.

2012-07-27 13:48:05 Enc#2 SLOT 07 Device Inserted
2012-07-27 13:47:52 Raid Set # 000 Rebuild RaidSet
2012-07-27 13:47:52 Enc#2 SLOT 06 Device Inserted
2012-07-27 13:46:55 Enc#2 SLOT 07 Device Removed
2012-07-27 13:46:55 TVRecordings RaidSet Degraded
2012-07-27 13:46:55 ARC-1880-VOL#001 Volume Degraded
2012-07-27 13:46:50 Enc#2 SLOT 06 Device Removed
2012-07-27 13:46:50 Raid Set # 000 RaidSet Degraded
2012-07-27 13:46:50 ARC-1880-VOL#000 Volume Degraded
2012-07-27 13:38:30 192.168.135.053 HTTP Log In
2012-07-27 13:20:39 H/W Monitor Raid Powered On
2012-07-27 13:17:43 RS232 Terminal VT100 Log In
2012-07-27 13:03:37 RS232 Terminal VT100 Log In
2012-07-27 13:03:05 H/W Monitor Raid Powered On
2012-07-27 12:51:12 H/W Monitor Raid Powered On
2012-07-27 12:39:22 H/W Monitor Raid Powered On
2012-07-27 12:37:23 H/W Monitor Raid Powered On
2012-07-27 12:21:44 Enc#2 SLOT 06 Time Out Error
2012-07-27 12:21:32 192.168.135.141 HTTP Log In
2012-07-27 12:17:00 TVRecordings Offlined
2012-07-27 12:14:18 192.168.135.141 HTTP Log In

**** Now I notice that the rebuild/expansion of the primary array isn't starting back up so I panic and start pulling out the new drives and trying different reboot strategies, none of which work.

2012-07-27 02:50:17 ARC-1880-VOL#001 Complete Rebuild 003:29:53
2012-07-26 23:20:23 ARC-1880-VOL#001 Start Rebuilding

***** The secondary array is able to rebuild itself onto the hot spare successfully (though undesirably since I want to use the large Samsung disks for the primary array only)

2012-07-26 23:19:21 192.168.135.053 HTTP Log In
2012-07-26 23:19:16 H/W Monitor Raid Powered On

**** The machine crashes spontaneously and I notice the problem and go and restart it.

2012-07-26 23:10:43 Enc#2 SLOT 03 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 02 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 15 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 14 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 13 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 12 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 11 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 10 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 09 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 07 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 06 Device Removed
2012-07-26 23:10:43 Enc#2 SLOT 05 Device Removed
2012-07-26 23:10:42 ARC-1880-VOL#000 Abort Migration 006:16:35
2012-07-26 23:10:42 ARC-1880-VOL#001 Abort Rebuilding 000:00:07

**** Now the system gives up the migration of the primary array and the rebuild of the secondary array because it can't find any drives left in the system. Of course they have not been touched by me and are still safely snug in the chassis running normally.

2012-07-26 23:10:42 Enc#2 SLOT 01 Device Removed
2012-07-26 23:10:42 Enc#2 SLOT 04 Device Removed
2012-07-26 23:10:42 TVRecordings RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#001 Volume Failed
2012-07-26 23:10:42 TVRecordings RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#001 Volume Failed
2012-07-26 23:10:42 Enc#2 SLOT 16 Device Removed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Failed
2012-07-26 23:10:42 TVRecordings RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#001 Volume Degraded
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Degraded
2012-07-26 23:10:42 Raid Set # 000 RaidSet Degraded
2012-07-26 23:10:42 ARC-1880-VOL#000 Volume Degraded
2012-07-26 23:10:42 ARC-1880-VOL#001 Start Rebuilding
2012-07-26 23:10:42 Enc#2 SLOT 01 Device Failed

**** At this point the entire enclosure seems to disappear causing all kinds of problems with both Raidsets

2012-07-26 23:10:42 Enclosure#2 Removed
2012-07-26 23:10:33 TVRecordings Rebuild RaidSet
2012-07-26 23:10:24 TVRecordings RaidSet Degraded
2012-07-26 23:10:24 ARC-1880-VOL#001 Volume Degraded

**** It decides the secondary array is broken and grabs my hot spare (which was not at the time reserved to a raidset) and kicks off a rebuild.

2012-07-26 23:10:12 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:09:58 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:09:40 Enc#2 SLOT 01 Reading Error
2012-07-26 23:06:09 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:05:52 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:05:16 Enc#2 SLOT 01 Reading Error
2012-07-26 23:04:50 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:04:31 Enc#2 SLOT 01 Time Out Error
2012-07-26 23:04:21 Enc#2 SLOT 01 Reading Error

**** After 7 hours of expansion, the trouble begins. It says the problem is with Slot 01 but I suspect it's with the entire array and Slot 01 is just the first drive on the list.

2012-07-26 16:54:07 ARC-1880-VOL#000 Start Migrating
2012-07-26 16:54:05 Raid Set # 000 Expand RaidSet

**** Start the expansion onto one of them and declared the other a hot spare

2012-07-26 16:49:24 Enc#2 SLOT 06 Device Inserted
2012-07-26 16:49:17 Enc#2 SLOT 07 Device Inserted

**** Inserting the two new Samsung drives

2012-07-23 01:27:37 192.168.135.053 HTTP Log In
2012-07-22 01:01:00 H/W Monitor Raid Powered On
2012-07-18 16:35:03 H/W Monitor Raid Powered On
2012-06-24 22:27:12 H/W Monitor Raid Powered On
2012-06-11 09:35:44 H/W Monitor Raid Powered On
2012-04-08 22:49:57 H/W Monitor Raid Powered On
2012-04-08 21:37:52 Proxy Or Inband HTTP Log In
2012-03-07 14:03:20 H/W Monitor Raid Powered On
2012-03-07 12:17:23 H/W Monitor Raid Powered On
2012-03-07 12:02:17 H/W Monitor Raid Powered On
2012-02-17 10:44:51 Enc#2 SLOT 01 Reading Error
2012-01-23 09:47:11 Enc#2 SLOT 16 Time Out Error
2012-01-23 08:07:44 H/W Monitor Raid Powered On
2011-11-30 16:45:49 ARC-1880-VOL#000 Complete Init 006:29:08
2011-11-30 10:16:41 ARC-1880-VOL#000 Start Initialize
2011-11-30 10:16:39 ARC-1880-VOL#000 Modify Volume
2011-11-28 09:56:48 ARC-1880-VOL#000 Complete Migrate 111:04:05
2011-11-23 18:52:44 ARC-1880-VOL#000 Start Migrating
2011-11-23 18:52:42 Raid Set # 000 Expand RaidSet
2011-11-23 19:31:04 Proxy Or Inband HTTP Log In
2011-10-30 12:28:14 H/W Monitor Raid Powered On
2011-09-29 19:12:48 Proxy Or Inband HTTP Log In
2011-09-29 19:11:40 Proxy Or Inband HTTP Log In
2011-09-06 09:46:30 H/W Monitor Raid Powered On
2011-08-28 14:41:52 H/W Monitor Raid Powered On
2011-08-15 20:06:45 Proxy Or Inband HTTP Log In
2011-08-12 22:59:41 H/W Monitor Raid Powered On
2011-08-11 18:14:05 H/W Monitor Raid Powered On
2011-08-11 17:33:53 H/W Monitor Raid Powered On
2011-08-03 12:11:47 H/W Monitor Raid Powered On
2011-07-18 22:08:09 H/W Monitor Raid Powered On
2011-07-18 22:03:07 192.168.135.121 HTTP Log In
2011-07-02 00:22:42 192.168.135.121 HTTP Log In
2011-06-21 11:33:39 H/W Monitor Raid Powered On
 
Which 1880 do you have? (Do you have a high port count 1880 or are you using an expander? Generally, for internal-multi-cable enclosures like the norco, if every drive drops out at once it is either a power problem (most likely), a bad card (less likely) or if you are using a basic 1880i with an expander then the cable between the HBA and the expander (or the expander). In any case, I don't have time right now to parse through the logs (stuck at work, will look more later) send a private message to Jus and odditory to ask for help and reference this thread. Please also post screenshots of your RAID config screen from the card.
 
Thanks mwroobel, will PM jus and/or odditory a little later (assuming they don't jump on this before I get there :)

The card I have is this one:
Areca ARC-1880IX-16 16-Port Full Height PCI-E x8 1GB SAS/SATA RAID Controller with breakout cables included

Motherboard: SUPERMICRO MBD-X9SCA-F-O LGA 1155 Intel C204 ATX Intel Xeon E3 Server Motherboard

OS Drive: Kingston SSDNow V+100 SVP100S2B/96GR 2.5" 96GB SATA II MLC Internal Solid State Drive (SSD)

Memory: Crucial CT2KIT51272BA1339 - 8 GB (2 x 4 GB) - 1333 MHz DDR3-1333/PC3-10600 - ECC - Unbuffered - 240-pin DIMM

Power Supply: CORSAIR CMPSU-750TX ATX 12v v2.2 / eps 12v v 2.91 750w ul fcc power supply-Retail

CPU: Intel BX80623E31235 Xeon E3-1235 3.20 GHz Processor - Socket H2 LGA-1155 - Quad-core - 8 MB Cache

And the aforementioned Norco Case: NORCO RPC-4220 4U Rackmount Server Case - Retail

Thanks!
Jason
 
Ok, since you are using the 1880 with the onboard expander, and all 16 drives dropped it is most likely either a bad card or a power problem. First thing I would try is a different set of power connectors off a different power tree from the PSU. I know you said you switched around the drives you added after you started the expansion.. Do you remember which drive was where before you switched them?
 
I can try a different power tree - I'm not near the server today, but as I recall, there were two power ports and I connected them both, but I will check tomorrow when I'm home. Of course, the dropouts are not reproducible. It has occurred twice now, but until I can force the primary array to continue or reverse the expansion, I can't really do much to test it. I suppose I could load it by failing one of the drives on the secondary array, but it has already rebuilt that array three or four times and it has never failed during THOSE rebuilds.

I am 90% certain that the drives are in the same positions they were in when I first tried the expansion. There's only one other position they could have been in (they could have been flipped) but I'm pretty sure they are in the same position. The array recognizes the one in slot 6 as the expansion drive and the one in slot 7 as the hot spare dedicated to the primary array.
 
Okay, so just a quick update - I changed to a different power chain, but only after it failed again. I THINK it's a reproducible problem. At least I was trying to watch the same recorded TV show when it went belly up. This time the drive in slot 1 showed a SMART failure before it disconnected the whole enclosure. So I switched to a different power chain AND replaced the HD in slot 01 and it is rebuilding the array. Once it finishes, I'll try watching the show again to see if it disconnects the whole enclosure or not. If it does, I'll start thinking that the problem is related to either the card or the content on the secondary array, so more testing to do there.

The primary array is still frozen in place - well actually I accidentally made it worse. I pulled the wrong drive (who knew "slot 1" was in the bottom right not the top left?) which caused it to show as failed. I shut the machine down, pulled all of the drives across all the arrays, shut it back down, pushed them back in, and then restarted the server hoping that the controller would recognize that the drive was not actually failed but was back. Buuuut, that didn't work. Now it lists that drive as "failed" in the Raidset hierarchy, and "free" in the enclosure listing. So I have two hot spares declared (one the original one that it was trying to expand to, and one that was always a hot spare) and one "free" drive which was a part of the original array that I pulled by accident. Absolutely no changes to the Raidset and Volumeset though. Still shows it stuck in "migrating" and on the web page it still says "failed migration."

Any suggestions for how to get the primary array to rebuild would be much appreciated. I PMed Jus and Odditory, but no answers yet.

Thanks!
Jason
 
I'm short on time to really decipher this whole thing from the OP downward since its a bit overload but main thing is stop playing around and experimenting with drives and raidsets that contain live data when you're trying to pinpoint what's causing instability. First thing I always have to question when multple drives appear to fail at the same time is does the card have adequate airflow or is it possibly overheating. This is a classic problem when an expander chip gets too hot.

I'll try to digest the rest of this info tomorrow and hopefully have a few more ideas. Main thing is patience and taking it slow, so you dont make things worse.
 
Thanks odditory - I've been pretty careful after reading through a bunch of similarly themed threads about not mucking with anything too heavy. I DID erroneously pull a drive from the primary raid set and put it back in which caused it to forget it existed, but when I pulled the right drive and replaced it with a fresh drive, the general instability of the system seems to have stopped. The secondary array has been accessible and loaded with streaming media going on and coming off it for 36 hours with no further issues to the overall enclosure.

So HOPEFULLY the enclosure disappearing instability was caused by the failing drive on the secondary array and I'm no longer trying to diagnose general instability. What I AM left with is a primary array that I cannot bring back online because it is stuck in the expanding phase it was running when the enclosure failed the first time. Logically, even if the data is being redistributed across the array to account for the newly added drive, a failure should be recoverable since all of the data is supposed to be able to survive the failure of TWO disks and I was only adding ONE disk. Plus as someone pointed out, you are supposed to be able to kill an expansion job and have it exit gracefully. All of the disks (including the expansion disk) are sitting in their same spots as before and all of them have the same data on them they had at the moment of failure. Well there's a 10% possibility that I mussed the expansion drive in my initial panic at the failure before wisdom took over and beat a "go slow and document what you're doing" approach into my head.

Currently the array sees the expansion disk as a hot spare and the second expansion disk I was putting in specifically to BE a hot spare as a hot spare. It sees 8 of the 9 original drives in the array as normally functioning parts of the array (though the array itself is in the Failed Migration state and inaccessible). It sees the drive that I pulled by accident (I was unclear on where each slot was and the Norco can't flash the LEDs unfortunately) as a free drive. But I have not done ANYTHING to it and it's still just sitting in it's slot where it always has been.

I appreciate the help!
Jason
 
Back
Top