• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Areca RAID5 Fail-- need help desperately please

I am really glad I saw this as I have a similar setup. Is there a workaround to having to "hot plug" an array?

I have a Server with a 1880x card that is connected to a Supermicro SC847-JBOD chassis. Because the Supermicro chassis is so loud, I only turn it on when I need to access the array. I think this is similar to "hot plugging" since the Server is generally already running when I turn on the Supermicro chassis.

My experience with this setup is that the array will generally show up in windows if my Areca card is running 1.49 firmware. When I upgraded to 1.51, it never shows up unless I power down the Server, turn on the Supermicro chassis, and then power up the server again.

Does anyone have a similar setup and know a workaround? Ideally, I would like to leave the Server on and power up and down the Supermicro chassis as needed.

P.S. I hope OP get his files back. Good luck.

Undervolt the fans from 12v -> 5v and have the drives just powerdown when idle? Still a ways from power effeciency and noise to being off but if you need to access it often I think this is the best way. This is what I do with my SC933's.

My SC846 is not that bad (loudness wise) with the fans at low RPM and the right PSU (some models of the PSU are super loud).
 
odditory: This is "business critical" data, but not time sensitive. Originally I was in a huge time crunch, but now I am not. I have a few weeks to recover this data. The files are recordings from measurement instruments, they are large (30GB++ each). These instruments write to these arrays but the arrays must be moved post-measurement to different machines for data processing work. That is why the logs from the Areca cards are so random and incomplete.

The responses I got were indeed from Kevin Wang. I offered to call him (I lived in Hong Kong for many years and spent decades working with Taiwanese hardware engineers, and I am very fluent with "Chinglish" and know some Cantonese and a little Taiwanese). He did not seem very interested. I just got mostly useless answers as you noted. He just wanted to rant on about my misbehaviour and about how the situation was hopeless without proper logs.

At one point he directed me to use the RESCUE option. That didn't do anything. I did try the "no init" step -- which I noted above when I did it. That didn't do anything useful: The array now says normal, but there is not trace of files or NTFS. I seem to remember feeling pretty confident that this step would just work, so I clicked through whatever options were defaults. I suspect that there was an option somewhere that I did/did not set/unset etc. because it didn't work. So that one may now be lost, but not to worry: There are still two other failed RAID5 sets besides this one.

These other failed RAID5's two are in their as-originally-failed state. I haven't done anything to them.

The original two RAID5's (of which only one is horked) were:

-- identical 8-disk RAID5's comprised of identical 300GB Seagate Savvio 2.5", 2400GB gross, 1900GB net.
-- The one functioning RAID was created identical to the failed one: I am certain of this.
-- The file systems are NTFS created by WinXP 32bit 4K workaround for >2TB
-- I am sure that these have been used with at least 4 different Areca RAID HBA's (some 1680's, some 1880's...) so the logs should be considered worthless (in fact they are the same logs as have already been posted)

These latter two are:
-- identical 12-disk RAID5's comprised of identical 600GB WD 2.5" RE's, 7200GB gross, 6600GB net (each).
-- The file systems are also NTFS but created by WinVista 32bit 4K workaround for >2TB
-- Same situation regarding the promiscuity and worthlessness of the logs...

Other important details:
-- I now have R-Studio installed and running (on Gentoo Linux, licensed copy).
-- I now have an LSI 9207-4i4e non-RAID HBA so that R-Studio can see the individual disks without having to put the Areca in JBOD mode (someone said that doing the latter can be counter-productive)
-- I still have one Areca 1882 HBA installed in case that one is preferable to use

I can see the member disks in R-Studio via the LSI HBA, but only the RAID set via the Areca HBA.

I have assigned them to a Virtual Block RAID in the order I believe they were created in (but I didn't create them so I really can't be sure). I have tried to scan this array, but it's taking forever.

It seems that I really don't know what to do next.

In order to learn R-Studio, I attached the one working RAID5 in an attempt to assemble a working virtual array and try to figure out all the offsets, block sizes, and whatnot and to see how to order/reorder the disks since I *know* these should eventually line up and work. I am still screwing with that at the moment. If you can help jump-start me with some of the parameters, that would be great. here's what I am seeing:

http://www.enlightenment.org/ss/e-514dc806d81353.24001120.jpg
http://www.enlightenment.org/ss/e-514dc8278ac572.91942428.jpg

The scan is not complete yet.

I apologize for being so detailed again, but we've wandered quite far from the OP so I wanted to make sure we're all on the same page.

You guys are great. Continued thanks to all. I can be pm'd at capt_dot_frito_at_g_m_a_i_l if anyone is interested
 
Last edited:
w/r/t Hammer!'s question, I thought I would share Areca's strange answer. BTW, I do not believe that it is an accurate answer in any real-world situation. I know it is inaccurate my situation:

[Areca's Kevin W:] i think your problem is not 'how to properly connect a raid after boot' but 'how to properly disconnect a raid after boot'

to answer your question 'how to properly connect a raid after boot', there have no special procedure needed to properly connect a raid after boot. array configurations are stored in array member drives not controller. you can connect array member drives to a running system without additional procedure. firmware can recognize these configurations in drives and combine them together into a raidset.

First, you can see that Kevin (Areca) thinks I do not know how to ask a question. That makes him not fully think through what I actually asked him and give an off-the-cuff answer.. *sigh*

Kevin says there is no special procedure; this does not match my experience. If I power up the array enclosure while it is already connected to running HBA, the staggered starting of the disks confuses the card, and it will put the array in an unstable state as far as the card is concerned. My experience is this: The disk enclosure has to be powered on and all the disks must be fully available -- only after this can I plug it in to the HBA, but not before.

In my case I have many RAID volume sets housed in several enclosures. Some enclosures have multiple RAID sets. If I 'offline' individual RAID sets, it powers down the associated disks. There is no way (that I can find) to power them back on without powering down the disk enclosure and powering the whole thing back on. Of course that means I have to reboot all the other systems that might be connected to the enclosure, and 'offline' the arrays before hand too. But there are other problems with the HBA where this does not always work either.

Also, in my case, if I plug in the array enclosure normally and the RAID volume mounts okay, and then I 'offline' it, I cannot plug in a different array into the same HBA -- it will not re-read the volume information correctly. For some reason, it presents stale volume information (like disk identifier) to the Linux kernel and it can't mount the new volume. I cannot seem to get Linux to get updated volume info, even if I rescan the SCSI bus (it is this predicament that started this whole nightmare in the first place).

This is why I wanted to know if there was a way to command the Areca card to 'online' an array -- that is either send the command to turn on all the disks associated within an array, or if there is a command to re-read the volume information and disk identifier etc. without powering everything down and doing a warm or cold boot.

Areca also thinks that I can live in suspended animation between 3-day email turnaround time lines. That's not helping much either.
 
@houkouonchi

This is something that has peeved me with RAID subsystems; technically, even if drives get moved around it shouldn't matter. Drives these days have UUIDs and Worldwide Names. By rights, there should be enough information to stamp a "write order" into the metadata so that even if a drive gets moved to another physical port, writing and recovery should be unaffected.

I fell victim to this with md RAID; I had moved some drives around and for some reason mdadm marked every drive as being a spare. During the restoration attempt I screwed up and ended up blowing away the array. Highly annoying.

If you are using a SAS card then you can move them around. It's only when you use SATA RAID cards that drive position is important (for both software and hardware RAID)
 
I am really glad I saw this as I have a similar setup. Is there a workaround to having to "hot plug" an array?

I have a Server with a 1880x card that is connected to a Supermicro SC847-JBOD chassis. Because the Supermicro chassis is so loud, I only turn it on when I need to access the array. I think this is similar to "hot plugging" since the Server is generally already running when I turn on the Supermicro chassis.

My experience with this setup is that the array will generally show up in windows if my Areca card is running 1.49 firmware. When I upgraded to 1.51, it never shows up unless I power down the Server, turn on the Supermicro chassis, and then power up the server again.

Does anyone have a similar setup and know a workaround? Ideally, I would like to leave the Server on and power up and down the Supermicro chassis as needed.

P.S. I hope OP get his files back. Good luck.

Only way to accomplish this is throw full PC hardware into your Supermicro chassis and use 10-40GBe network cards to connect the two. Then you can turn them off and on to your hearts desire and not risk loosing anything.

The cheap easy way would be to simply power off both boxes before unplugging and re-plugging in.
 
@Capt.Frito

I'd love to help with specific info but its been a while since I have used r-studio. I'll try to find some spare disks and play around so I can get a feel for it again.
 
Practical question: When an NTFS filesystem is created, does it always reserve space for an "MBR"? that is the implication in the steps in R-Studio's docs but I am not sure if this assumes that the disk is a boot disk (array) -- none of my arrays are bootable arrays, but I didn't create them array -- they came to me already created...

I found the NTFS header on two of the 8 disks, both matched at sector 1024. Not sure how that helps. I also found the MBR (or at least the pattern match for both). However the template for the NTFS record and the MBR do not "fill in" like they do in R-TT's docs and examples. I am not sure if it's because I installed the Linux package on Gentoo by hand (from the Debian .deb pkg) but it's not working the way I thought it should :confused:

Does anyone know?
 
Capt.Frito, thanks for your help. I too tried to contact Kevin for help and didn't get anything useful. What I ended up doing is similar to what staticlag suggested which is to chain my JBOD chassis to my server using an Add2psu adapter which turns on/off the JBOD chassis with my server. The only problem with this is that I cannot do a reboot as the Areca card in the server will hang in the bios screen during a reboot, if I did (maybe because the Areca card does not "reset" the drives?). I'm not sure why - so if anyone reading this has any suggestions to correct this, it will be much appreciated.

I actually thought about putting a motherboard in the JBOD chassis and then connect that to my server using infiniband, but only mini-ITX boards will fit in the chassis and those only have 1 PCIe slots which is not enough for both the Areca card and an infiniband card. So, I can connect using the built-in 1Gbe ethernet, but speed is more important to me than being able to reboot - although it is a pain to have to go physically shutdown the server if I forget and do a reboot by mistake.
 
After much consternation, I was able to reconstruct the first failed array. I was able to figure out which disks should be in what order using rstudio and tracing through the MBR, the NTFS file system fragments and so on. It was painstaking work just like everyone said it would be.

However, I have learned quite a bit that I will share, in case anyone comes along who might benefit.

The biggest thing I learned was that Areca's support is useless, and in some cases quite harmful. Most often I just waited forever to get some utterly useless words, like, "use the block size you created the array with." That would be helpful except that the question was "I did not create the array, and this is not the card it was created with -- how can I figure out the block size in this case?" When I pointed this silliness out Areca (Kevin W) just accused me of asking the wrong question. This nonsense repeated itself time and time again :rolleyes:

Areca gave me information that was dangerously wrong on several occasions. For example, I asked them how to properly "online" an array after issuing an "offline" command. They said just plug it back in. I am here to tell everyone that if you follow this advice from Areca you will BLOW UP YOUR RAID for sure. As happened in my case, the "offline" command powers down the member disks. Pulling the SAS cable and plugging it back in did indeed cause them to power up, just like Areca said, but the disks were set to stagger their starts, so when the Areca thought it was time to see if the array was there, only 6 of the 8 disks were available. Kaboom. Nice work Areca.

There are many more examples, but suffice it to say, that you will get bad advice from them if your case does not fit their expected profile. That is:
1. You must be the person who created the array, and you must have meticulous notes on what you did -- or -- be able to find depose/recover the notes from whoever did create it;
2. You must have only used one controller card ever;
3. The controller that you have now MUST be the controller you created the RAID with originally;
4. You cannot have cleared your logs ever or have had the initial log entries covering the creation sequence rotated out
5. You better not have had the "auto-fix" or "consistency check options turned on when things went bad;
6. You must be prepared that if they do not understand the situation it is because either a) you do not know how to ask a question or b) that you are stupid.

If you do not fit the above profile to a "T", don't waste your time with Areca. Send them a note saying what happened so they know, but don't waste time listening to them, at least not Kevin Wang. If you do decide to try, assume their advice is unreliable and think it through, get second opinions from this site, etc. Be skeptical. Very skeptical.

In the end, this array was an 8-drive RAID5, where 2 disks were shown as failed -- because of hotplugging the RAID without offlining it first and then hotplugging it without powering up the disks and letting them settle first. (Please note that the instructions I got from Areca's after-the-fact have caused this to happen, regardless -- so asking them how to do it in advance would have made no difference in what happened to me!)

I spent lots of time screwing with rstudio. It was fascinating and educational. But in the end it did little to help me understand how to actually order the disks as far as the Areca card was concerned. This was because rstudio ordered the disks based on SCSI parameters that were useful within its own environment, but useless in Areca's world. If there was a way to relate the two, I cannot say for sure, but what I can say is I never found it.

I had thankfully earlier followed houkouonchi's advice to go the "no init" route. After days and days with rstudio and learning that the data was still there and getting frustrated trying to figure out how to translate the rstudio parameters in to ones that made sense in the Areca card environment, I decided to give up.

THEN THE MIRACLE OCCURRED

I hooked the array up to the only system I had that could reliably create the closed/proprietary WinXP NTFS system I needed to put my RAID5 system back into service, (this is after the no-init action which restored the volume set but not the NTFS file system, despite all the flailing with rstudio). Windows tried to boot, saw the array, and ran an old-school chkdsk when I wasn't looking (mercifully -- I probably would have stopped it instinctively). When it emerged Windows booted normally. For grins, I looked at the array's contents, and lo-and-behold, there were my files.

With renewed faith in the people on this list and renewed contempt for paid-for tech support, I will now attempt to recover the other two RAID5's.

You guys were/are terrific. Areca? Caveat Emptor...
 
Honestly from past experience 3ware/LSI are not much better.

3ware locks up their cards so much its almost impossible to recover yourself. They basically told me they couldn't do anything for me a couple years ago when an array (24 drive raid6) with 1.5 TB disks failed because their tools could not support 1.5 TB disks.

I spent a week trying to use their tools they sent me to recover the array when they neglected to tell me I needed a special firmware to even use their tool (which was DOS based and like 20MB so a huge PITA).

Luckily the script their tool used was in plain text and I was able to find the problem (It was not setup right for disks on SAS expanders because aparrantly they don't ever deal with people using SAS expanders) and I was able to modify it tor recover the other array that died at the same time on the same controller (1TB disks). From reading their bat file I was able to figure out where the raid metadata was stored on the actual disks (1020 sectors from the end of the drive) and its size. I was then able to use their utility on a 1 TB disk system to dump the metadata and then do the same with a JBOD controller and linux/dd to verify i got the same data (via md5sums).

I then had to create an entire new array (with 24 new 1.5 TB drives). Wait for the init to complete. Then hook it up to a JBOD controller to dump the metadata and then hook u p the original disks to write back to the array.

A huge pain in the ass which could have been avoided had 3ware had a 'no init' option for rescue like Areca and LSI can do.

Its a decently complicated process. If you want you can check out an internal wiki article I made back when I did it in 2009:

http://box.houkouonchi.jp/3ware_dcb_corruption.html

I was able to recover 100% of the data off that 3ware array where 3ware literally told me there is nothing they could do and the data was unrecoverable... LAME.

Edit:

One of the reasons I had you run the vsf info vol=1 and what not is so you can find out what the stripe size is and stuff. Also you can tell if logical/physical order is correct from the web-interface (or on newer cli64 with cli64 rsf info raid=X after I got areca to add that ability from the cli) because in the web-interface it will list the ports in the top part in a bad order. For example through the new cli:

Code:
# ./cli64 rsf info raid=1
Raid Set Information
===========================================
Raid Set Name        : Raid Set # 000
Member Disks         : 8
Total Raw Capacity   : 8000.0GB
Free Raw Capacity    : 0.0GB
Min Member Disk Size : 1000.0GB
Supported Volumes    : 128
Raid Set State       : Normal
Member Disk Channels : E1S1.E1S5.E1S3.E1S4.E1S7.E1S6.E1S2.E1S8.
===========================================
GuiErrMsg<0x00>: Success.

So this machines logical order is 1, 5, 3, 4, 7, 6, 2, 8

which means:
slot 7 would need to physically be moved to slot 2,
slot 2 physically moved to slot 5.
slot 5 physically moved to slot 7.

if you wanted to recover it via the 'no init' option.

In the web-interface the order listed in the top part will be out of order like that (counting top to bottom)

The logical/physical order getting messed up is one of the reasons I higly recommend against hotspares.
 
Last edited:
@houkouonchi: I have an old 3ware card & operating RAID5 from ages ago. The array has been running fine for all that time. I guess I should retire the whole thing. If not, I will try to save a copy of the DCB before it's too late :D So, once again, you have helped immensely,

I have always had great success with software RAIDs, 5's and 6's. For performance I suppose that hardware cards are still the way to go. It also seems true that for all the hardware-based solutions Areca is probably the easiest to recover from a given disaster. It just irked me a little that the "help" I got from Areca varied from utterly useless to downright wrong. I didn't help that it was all penned in a very condescending tone. :mad:

If it weren't for this list my situation would most certainly have been hopeless. Forever grateful :)
 
Thanks for sharing your experience and congrats for being able to recover your data.

So how does one bring a raidset back online after an offline command? Is a complete powerdown and then poweup the only way?
 
Thanks for sharing your experience and congrats for being able to recover your data.

So how does one bring a raidset back online after an offline command? Is a complete powerdown and then poweup the only way?


I think the problem is when you have an expander and areca firmware bug. If you have local disks I wouldn't think its a problem. You definitely want to offline it if your gonna disconnect it though.
 
@Hammer! Assuming the offlined array is external via an expander, I believe that the only way to bring an offlined array back online is to power down the chassis and power it back up again AND you must wait until all the member disks are fully spun-up and stable before connecting the SAS cable. At least this is the only way it works for me. And please note that I only have SAS experience.

And again in my case, more often than not, the OS (usually Linux 3.6.11 kernel w/arcmsr driver but also Win Vista 32bit, WinXP 32bit) does not detect that the volume has changed. I have not found any forced-update technique that will make the OS get valid file system data. As houkouonchi mentioned, this is likely a driver bug but it seems that SCSI in general has a history of being stubborn like this.

This file system detection bug is not an issue if you offline, disconnect, power-up, and reconnect everything the same way. However, if I move RAID sets around, then I have to reboot to get the Areca card to understand what is going on. This is not a problem for the first RAID volume you present to the Areca card after a reboot regardless of any time delay in connecting it -- the first volume to get attached always works fine (assuming the disks were fully awake). This bug seems to affect only subsequent RAID volumes.

For me it is a bit of a hassle to do this because I have multiple RAID sets in any given SAS drive enclosure. This means I have to offline then all, and then initiate the whole power down/power up cycle if I am to be successful. That's life, i guess.

@houkouonchi -- thank you for the edit. I am sure that will help me with the other two RAID volumes I jacked up.

When I start that process, I will likely start another forum topic to follow the process a little more cleanly, especially the parameter discovery/member disk order parts using cli64.
 
Last edited:
And again in my case, more often than not, the OS (usually Linux 3.6.11 kernel w/arcmsr driver but also Win Vista 32bit, WinXP 32bit) does not detect that the volume has changed. I have not found any forced-update technique that will make the OS get valid file system data. As houkouonchi mentioned, this is likely a driver bug but it seems that SCSI in general has a history of being stubborn like this.

There are a few ways you can re-detect a volume if it changed sizes or went off/on etc..

One is use this script:

http://box.houkouonchi.jp/rescan-scsi-bus.sh

In some cases you have to run:

echo "0 0 0" > /sys/class/scsi_host/host0/scan

although host0 might need to be changed to whatever the areca controller is. I have seen it not work automatically before but ive never had neither of these two methods work although in one rare case I think I had to change the LUN with like a failed array to be 1 digit higher so it detected as a new block device. Just rescan-scsi-bus.sh worked with changing the LUN.
 
Anyone else amazed at the technical know how in this thread?

Seriously, two documented occasions on how support failed from a 'reputable' hardware vendor, and yet the customer persevered and recovered from 'unrecoverable' circumstances.

As a person who's getting read to build a NAS for my company, I've learned a lot in the past 30 minutes of reading this thread.
 
Anyone else amazed at the technical know how in this thread?

Seriously, two documented occasions on how support failed from a 'reputable' hardware vendor, and yet the customer persevered and recovered from 'unrecoverable' circumstances.

As a person who's getting read to build a NAS for my company, I've learned a lot in the past 30 minutes of reading this thread.

It's pretty common around here. I would say that [H] has the most informative forum out there, It's the only one that includes information on a wide variety of operating systems, cases, storage options , processors, video cards and audio. It's a forum of builders, developers, OEM's, and regular end-users.

Over the years I've begun to get to the point where there's really no reason to post anywhere else.
 
Not to go off topic, but I've found the [H] my favorite tech forum over the years, but some of the treatment of the less knowledgable members has been a turn off and bordering abuse at times.

Luckily, Ive had mostly psotive experiences here, but as good as [H] is, it has that much more room for improvement.
 
Anyone else amazed at the technical know how in this thread?

Seriously, two documented occasions on how support failed from a 'reputable' hardware vendor, and yet the customer persevered and recovered from 'unrecoverable' circumstances.

As a person who's getting read to build a NAS for my company, I've learned a lot in the past 30 minutes of reading this thread.

I get the feeling that a lot of people use components for purposes they are not designed for and then blame tech support for not supporting those purposes.

The solution seems to have been found by dumb luck - not stopping ChkDsk, not technical skill. No one made that suggestion.

---

We should note that one of the problems was that the controller was set for staggered spin up. Despite that being a problem no one here suggested changing to non-staggered spin up.

I could go on, but ...
 
We should note that one of the problems was that the controller was set for staggered spin up. Despite that being a problem no one here suggested changing to non-staggered spin up.

I could go on, but ...

Areca controllers do not support non staggered spinup. Trust me I wish they did. On my external 90 TB raidset that is two enclosures (each 37 ports on each) means that for 30x3 TB disks I have to wait for a staggered spinup attempt on 74 disks which means even when its set low it takes a while. I know my areca controllers (for both controllers) on my machine (54 disks) takes about a minute to start up during POST.
 
I get the feeling that a lot of people use components for purposes they are not designed for and then blame tech support for not supporting those purposes.

The solution seems to have been found by dumb luck - not stopping ChkDsk, not technical skill. No one made that suggestion.

---

We should note that one of the problems was that the controller was set for staggered spin up. Despite that being a problem no one here suggested changing to non-staggered spin up.

I could go on, but ...

FWIW, in 7 years of dealing with Areca I've never found them to be rude or condescending via email. I guess its all in the approach. There have been times I was frustrated that response took a while or they were brief but just took some patience with some additional back and forth.

In my experience dealing with just about every brand of controller, believe it or not Areca's have actually been the most flexible when its come to disaster recovery scenarios. Even when meta's been lost and drive order is lost, worst case it *can* be done with R-Studio, but its not for the faint of heart as it requires some data recovery experience (the ability to analyze disks with a sector editor to determine disk order, lay in a partition offset as well as determine block pattern - which in Areca's case is proprietary or at least nonconventional) - unfortunately its somewhat beyond the scope of a forum thread but there are people that will do these types of recoveries on a consulting basis (DR-Kiev at hddguru.com is highly recommended, I'll do them on occasion if I have time).

In any case OP good luck with the rest of your arrays.
 
Last edited:
@GeorgeHR -- I believe that you missed the point entirely. The Areca controllers are being used for exactly what they were designed to do. I am using these to create RAID5 volume sets that are movable across a family of RAID controllers and which can be hot-plugged. All of this is well within Areca's advertised capabilities. Edit: chkdsk would not have worked if I did not successfully complete painful procedures with rstudio and rebuild the array via Areca's "no-init". To imply that a simple chkdsk was all that was ever needed is wrong and misleading -- it happened when the volume was reporting as good but the NTFS volume could not be read. I plugged the array into a Windows box to reformat and that's when the chkdsk occurred. Fortunately I booted the Windows box with the array attached -- if I hadn't the chkdsk would not have run at all. It was the last thing that I needed to do, and it was not mentioned in this thread, that is true. But to say that's all I needed to do is nonsense. It is just that when I completed the "no-init" process the card was set to do a consistency check, and so it kicked it off without any help from me. I should say that I got the Areca card as part of a purchase from a third-party who had done some changes to Areca's base defaults. I say this, not to place blame, but rather to caution anyone from drawing the conclusion that whatever Areca didn't do I must have done. Not so. There were four companies involved in this thing along the way. I just so happened to be the one to whom the restoration job fell. I have no desire to place blame on anyone but myself for the cause, and have repeatedly stated this throughout this thread. Your defence for "downtrodden tech support everywhere" and "bad end-users everywhere" is a commonplace straw man argument.

In the end I used rstudio, painstakingly went through the disks, and so on, so I don't think it is beyond my ability or that I am faint of heart or that I lack patience. I did get the first array data back, by complete happenstance. Windows chkdsk of all things found and repaired the file system after I had done everything else like reorder and no-init the array, etc. No one mentioned to do something like this, ever, it was a total fluke that it even got run. But there are points along the way where the process can be greatly aided or seriously hindered by getting information known only to the RAID card vendor. If they don't share it, or if others that might know keep it hidden for whatever reason, it can make a miserable situation become even more so. I for one would have been just as happy to pay the $? for the answers -- this was not a question of cost, but one of expedience.

I should say that I have software RAIDs, plus hardware raids by three vendors: Areca, Adaptec and 3Ware. of the hardware variety, I prefer Areca and won't be changing, so if anyone thinks my intention is to Areca-bash, you are wrong. In fact as a direct result of this disaster have decided to retire non-Areca RAIDs and convert them to Areca.

The issue I have with Kevin @ tech support is that I was repeatedly told wrong things, or even worse, my questions were answered with questions. By the end of it, Areca told me not to trust any answers gotten via the Internet -- i.e., on forums like these. odditory: They are impugning your advice right along with the rest of it, of which your advice (among others) I have learned to hold in the highest regard, ironically marking yet another piece of bad data from Kevin. So it is completely fair play that when I get bad advice from Areca that the Internet be advised of it. And I have posted this explicitly, not to embarrass Areca, but to be free of accusations that I was taking things out of context, etc.

I asked for, and needed very specific details that only Areca would know owing to the behaviour of their controllers. I was not asking them to fix my problem, per se, but mainly to fill in gaps in my understanding and provide details about their raid creation. For example, in order to use rstudio effectively, I needed to know Areca's block ordering scheme: left hand sync, right hand, etc. In answer, Kevin W simply asked me if I was asking the right question. I never received an answer to that one. I suppose in the end I would have been better off just asking single one-liner questions with multiple choice answers, but that would have required me to be the RAID expert and not Areca. And if Areca would have told me up front that I could only ask one question per day and there would be a 72-hour turnaround time I would have been a lot less frustrated by the process. All those things I know now, after now having gone through it. And some have been through it and their approach can be tempered by their experience gained over their years of personally dealing with Areca. As for me, it was all on-the-spot training. For the record, I have spent the last 30+ years providing tech support to Asian customers on highly technical telecommunications products so I have great empathy for both sides of this support and language barrier issue.

For example if someone asks me, "Is this signal is modulated as 64QAM or 256QAM, or what is the minimum SNR needed?" I don't wait three days and then answer with things like, "What did you set it up to be? Are you sure you want to know what minimum SNR should be? You probably are really wanting to know symbol rate. Check your logs for symbol rate. It will be whatever you set it to... I am sorry, Taiwanese is not my original language."

Thus I too am also sensitive to "approach" and know that customers in technical distress are edgy on account of circumstances.

But since this has now turned into a question of who said what and how id they say it, I will quote the entire bodies of the emails in the thread, and you all can judge for yourselves.

First, I had asked Areca the basic questions about what to do, gave a status, relevant facts, logs, etc. We got to the point about my mistakes not disconnecting the RAIDs without offlining them first. Then I had a problem related to the half-shut-down array enclosure and the spin-up problem:


On Mon, 2013-03-25 at 20:12 +0800, Areca Support wrote:
Dear Sir/Madam,
you had told me that you have three raidsets with problems, but recover one raidset at one time will be better to avoid mistake. could you please focus on recover one raidset first ?
or do you mean the initial screenshoots you provided (at 3/18) with degraded volume belongs to one raidset on one card and these photos you provided next day with a failed volume belongs to another raidset on other card ?
to answer your question 'how to properly connect a raid after boot', there have no special procedure needed to properly connect a raid after boot. array configurations are stored in array member drives not controller. you can connect array member drives to a running system without additional procedure. firmware can recognize these configurations in drives and combine them together into a raidset.
if the array signatures in these drives are mismatch, these drives will not be listed in same raidset.
the event log shows drives been removed while running, and drives been removed while running result array broken,broken array mean array signature change. so i think your problem is not 'how to properly connect a raid after boot' but 'how to properly disconnect a raid after boot'.
Best Regards,

Kevin Wang

It is important to note that we had already established that initial causation was my error in not offlining the array, and that were had already decided to work on only the failed 8-disk array (the other two were 12-disk arrays). This is why I had asked about onlining because of the spin-up issue -- this was in answer to the third time I asked about it. But Kevin does not address this, he just assumed that I asked the wrong question. So I responded with this -- note I may have been frustrated a little at the fact he didn't think I understood even the most basic difference between connecting something and disconnecting it. I had sent my email asking about this on March 22, 2013 9:01 PM. So it was three days before I got the reply I posted above. And this time will be the fourth time asking about the online/spin-up issue, and by now I admit to being a little weary at the task. Anyway I responded thusly:

To: Areca Support
Sent: Monday, March 25, 2013 9:59 PM
Subject: Re: Multiple RAID5 failures using 1882 cards -- need help

You wrote: "so i think your problem is not 'how to properly connect a raid after boot' but 'how to properly disconnect a raid after boot'."

You are wrong -- I asked the question I wanted answered! I understand how to use the "offline" setting already. So please do not tell me again.

the issue is this: When I "offline" an array it powers down the disks in the disk enclosure. But there are more than one raid set in the disk enclosure. I am asking, is there a way to turn the disks back on without setting all disks "offline" and powering everything down and powering it all back up? The 1882 card does not properly recognize a NEW array unless I power cold-boot the host with the 1882 HBA. Is this the only way?

Also, I have three RAID volumes that are in this state. I use the same 1882 HBA for all of them, so the logs will always be very confusing. I followed your advice with one of the RAID sets (the first one we began discussing ("no init"). But it did not work. I used all the defaults which included some setting "assume data is good, recompute parity". When it finished it said that the RAID volume set was "normal" but there were no files on it when I mounted it (read-only). So whatever you told me to do failed. I do not know if there is any hope with that volume set anymore. But I have not touched the other two failed sets at all.

I have switched to using R-Studio to try recovering this, but I need some critical information (some I know, but perhaps you could fill in what I am missing. If you don't know directly it would help if you could tell me how to find it:
1. Number of disks: 1st bad array = 8 disks, all 300GB; 2nd & 3rd bad array: 12 disks, all 600GB
2. RAID type: All RAID5 (no spares, never had any)
3. File System: All NTFS (created by Windows XP/2003) with 4K option to break 2TB limit in XP
4. Type: Basic volume

Its unknown parameters that must be found are:
1. Disk order
2. Block size
3. Block order
4. Disk offset

I have multiple RAID volume sets I need to access, not just one. I cannot just leave one array plugged in for days while I wait for you to answer questions, especially when you provide answers to question I did not ask. So when finally do answer me, I have to shut down servers and arrays, move cables and re-start everything. It is not easy, and it fills up the logs with even more confusing information.

The questions I asked at the end came straight from R-Studio. I have seen these parameters with software RAIDs i have created, so I don't think these are out-of-scope. The good folks at R-TT had told me to ask Areca for these answers explicitly. Even odditory says that these are 'proprietary or unconventional'. So if Areca doesn't spill the beans, it becomes very difficult. Anyway, here are the non-responses I got back after the obligatory 3-day wait:

On Thu, 2013-03-28 at 15:43 +0800, Areca Support wrote:
&#65279;
Dear Sir/Madam,
two methods to added a offlined raidset :
1. restart the host with raid card as you said.
or
2. remove these drives belongs to offlined array and reconnect them back. you do not have to power off the entire enclosure with some drives belongs to another raidset.

and regarding the question about these needed information for recovery :
1. disk order, the disk order stored in array configuration field which had been over written by the new array you created.
2. block size, what it exactly mean ? block for each member drive ? if so, it will be the stripe size you configured
3. block order, what is the difference between disk order and block order ?
4. disk offest, controller stored configuration with various space depends on conditions. so the offset for data storage not fixed.

and regarding the not work raidset you created.
if new created raidset do not work, it could be caused by following reason
1. the drive order in new raidset is incorrect
2. the volume configuration is incorrect
3. one or more drives do not have correct data inside.

if you have older screenshoots for this array from controller web manager, maybe you can refer these photos to find out the possible cause.
i do not know why you mentioned the setting 'assume data is good, recompute parity', this setting should available in volume check only but you should not need to execute a volume check while recovering a broken volume.
if you execute volume check on the new created volume, it may corrupt the data inside if the new volume use incorrect settings.
for example, if you use wrong raid level to create the new volume and execute a volume check next, you will lost huge data inside.
and array recovery depends on conditions, we need to understand what happen before to find out the most safety recovery procedure. because improperly reaction may corrupt the array configuration easily and configuration corruption may result data lost.
so we have no any document can allow customer recover every broken array without our analyze.in our experience, many customers trying to recover array by these posts they found on internet, the result is data lost.
Best Regards,

Kevin Wang
Areca Technology Tech-support Division

The highlights are mine. The first highlight is just wrong information. When I did what he said, I ran into the spin-up issue, just as I worried would happen -- and previously asked about. Fortunately for me, it only dropped three more of the member disks out of the array, which by now was getting to be routine, so it didn't make anything any worse, but if this had been a good array, it would have broken it just as if I had not offlined it.

The next highlight shows yet-another reference to the incomplete logs. Without those, he kept contending, the situation was lost.

And then in some bitter irony, the last highlight shows him telling me not to trust forums like this one, but rather rely solely on Areca (in effect).

In all that he wrote, he did not supply a single answer or say anything that he had not already written. Please note that the turn-around was another three days, even though my replies to him were within minutes in almost every case. I suspect Areca has a 72hr turn-around and he was happy just to answer questions with questions to comply with job policy, but I admit I am speculating now.

Here was my not-so-nice reply to Areca. I should not been more civil, but by now this had been almost two weeks, with 72 hours spacing each email from Areca, and pressure on me mounting. And I should say again that I was not blaing Areca, but by now R-TT were telling me taht there were details I needed that only Areca would know and that without these the situation was not going to get resolved. So, I replied:

To: Areca Support
Sent: Thursday, March 28, 2013 11:27 PM
Subject: Re: Multiple RAID5 failures using 1882 cards -- need help

You advice has no practical value for me, and in many cases is just wrong.

For example, your suggestion that I just can plug in the offlined RAID again and it will recognize, is not what happens. The offline action powers down the disks. Plugging the RAID back in powers up the disks, but they are a staggered start, so the 1882 waits some amount of time and then reads whatever disks are available but at a time when all the disks are still spinning up, so it sees the RAID as incomplete. This is exactly what caused my problem in the first place. So your advice is wrong and your understanding is wrong. What you have instructed me to do will CAUSE a RAID to fail.

I told you I did not create the RAID in the first place, so telling things are whatever I used when I created the RAID is just useless. I need to know how to find out now. In fact, it was not created using my 1882 card -- the RAID was created by another company entirely. That is why the logs do not have the history you are looking for -- my cards were not used to create the arrays.

As for your comment, "i do not know why you mentioned the setting 'assume data is good, recompute parity', this setting should available in volume check only..." I do not know why simply cannot understand that I am telling you what happened. All you seem to be interested in doing is trying to convince me that I am stupid. As it happens the card was set schedule consistency checks, so it did it automatically. I did not realize what was happening at the time.

You have sent me many emails -- but all these say is why you cannot help me, and why this is my fault, and why everything I am doing is wrong, and that I keep asking the wrong questions, and that I am confused, why Areca is perfect. This is just a huge waste of time.

And Areca's most recent reply:

Dear Sir/Madam,

the status 'incomplete' is used for avoid unnecessary raidset degradation just because some members not be detected yet.
if you have a raid failure case caused by insert members while system running, please inform me more detail about the case to find out the root cause.

after schedule volume check completed, does the check completed event shows any error ?
if the check result contain no error, the volume is healthy. or if the check result contain huge errors, the data in volume are currupted.


and i am sorry, english is not my native language, so i may not able to fully understand your descriptions and asking improperly questions.
maybe you would like to contact with our locak distributor for assist.
distributors can be found in our web site, the where to buy page :
http://www.areca.com.tw/wheretobuy/main.htm


Best Regards,
Kevin Wang
Areca Technology Tech-support Division

Maybe this is where GeorgeHR entered in the picture: "maybe you would like to contact with our locak distributor for assist. distributors can be found in our web site, the where to buy page :http://www.areca.com.tw/wheretobuy/main.htm"

Regarding the language barrier deflection: I lived full-time in Hong Kong for 10 years, and got along quite well. I travelled the region quite extensively, China, all of SE Asia, and more. I can almost always deal with language barriers -- and I can even speak and read/write (complex) some Cantonese, a little Taiwanese, Putonghua, etc. I can also eek out some Tagalog, Ilocono, Thai ... and some Spanish.

That's the story on Areca's support from this case's point of view. I am sure after 7 years of dealing with Areca's support, odditory has developed some rapport with them, and that's great. But as for my approach, I asked very straight-forward questions, but only got only questions back. The only hard answers I got back were either simply restating that the original cause of it all was not offlining and therefore not Areca's fault, or worse, was just plain wrong next-step to-do's.

I don't know how my approach could have provoked such a thing, but maybe it did. Judge for yourselves.
 
Last edited:
Back
Top