• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Areca 1880i RAID Set/Volume Set Missing - Help Needed

Since the controller has already found some errors, I'd stop the check and re-run it with scrubbing and parity re-calculation. See if it makes any difference. With all drives being in place there is a high probability that scrubbing and parity re-calculation could take care of the errors. Do not give up on it just yet. :)
What Background Task Priority and Disk Write Cache Mode are set to?
BTW, check the HDD Power Management settings. What do you have in there? And please post a screenshot of the SMART data page of any drive.
 
Since the controller has already found some errors, I'd stop the check and re-run it with scrubbing and parity re-calculation. See if it makes any difference. With all drives being in place there is a high probability that scrubbing and parity re-calculation could take care of the errors. Do not give up on it just yet. :)
What Background Task Priority and Disk Write Cache Mode are set to?
BTW, check the HDD Power Management settings. What do you have in there? And please post a screenshot of the SMART data page of any drive.

Done, restarted it with both options selected. It is at 16.7%, 7 errors found

Background Task Priority: High (80%)
Disk Write Cache Mode: Auto

Here is the smart status for one of the drives:

Device Information
Device Type SATA(5001438006223784)
Device Location Enclosure#2 Slot#5
Model Name Hitachi HDS722020ALA330
Serial Number <insert serial>
Firmware Rev. JKAOA3EA
Disk Capacity 2000.4GB
Current SATA Mode SATA300+NCQ(Depth32)
Supported SATA Mode SATA300+NCQ(Depth32)
Error Recovery Control (Read/Write) Disabled/Disabled
Disk APM Support Yes
Device State Normal
Timeout Count 0
Media Error Count 0
Device Temperature 56 ºC
SMART Read Error Rate 100(16)
SMART Spinup Time 119(24)
SMART Reallocation Count 100(5)
SMART Seek Error Rate 100(67)
SMART Spinup Retries 100(60)
SMART Calibration Retries N.A.(N.A.)
 
Last edited:
Not sure that it really matters now but I had sent an e-mail reply a bit earlier to 'Benjamin' at Areca support detailing what I did more or less to get this point (volume operating, running check to fix errors). This was his reply:

What you have done was very dangerous to your raid set. I meant by removing one disk a time and reinserting it back. Anyhow, the current hierarchy (after SIGNAT) looks good.

The error from the volume check could be resulted from previous drive errors. You can let it finish as long as you have a good backup. Otherwise, you can stop it and backup the data again on the other copy.

Is what he is saying here at all accurate? Either way I guess it worked out but it is a bit troubling that Areca support thought we shouldn't do that.
 
One thing that jumped out at me right away - the drive's temp! It has exceeded 55C (out of its operational specs)!
Get better cooling for your system!

Secondly, never keep any setting at Auto. Choose either Enabled or Disabled option. This way it would be you making a decision, not the internal logic. If you left it at Auto and later connected a BBU, then the controller would disable Write Cache Mode automatically without you knowing it. And you'd be wondering why your performance has dropped all of the sudden.
 
One thing that jumped out at me right away - the drive's temp! It has exceeded 55C (out of its operational specs)!
Get better cooling for your system!

Secondly, never keep any setting at Auto. Choose either Enabled or Disabled option. This way it would be you making a decision, not the internal logic. If you left it at Auto and later connected a BBU, then the controller would disable Write Cache Mode automatically without you knowing it. And you'd be wondering why your performance has dropped all of the sudden.

Hmm not a bad idea, I may have to look into some better fans for the Norco 4020 backplane area (argh what a pain it was to get those installed). Despite the high temps, I will say that I have had those fans in there since essentially a year ago without any drive failures. Even before that in a previous incarnation using this case and the panaflow 80mm fans I put in (to replace the stock fans) it was running with a different set of drives (12 or so) without any real issues (that was a WHS server with no RAID controller card). Anyhow, for now I put a largeish external fan in front of it and drive temps are down to 45-52 C.

I will look into setting write-cache mode to something other than auto; good point on that.
 
Is what he is saying here at all accurate? Either way I guess it worked out but it is a bit troubling that Areca support thought we shouldn't do that.

What I had recommended to you was once recommended to me by Kevin Wang.
However, if you checked the history, you would find that after you got an answer from Benjamin my next reply was to update the F/W and then use Benjamin's suggestion. :) It does not say a thing about pulling the drives out.

At any rate, do not worry. I used this technique (with drives being pulled half-way out) a number of times and it has never failed. Also, I do not suggest something that has not been either tested by me personally or unless comes as a recommendation from a person I trust. Been working with Areca controllers, as well as reselling them, for more than 6 years already, used or owned pretty much every generation of Areca cards. What Odditory does here at [H]ard Forum (and I have a huge respect for), I was doing primarily at 2CPU.
You have not seen me hesitating making suggestions, have you? :) Well, it should speak for itself. So, make your own decision as to whether to be concerned about what you did to recover your data. ;)
 
Hmm not a bad idea, I may have to look into some better fans for the Norco 4020 backplane area (argh what a pain it was to get those installed). Despite the high temps, I will say that I have had those fans in there since essentially a year ago without any drive failures.

That's somewhat irrelevant, just because you haven't had any failures until now doesn't mean you weren't pushing your luck albeit unknowingly. The Hitachi's will take a lot of punishment, yes, but you're decreasing their lifespan and increasing the odds for array failure or multiple drives dropping within a short timeframe. I have experienced drops when drives got too hot, which was a rare occasion that a fan controller board power connector came unplugged due to vibration over time.

Jus made an excellent catch noticing the 55C temps. Way too high. If mine even get above 45C I know something's wrong. And if your drives are getting that hot due to lack of case airflow I don't want to know what the sas expander chip temp is sitting at, although that does have a higher thermal ceiling. Remember the HP sas expander was designed for high-airflow HP server cases, so while we co-opted them for home server use we have to be mindful of why they don't have a fan on the heatsink -- not because they don't need it, but because they don't need it in an HP server case.

Anyway good to see you got your array back and props for being the rare instance where someone doesn't panic and takes things one step at a time. It's always scary the first time you get hit with this situation, unfortunately this phenomenon is sort of the Areca's "personality" as anyone that's used these cards for any length of time can attest, but the Areca being superior in so many other ways more than makes up for it. And the main reason I suggest people to email Areca first is so they can get an earfull of what I think is something they could fix in their code or provide a tuning option in the WebGUI as far as turning sensitivity down and retries up before the card times a drive out. Areca indicating the HDS7220ALA330 is a desktop class drive and therefore 'not recommended' is B.S., thats more of a CYA (cover-your-ass) statement -- they should be grateful desktop class drives work great with their controller or they would have sold a whole lot less cards than they have. /rant
 
Last edited:
What I had recommended to you was once recommended to me by Kevin Wang.
However, if you checked the history, you would find that after you got an answer from Benjamin my next reply was to update the F/W and then use Benjamin's suggestion. :) It does not say a thing about pulling the drives out.

At any rate, do not worry. I used this technique (with drives being pulled half-way out) a number of times and it has never failed. Also, I do not suggest something that has not been either tested by me personally or unless comes as a recommendation from a person I trust. Been working with Areca controllers, as well as reselling them, for more than 6 years already, used or owned pretty much every generation of Areca cards. What Odditory does here at [H]ard Forum (and I have a huge respect for), I was doing primarily at 2CPU.
You have not seen me hesitating making suggestions, have you? :) Well, it should speak for itself. So, make your own decision as to whether to be concerned about what you did to recover your data. ;)

Ahhhh, I see I simply misunderstood what you said. I just assumed that you were not being that specific and really meant "do what I suggested and then do what Benjamin suggested." I will make sure to note that when I compile some of these tactics for future. I just was curious why he might have been concerned with me doing that; now I know why, thank you for the clarification. Anyhow, it sounds like it wasn't a big deal and it is quite clear that you know what you are talking about at any rate :). It is still chugging along at 29.9% and only 7 errors found so I am feeling fairly good about it.

That's somewhat irrelevant, just because you haven't had any failures until now doesn't mean you weren't pushing your luck albeit unknowingly. The Hitachi's will take a lot of punishment, yes, but you're decreasing their lifespan and increasing the odds for array failure or multiple drives dropping within a short timeframe. I have experienced drops when drives got too hot, which was a rare occasion that a fan controller board power connector came unplugged due to vibration over time.

Jus made an excellent catch noticing the 55C temps. Way too high. If mine even get above 45C I know something's wrong. And if your drives are getting that hot due to lack of case airflow I don't want to know what the sas expander chip temp is sitting at, although that does have a higher thermal ceiling. Remember the HP sas expander was designed for high-airflow HP server cases, so while we co-opted them for home server use we have to be mindful of why they don't have a fan on the heatsink -- not because they don't need it, but because they don't need it in an HP server case.

Anyway good to see you got your array back and props for being the rare instance where someone doesn't panic and takes things one step at a time. It's always scary the first time you get hit with this situation, unfortunately this phenomenon is sort of the Areca's "personality" as anyone that's used these cards for any length of time can attest, but the Areca being superior in so many other ways more than makes up for it. And the main reason I suggest people to email Areca first is so they can get an earfull of what I think is something they could fix in their code or provide a tuning option in the WebGUI as far as turning sensitivity down and retries up before the card times a drive out. Areca indicating the HDS7220ALA330 is a desktop class drive and therefore 'not recommended' is B.S., thats more of a CYA (cover-your-ass) statement -- they should be grateful desktop class drives work great with their controller or they would have sold a whole lot less cards than they have. /rant

Anyway good to see you got your array back.

Completely true Odditory, no denying that at all; highly likely that I have just been lucky the last couple of years. Also, it is possible that the drives haven't always at that level so it may not have been all luck (it generally doesn't run for that many hours straight). Absolutely true about the SAS expander, an HP case will certainly have more effective (and much louder as a consequence because in a server room who really cares) cooling.

In any rate, any good suggestions on cooling the 4020 down? How do you keep yours at below 45C or so? I am disappointed that the panaflow 80mm fans might not being doing their job well; they are so quiet. It might be that those just need replacing but it might also be that I (like several others around here I recall) 'turned around' the fan panel so that there was actually some room for power cabling next to the backplane so the fans are maybe an inch further back from the backplane. I may try to find some 'nice' way of putting an external fan in front of the drive bays like I have now sitting outside of the Huntec rack but I can't yet imagine a pretty way to do that so we'll see.

As far as the SAS expander goes I was thinking about putting a slot cooler in but I haven't decided.

Thank you for the props; it was indeed a bit scary. I would expect a drive failure before something like this. After having read Nitrobass24's recovery thread I knew to take this quite seriously and I wanted to be cautious. Also, from several years of previous experience, working slowly and documenting has served me well in allowing others to assist me if it is something I don't have any experience in (or in undoing what I have done!).

Good call on suggesting contacting Areca, hopefully they indeed take these sort of occasions seriously and work on improving how the controller handles the situation. I had the same thought as far as the desktop drive statement went. I could already feel that that was one of those "Hey well if this troubleshooting goes south or we really can't figure out what to do, we can just say 'upgrade all your drives and then we can help you and move on because the end user will never do it."
 
Last edited:
Well the volume check completed with 'scrub bad block'/'recompute parity' options finished, it ran for 9 hours and found 32 errors. Is that anything to be concerned about?

I am running another check again and I will probably do a chkdsk after that.

Besides looking into a slot cooler or heatsink fan (mounted somehow?) for the SAS expander I am also taking a look at rack fans (http://cableorganizer.com/computer-cabinets/rack-fans.htm#fans). The only issue is that what really would be the best would probably be a fan that literally sits right in front of the 4020 and pushes air over/past the drives out the back. I may just end up with my lame external fan solution sitting on the floor in the front of the case. ;)

The rack case I have for reference (I have the door on):
http://www.huntec.com/product_detail.php?CatID=49&ProductID=861
 
What fans do you have in your norco?
I replaced my stock fans w/ Scythe DFS123812L-3000. They are quieter than stock but have good static pressure and move a lot of air.
 
What fans do you have in your norco?
I replaced my stock fans w/ Scythe DFS123812L-3000. They are quieter than stock but have good static pressure and move a lot of air.

Hi Nitrobass24,

Panaflo 80mm fans are what I installed to replace the stock fans.

How many drives do you have installed and/or what type of drives are they? What do your drive temps look like?

Aren't those 120mm fans? Did you get a special fan mount that allows that for the Norco 4020? Do they sell a replacement fan divider now that supports 120mm fans?
 
I had the Panaflo 80s but i bought a 120mm fan bracket from Cavediver. Norco now sells one too.

I have Hitachi 5K3000s 10 of them and they run about 38 degrees in my rack.
 
I had the Panaflo 80s but i bought a 120mm fan bracket from Cavediver. Norco now sells one too.

I have Hitachi 5K3000s 10 of them and they run about 38 degrees in my rack.

Excellent, I purchased a fan bracket from Norco; I am already not looking forward to undoing all my zip ties and installing it ;).

As far as the Scythe DFS123812L-3000's go, how would you rate the noise? Unfortunately, for better or worse the server sits about 5 feet away from me so it is preferable that it not be a huge noise generator (which is why I had the panaflo's).

As an aside the second volume check is at 42% with 0 errors; so far so good.
 
Hard to say w/o knowing what model# you have of the panaflow.
The 3000rpm ones that i have are not quiet by any standard just more quiet than stock. I would imagine the 2000rpm version is acceptable and still moves a lot of air.

I also don't notice it cause mine is in a closet, and the fan on my UPS is much louder than anything else in there.
 
Hard to say w/o knowing what model# you have of the panaflow.
The 3000rpm ones that i have are not quiet by any standard just more quiet than stock. I would imagine the 2000rpm version is acceptable and still moves a lot of air.

I also don't notice it cause mine is in a closet, and the fan on my UPS is much louder than anything else in there.

Ahh I see. Maybe I will try to find something that possibly moves 100 CFM or so.

Panaflo L1A's (80mm) are what I have, I believe these are the specs (the site I bought them from is now defunct):
Air Flow: 24 CFM
Fan Speed: 1900 RPM
Fan Size: 80x80x25 MM
Noise Level: 21 db
 
I am not thrilled with what the noise will potentially be like but I decided to order the 2000 RPM version:

http://www.newegg.com/Product/Product.aspx?Item=N82E16835185060 - Scythe SY1225SL12SH 120mm "Slipstream" Case Fan

I will just have to test it out and see how loud 3 of them sound together. Ultimately, based on reading quite a few reviews/posts, if you want something that moves a decent amount of air, you just can't have it too quiet.
 
I received my 120mm fan bracket today. I have not attempted to install it yet but I am looking at the case and the holes on the 120mm and it appears that it does not fit my case. Is this possible? I thought the 120mm fan bracket was meant for the 4020 but it looks like based on some research that there was a case revision to the 4020 and I have an older version? If so, how do I tell what version of the case I have? I have a Norco 4020 case that I purchased on 03/05/09.

In ways, if I can't use the 120mm fan bracket it is almost for the best; it was going to be a pain in the ass to recable everything (because you can't take out the fan bracket without re-cabling).

As it appears I am potentially stuck with the 80mm fans, it is clear that I need to replace the Panaflo L1A's I have. Based on a fair bit of reading (thanks Odditory) it sounds like the Masscool 80mm fans served him well: http://www.newegg.com/Product/Product.aspx?Item=N82E16835150007

It's either that or an arctic cooling 80mm: http://www.newegg.com/Product/Product.aspx?Item=N82E16835186031 but the Masscool's are cheaper and it looks like move more air.

I also installed a slot based PCI Vantec cooler nearby the SAS expander to hopefully cool that down a bit. Is there any way to view the HP SA expander temperatures? It is a couple of slots away but hopefully it should make a difference. I may also just move to SAS expander one slot closer to the PCI slot cooler, it's not a lot but it might make a difference.

I am considering attaching a tiny fan to the SAS expander using some small gauge wire or a zip tie if I can find one long enough (none of what I have can reach that long diagonally) to cool off the heatsink.
 
So, up until this point I thought I was out of the woods. I have ran multiple volume checks until they were clean, ran a couple of chkdsk's until those were clean (the first one fixed a few issues) and then literally copied EVERYTHING over from backups just to be safe (and to test the backups). This all went without a hitch and I was finally starting to use the RAID array again and had copied over some new data last afternoon (which is now technically not backed up).

Well, unfortunately I had some additional issues that are a bit of a concern. So last night when the volume wasn't doing anything (that I was actively aware of at least) it started beeping again and I checked the logs and it again started timing out to hard drives and then finally removed the 'enclosure' (ie HP SAS expander) and then degraded the volume/RAIDSet and finally 'removed' all 11 devices (I will paste the log below, I wanted to explain in 'English' first). So, I calmly rebooted expecting to have to do everything the same as before to get the controller to view the RAIDSET again. Luckily, after a reboot the array showed up, volume there, no issues, so not quite the same as before.

I started another volume check but then aborted it because I wanted to install a PCI slot cooler next to the HP SAS expander (it had not yet found any errors). I then left it on overnight doing a volume check and went to sleep. I was kindly awoken by the beeping again at 8:30AM or so. Looking at the logs, it appears that it had found 7 errors (no idea if it did indeed fix them or if it did not because of the time out issues). However, again it had a similar issue to what occurred last night. Time out to a particular slot/reading error, then time out to the SAS expander, then another drive time out and finally the SAS expander was 'removed' and the RAID set/volume was degraded/failed and the devices in all slots 'removed'. Again, another reboot and the raidset/volume was there and happy.

This time (in the morning) right after the errors/beeping occurred and before I started another volume check (which is going on no, 34.0% 0 errors), I looked at temps briefly. The drives all ranged from 36-48 C at the highest; not nearly as high as before (slot 11 drive was 40 C in case it matters). I also touched the HP SAS Expander heatsink and it was no-where near as hot as it was previously. It was warm but I would not say hot to the touch (you can leave your finger on it without feeling like you need to take it off).

So, I have e-mailed the above description and the event log to Areca but I was curious what you esteemed gentlemen thought. Are my drives having the timeout issues?

My HDD power management settings are:
Stagger Power On Control: 0.7
Time To Hdd Low Power Idle: disabled
Time To Hdd Low RPM Mode: disabled
Time To Spin Down Idle HDD: 60

Is my HP SAS expander getting too hot (seems unlikely given how much cooler it is now)? Is one of my drives failing (also seems slightly unlikely given that you reboot and everything is fine and that other drives do also time out)? Or is my HP SAS expander failing and I simply need to replace it with another one (maybe one that comes with FW newer than 2.02)? In case it matters I believe the P/N for my HP SAS Expander is 468406-001, and all the HP SAS Expanders I can find listed online are P/N: 468406-B21 so it is possible I have a slightly older one. It seems slightly odd for the SAS expander to be failing, if it was failing it seems like it would simply not work more consistently; not fail after 6 hours or fail after working for a few days straight while I restored from backups.

Anyhow, here is the part of the controller event log that should essentially match what I mentioned above (newest even first, scroll down to the bottom to see the oldest, let me know if you want me to take off the code tags):
Code:
2011-11-05 08:31:08 	Enc#2 Slot#15 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#14 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#13 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#12 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#11 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#10 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#9 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#7 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#6 		Device Removed 	  	 
2011-11-05 08:31:08 	Enc#2 Slot#5 		Device Removed 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Abort Checking 007:28:02 7
2011-11-05 08:31:08 	Enc#2 Slot#16 		Device Removed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001	Volume Failed 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-05 08:31:08 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-05 08:31:08 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-05 08:31:07 	Enclosure#2 		Removed 	  	 
2011-11-05 08:30:59 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-05 08:30:52 	Enc#2 Slot#14 		Time Out Error 	  	 
2011-11-05 08:30:42 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-05 08:30:00 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-05 08:29:49 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-05 08:29:14 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:28:47 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:28:35 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:26:18 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:25:20 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:25:00 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:24:49 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:24:26 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:24:15 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:24:06 	Enc#2 Slot#11 		Reading Error 	  	 
2011-11-05 08:23:59 	Enc#2 Slot#11 		Reading Error 	  	 
2011-11-05 08:23:47 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:23:19 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 08:23:07 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 01:48:28 	192.168.001.002 	HTTP Log In 	  	 
2011-11-05 01:03:06 	ARC-1880-VOL#001	Start Checking 	  	 
2011-11-05 00:57:06 	192.168.001.002 	HTTP Log In 	  	 
2011-11-05 00:55:38 	H/W Monitor 		Raid Powered On 	  	 
2011-11-05 00:54:24 	192.168.001.002 	HTTP Log In 	  	 
2011-11-05 00:27:08 	ARC-1880-VOL#001	Abort Checking 001:20:59 0
2011-11-05 00:13:55 	192.168.001.002 	HTTP Log In 	  	 
2011-11-04 23:06:09 	ARC-1880-VOL#001 	Start Checking 	  	 
2011-11-04 23:04:54 	192.168.001.002 	HTTP Log In 	  	 
2011-11-04 23:04:25 	H/W Monitor 		Raid Powered On 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#12 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#10 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#16 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#15 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#14 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#11 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#9 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#7 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#6 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#5 		Device Removed 	  	 
2011-11-04 23:00:43 	Enc#2 Slot#13 		Device Removed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Failed 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-04 23:00:43 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 23:00:43 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-04 23:00:42 	Enclosure#2 		Removed 	  	 
2011-11-04 23:00:29 	Enc#2 Slot#13 		Time Out Error 	  	 
2011-11-04 23:00:29 	Enc#2 Slot#12 		Time Out Error 	  	 
2011-11-04 23:00:29 	Enc#2 Slot#10 		Time Out Error 	  	 
2011-11-04 23:00:23 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 22:59:41 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 22:59:29 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 22:59:08 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-04 22:57:47 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 22:57:33 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 22:57:20 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-04 16:53:06 	192.168.001.002		HTTP Log In 	  	 
2011-11-04 16:49:02 	H/W Monitor 		Raid Powered On 	  	 
2011-11-04 12:13:10 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-04 12:13:10 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-04 12:12:46 	Enc#2 Slot#14 		Time Out Error 	  	 
2011-11-04 12:12:40 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-04 12:12:19 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-04 12:07:49 	Enc#2 Slot#11 		Reading Error
 
Last edited:
Try disabling Spin Down Idle option and see if it helps. With 11 drives configured for staggered spin up at 0.7 the controller may still think the drives (the RAID set that is) are not read for longer than 7 sec and starts dropping them off.
 
Try disabling Spin Down Idle option and see if it helps. With 11 drives configured for staggered spin up at 0.7 the controller may still think the drives (the RAID set that is) are not read for longer than 7 sec and starts dropping them off.

Done, thank you for the feedback Jus. I really hope that resolves the issue, that is better than replacing drives/SAS expander or thinking there is some issue with the backplane, cables or port on the SAS expander somehow.

It is 66% done error checking so far, no errors found, no timeouts/resets since 8:30AM or so.
 
So, as the volume check was nearing completion, 96% or so, it finally found 4 errors. Luckily/unluckily, I happened to be watching the output and noticed that slot 11 was again timing out/having a reading error and then the enclosure was timing out. So, instead of letting it go through the entire process of failing/degrading the RAID volume, I simply rebooted the server.

Code:
2011-11-05 17:40:21 	H/W Monitor 		Raid Powered On 	  	 
2011-11-05 17:39:52 	192.168.001.002 	HTTP Log In 	  	 
2011-11-05 17:39:30 	ARC-1880-VOL#001 	Abort Checking 	008:58:46 	4
2011-11-05 17:39:09 	Enc#2 Slot#11 		Reading Error 	  	 
2011-11-05 17:37:06 	Enc#2 Slot#11 		Reading Error 	  	 
2011-11-05 17:36:51 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-05 17:36:31 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 17:35:21 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 17:35:00 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 17:34:52 	Enc#2 Slot#11 		Reading Error 	  	 
2011-11-05 17:34:42 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 17:34:22 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 17:33:55 	Enc#2 Slot#11 		Time Out Error 	  	 
2011-11-05 11:33:39 	H/W Monitor 		Test Event 	  	 
2011-11-05 11:31:22 	H/W Monitor 		Test Event 	  	 
2011-11-05 11:29:44 	H/W Monitor 		Test Event 	  	 
2011-11-05 08:40:44 	ARC-1880-VOL#001 	Start Checking 	  	 
2011-11-05 08:34:35 	192.168.001.002 	HTTP Log In 	  	 
2011-11-05 08:33:46 	H/W Monitor 		Raid Powered On

Once the server booted back up and the array was fine I shut the PC down and took all of the drives out of their hot-swap drive bays and firmly re-inserted them ( also noted the serial of the drive it considers to be in 'slot 11' connected to the SAS expander). I also made sure all of the SATA cables felt firm connecting to the Norco 4020 backplane. Then I, once again started the 9 hour volume check process.

It almost feels like it is an issue with the drive itself but it does seem odd that it would time out the entire SAS expander if that was the case.
 
The drive #11 might be at fault here and/or its cable. Remember, they are SATA drives and communicate in simplex mode with the expander, whereas SAS do it in full-duplex mode. Expander has to wait for a SATA drive to respond back once a command (I/O request) was issued and may as well go belly up if the drive does not communicate as expected. That would be my guess. See if you can replace the drive.
 
The drive #11 might be at fault here and/or its cable. Remember, they are SATA drives and communicate in simplex mode with the expander, whereas SAS do it in full-duplex mode. Expander has to wait for a SATA drive to respond back once a command (I/O request) was issued and may as well go belly up if the drive does not communicate as expected. That would be my guess. See if you can replace the drive.

Interesting point. I am going to let it run the volume check and see if it has an issue again while it is doing that, In the meantime I did go out and buy another HD tonight so that I have one on hand. It's another 2TB Hitachi, HDS723020BLA642. Another 7200RPM 2TB Hitachi but it is 6Gbps (all they had that was 7200RPM and Hitachi) but I assume I will be able to add it into the current array with the other older 2TB Hitachi's and the Areca 1880i/HP SAS expander without any issues.

What would be the easiest/best way to get it into the array? Add it in as a hot spare to start with and then pull the drive in slot 11? Or try to expand the array to it? I am slightly afraid of what might happen if I tried to expand to it and the drive in slot 11 timed out.

--UPDATE--

Oddly, while it was running the volume check (9% in, 0 errors, it is running at like 3% per hour for some reason), it did have another time out on that same slot except this time it didn't have a read error, didn't have any additional time-outs and didn't cause a timeout on 'ENC#2'. It was just the two below and that was it; bizarre. Maybe fully inserting all the drives again helped in some small way?

2011-11-05 21:37:32 Enc#2 Slot#11 Time Out Error
2011-11-05 21:37:20 Enc#2 Slot#11 Time Out Error

--UPDATE--
It was only 9% in after 6 hours so I rebooted the server and started the volume check again.
 
Last edited:
I am having very similar problems with an 1680-ix8, HP SAS, Norco 4224 and 15xHitachi 7200RPM (older drives) [Raid6] + 9 x Hitachi 5300RPM (newer drives) [Raid6].

I also see time outs and then 'device removed' warnings for anywhere between 1 and 4 drives. Sometimes the drives are on the same Norco backplane, sometimes not.

I went through the process, like you, with Benjamin at areca and we recovered the array at least 3 times but still no resolution - i.e. he was not able to help me nail down exactly what the cause is.

I have ordered a new power supply (I'm currently using 750w, will try more juice) and also a new norco backplane (just in case). I didn't think to check the HP SAS temperature so I'll try to do that this weekend and then buy a cooling card if it feels hot. Benjamin also mentioned using some command line tool to change the timeout options or somesuch (email not to hand, haven't done this yet). Did you get a suggestion along these lines also?

Right now, I have a recognized raidset for raid_01, but no volume set associated with it. I am not sure if my data is toast. Was going to try the level2rescue and see but I probably should mail Benjamin again first.

Also, I got the same "thoughts" as you did, namely Hitachis are consumer drives, chassis vibration, etc etc. Sounded to me like he was out of ideas :( That said, Benjamin's support has been excellent - can't fault them for Areca's helpfulness. I'm just no closer to working out what the issue is :(
 
@mrwilby

Interesting. It does sound like I was in a similar though mine appears to have changed. My volume/RaidSet no longer disappears like that and the data is accessible. After I ran the level2rescue and signat commands my data has been more or less accessible since. It sounds like it is a good idea to run that command but I would wait to hear from Benjamin before you go ahead and do that to make sure your data is safe.

Let me know how replacing the PSU and the backplane goes, I would be curious as to how/if that works out for you. I am also curious how much of a pain it is to replace the backplane; at the very least the cabling will be a hassle.

As far as a command line tool, no Benjamin has not mentioned anything like that to me. I changed the 'time to spin down idle' value per Jus here on the forums but that was in the GUI. I have all the HDD power management options essentially disabled now.

It sounds like your issue is a bit tougher because it's not being consistent as to which drive starts the time outs and the device removed. Mine appears to START with device 11 pretty consistently which at least points to that drive/SAS->SATA cable/backplane/port on the HP SAS expander. Though, thinking about it now, I think I can probably rule out the port on the SAS expander potentially because technically each port has 4 drives connected to it. So if there was an issue with the SAS expander port it stands to reason that it would be 4 drives at a time, not one.

I wish I could provide you some additional help but unfortunately I am in the same, relatively clueless, boat as you are when it comes to diagnosing these sorts of issues. I am going to type a follow up post shortly that details some additional facts I have tracked down that again, I believe, point to something to do with drive #11.
 
Thanks Pirivan. I will continue to follow your thread (and try not to take it too much further off course :).
 
So, after rebooting the server and starting the volume check because it had only gotten to 9% in 6 hours, it ran overnight and finished. Here was what happened during the volume check.

It got to 96% then found two errors and showed this in the logs:

2011-11-06 10:41:12 ARC-1880-VOL#001 Complete Check 009:17:19 2
2011-11-06 10:18:36 Enc#2 Slot#11 Reading Error
2011-11-06 10:18:27 Enc#2 Slot#11 Reading Error
2011-11-06 10:15:12 192.168.001.002 HTTP Log In
2011-11-06 01:23:53 ARC-1880-VOL#001 Start Checking
2011-11-06 01:18:40 192.168.001.002 HTTP Log In
2011-11-06 01:17:36 H/W Monitor Raid Powered On

Usually it goes Time Out Error to slot#11 -> Reading Error for slot #11 then time out of Enc#2 -> Enc#2 being removed and then the RaidSet/Volume Degraded but did not this time (that is a loose description of the past pattern, it has also timed out to other drives as well after timing out FIRST to #11).

So, this time, while it had some reading errors, it moved passed them and did not time out the HP SAS expander or any other drives and finished. It is interesting also that it usually only sees errors when it gets to 90%+. It looks like on the last volume check shown above, it got to 90%+ like it had previously and then had two reading errors, both on drive 11 and found two media errors, also on drive 11 (see below). It is curious that this time (as well as the last time) when timeouts or reading errors occurred on drive 11, nothing else was affected and the array continued to work. Maybe this is due to the power management change that was made previously (disabling time to spin down idle) or removing/re-inserting all the hot swap drives?

So, I decided to take a look at the drives and compare to drive 11 to look for differences. For brevity I have REMOVED all the fields that had the same values for each drive so first here are the 'common' values between all the drives that I removed from the comparison:

Code:
Disk Capacity 	                2000.4GB
Current SATA Mode 	        SATA300+NCQ(Depth32)
Supported SATA Mode 	SATA300+NCQ(Depth32)
Error Recovery Control (Read/Write) 	Disabled/Disabled
Disk APM Support 	        Yes
Device State 	                Normal
Timeout Count 	                0
SMART Seek Error Rate 	100(67)
SMART Spinup Retries 	100(60)
SMART Calibration Retries 	N.A.(N.A.)

Now here are all the drives:
Code:
Device Type 		SATA(5001438006223784)
Device Location 	Enclosure#2 Slot#5
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	43 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	119(24)
SMART Reallocation Count100(5)

Device Type 		SATA(5001438006223785)
Device Location 	Enclosure#2 Slot#6
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	35 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	116(24)
SMART Reallocation Count100(5)

Device Type 		SATA(5001438006223786)
Device Location 	Enclosure#2 Slot#7
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	35 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	115(24)
SMART Reallocation Count100(5)

Device Type 		SATA(5001438006223788)
Device Location 	Enclosure#2 Slot#9
Firmware Rev. 		JKAOA20N
Media Error Count 	0
Device Temperature 	48 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	118(24)
SMART Reallocation Count85(5)

Device Type 		SATA(5001438006223789)
Device Location 	Enclosure#2 Slot#10
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	39 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	116(24)
SMART Reallocation Count100(5)

Device Type 		SATA(500143800622378A)
Device Location 	Enclosure#2 Slot#11
Firmware Rev. 		JKAOA3EA
Media Error Count 	2
Device Temperature 	40 ºC
SMART Read Error Rate 	95(16)
SMART Spinup Time 	118(24)
SMART Reallocation Count98(5)

Device Type 		SATA(500143800622378B)
Device Location 	Enclosure#2 Slot#12
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	41 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	116(24)
SMART Reallocation Count100(5)

Device Type 		SATA(500143800622378C)
Device Location 	Enclosure#2 Slot#13
Firmware Rev. 		JKAOA28A
Media Error Count 	0
Device Temperature 	44 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	124(24)
SMART Reallocation Count100(5)

Device Type 		SATA(500143800622378E)
Device Location 	Enclosure#2 Slot#15
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	39 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	112(24)
SMART Reallocation Count100(5)

Device Type 		SATA(500143800622378F)
Device Location 	Enclosure#2 Slot#16
Firmware Rev. 		JKAOA3EA
Media Error Count 	0
Device Temperature 	41 ºC
SMART Read Error Rate 	100(16)
SMART Spinup Time 	116(24)
SMART Reallocation Count100(5)

Let me know if you want me to remove the code tags, I put them in there for formatting reasons. In any case, looking at it, none of the SMART errors are near the threshold but it is interesting that drive 11 has the 'lowest' SMART Read Error Rate. It is at 95 and everything else is 100. Also, it has a slightly lower SMART Reallocation Count than the others (though so does drive 9, which hasn't seemed to have been a problem).

I am not sure how much of any of this is a real concern as the SMART rates seem to vary a bit from time to time. For example, drive #11 is now at SMART Read Error Rate 100(16) but was at 95(16) when I copied the above. And drive #09 is now at SMART Read Error Rate 98(16) but was at 100(16) when I copied the above. Looking at the error logs since the all started happening while drive #09 has been 'removed' from the array it has never timed out, so maybe the SMART values aren't of a huge concern.

While I wait for the new Hitachi 2TB drive to do a full format (just to be safe) what would be the best way to plan to remove drive 11 and add in the new drive? Just take out drive 11 and 'fail' it and put in the new drive? Or add the new drive as a hotspare, let it get setup and then remove drive 11? I'm not totally convinced that drive 11 is bad but it seems like a good troubleshooting step to remove it based on the behavior so far.

--EDIT--
As there is some data I need to copy over the backup, after the volume check finished, I shutdown, attached a small fan to the HP SAS expander heatsink using a lame zip-tie (seems to work I guess) and started a chkdsk. Once that finishes I am going to try to update my backups before I proceed (assuming I don't have additional timeouts/issues).
 
Last edited:
Well unfortunately it looks like I have to speed up my intended plan. While the chkdsk was going on I got another beep from the server and slot #11 again timed out, followed by slot #9 and then the SAS expander. The only difference was that this time slot #11 also failed (had never done this before) and the RaidSet/volume degraded.

So, even though it wasn't done doing a full format (I liked to do this to test a new drive) I decided to throw the Hitachi HDS723020BLA642 (different than all my other older 2TB 7200RPM Hitachi drives) into there and rebuild. The rebuild process is not as clear as I would like it to be I wish Areca would spell it out somewhere in the UI or anywhere in the manual. It didn't startup after a reboot so I just added the new drive in as a hot spare for the RaidSet and as soon as I did that it started the rebuild. The rebuild is at 2.1% now so it may be a while.

Here is the log output from what I just described above happening:

Code:
2011-11-06 15:07:09 	ARC-1880-VOL#001 	Start Rebuilding 	  	 
2011-11-06 15:07:07 	RAID6 2TB DRIVES 	Rebuild RaidSet 	  	 
2011-11-06 14:53:55 	192.168.001.002 	HTTP Log In 	  	 
2011-11-06 14:53:44 	H/W Monitor 		Raid Powered On 	  	 
2011-11-06 14:32:26 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-06 14:32:26 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-06 14:32:10 	Enc#2 SES2Device 	Time Out Error 	  	 
2011-11-06 14:32:01 	Enc#2 Slot#9 		Time Out Error 	  	 
2011-11-06 14:31:53 	Enc#2 Slot#11 		Device Failed 	  	 
2011-11-06 14:31:52 	RAID6 2TB DRIVES 	RaidSet Degraded 	  	 
2011-11-06 14:31:52 	ARC-1880-VOL#001 	Volume Degraded 	  	 
2011-11-06 14:31:33 	Enc#2 Slot#11 		Time Out Error

Right now I am feeling pretty confident that drive #11 has an issue. I put it into a USB carrier and connected it to another machine just to see if it was totally dead and it was viewable in disk management. So it is possible that it has SOME issues but is not totally dead. Hopefully I will still be able to RMA it if it is still technically accessible but causes issues in my array. I will wait until the rebuild finishes with the new drive before I try to get it hooked up to another box and run Hitachi diagnostics on it (I hate that tool by the way, it will only detect hard drives connected to ONE motherboard I have and you can't use it with anything in a USB enclosure).

At the very least if the new drive in slot #11 has issues it will point to that it is potentially that part of the backplane/cable and not the drive itself.

I admit I am also a little concerned about the drive in slot #9 as well because it timed out when #11 and the expander did AND it has this in the status information for the drive:
SMART Reallocation Count 85(5)

Drive #11 looked like this when it 'failed' before I took it out:
Media Error Count 0
Device Temperature 34 ºC
SMART Read Error Rate 94(16)
SMART Spinup Time 117(24)
SMART Reallocation Count 98(5)

I have two more drives coming in, but they won't get here until next week. I may have to go down to Fry's again and pay the exorbitant cost for another 2TB Hitachi ($190 or so with tax, now is a terrible time to need drives).
 
Last edited:
Well with the drive that was in slot #11 removed and a new drive installed the rebuild finished without any issues AND a follow up volume check finished; 0 errors no timeouts etc. What is truly bizarre is that the drive appears to have been causing a lot/many of the issues but I am not sure that it is failed/bad. I am running a chkdsk on the volume now which is going INCREDIBLY slowly. It is running at approximately 1 file per second and is using ALL available RAM on the server for some reason. I am guess I will just let it run; however tedious (it will probably take like 2-3 days at this rate).

In any case, here is what drive 11's status was according to the RAID array before I took it out:

Device Type SATA(500143800622378A)
Device Location Enclosure#2 Slot#11
Disk Capacity 2000.4GB
Current SATA Mode SATA300+NCQ(Depth32)
Supported SATA Mode SATA300+NCQ(Depth32)
Error Recovery Control (Read/Write) Disabled/Disabled
Disk APM Support No
Device State FAILED
Timeout Count 1
Media Error Count 0
Device Temperature 34 ºC
SMART Read Error Rate 94(16)
SMART Spinup Time 117(24)
SMART Reallocation Count 98(5)
SMART Seek Error Rate 100(67)
SMART Spinup Retries 100(60)
SMART Calibration Retries N.A.(N.A.)

As you can see, it was marked as FAILED and it has some SMART Read Errors and SMART Reallocation Count issues.

I ran an HDDScan SMART status check on the drive with it connected to another PC and found the following:
Code:
	Num	Attribute Name			Value	Worst	Raw (Hex)		Threshold
	005 	Reallocation Sector Count 	098 	098 	0000000000-00C2 	005
	196 	Reallocation Event Count 	098 	098 	0000000000-00CA 	000
	197 	Current Pending Errors Count 	100 	100 	0000000000-000C 	000

This again confirms the SMART errors the RAID controller saw. However, so far doing an HDTune scan on it has not turned up any bad blocks. The drive was clearly causing issues with the array but I am not sure how much 'evidence' I will need to get Hitachi to RMA it.

This also makes me nervous about the drive in slot 9. It has some slightly similar SMART status on the RAID controller:
SMART Reallocation Count 85(5)

I am afraid that it might also start to timeout/have read errors. As soon as I get some other drives in I will likely make sure the drives are good and then add another one in as a hotspare, pull slot #9, rebuild and run HDDScan on the drive from slot 9 to see if it has similar SMART values. Is there a way to do this more gracefully or is that really the best way to get a drive out of the array?
 
Well the chkdsk also finished, no errors. I am a bit worried about the SMART status on drive 9 but it will be a few more days until another Hitachi drive arrives so I will just keep an eye on it and keep my backups updated.

I sent more or less all the information from my previous page or so of posts over to Areca and they had this helpful tidbit to add:

"Thanks for the detailed information. Again, you have large array built from desktop version hard disk drive. Reliability would be my concern. "

I kind of figured I would get to the point where they would harp on this, I just wish they had answered a few of the questions. At least they can possibly tell me the best way to migrate away from drive #9 without failing it; we shall see. It hasn't shown any problems since I took out drive #11 but the SMART Reallocation Count 85(5) is troubling.

I also contacted Hitachi about the SMART errors on my drive to see if an RMA was appropriate, so far no response. I may have to submit an RMA request or call in because the drive is literally useless to me. I'm not going to put a drive that was causing timeouts/read errors and causing the entire SAS expander to timeout and the array to potentially fail back INTO my array or into any other PC for that matter. We'll see what they say.
 
I am not sure if this matters to anyone but in case it does or for future reference, everything has been working great without drive #11 installed. Hitachi was not at all helpful regarding RMA'ing drive #11 (it passed their DFT, which is all they care about, they claim that 'SMART' is not indicative of issues for their drives and they simply did not address any of the points I made about the RAID issues the drive caused) or drive #9 but I went ahead and submitted drive #11 for an RMA with detailed information inside the RMA about how it had caused significant issues with my RAID array all of a sudden and the SMART errors.

In ways I wish drive #11was actually failing DFT because then it would be a more straightforward RMA but we shall see. As soon as I get in some replacement drives and test them with DFT/look for SMART status errors in HDDScan, I will probably pull drive #9 as well and let the array rebuild onto the new drive as a hotspare (apparently that is the 'recommended' way to replace a non-failed drive according to Areca). If it has some of the same SMART warning signs as drive #11 I don't want it to cause all the same issues. We'll see how much of a hassle Hitachi gives me in RMA'ing that as well.

Beyond that I have some new masscool 80mm fans coming in (my alternative to the 120mm fan bracket that doesn't fit my Norco 4020 case) that I will install to see if I can keep the drive temps a bit lower.
 
Well, with the 80mm Masscool's installed the drive temps are quite a bit better. Not great but better.

Per Jus's suggestion I also set "Disk Write Cache Mode" to "enabled" as I don't have a BBU and with it set to automatic it should have been enabled already.

In any case the only other real issue I am facing is that somewhere along the way whenever I now turn the system on it does the following:
"Waiting for fw to become ready" and then goes the max (300 seconds) times out and reboots.

Then when it gets to "Waiting for the FW to become ready" the 2nd time it takes about 20 seconds (normal) and then moves on and boots fine. I am not sure why it is timing out on the first boot attempt with the RAID card.

I did upgrade the firmware and change the setting "Time To Spin Down Idle HDD" from 60 to disabled during my troubleshooting process previously but I am not certain that either setting caused the timeout issue on first boot.
 
Back
Top