• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Drive failing in my RAID5 array?

SilverMK3

[H]ard|Gawd
Joined
Dec 15, 2002
Messages
1,346
I woke up the other night to a horrible 'alarm' noise coming from my computer room. My server was hard-locked and beeping so I rebooted it. Here's what I've been able to pull from the logs:

HighPoint RocketRAID Admin console:
Code:
Event View 
 Date Time Description 
 2006/12/14 17:46:0 Disk 'ST3320620AS' at Controller1-Channel2 failed. 
 2006/12/14 17:48:33 Array 'CORTEX_TERABYTE' rebuilding started. 
 2006/12/14 21:29:39 Array 'CORTEX_TERABYTE' rebuilding completed. 
 2006/12/14 21:42:55 Disk 'ST3320620AS' at Controller1-Channel2 failed. 
 2006/12/14 21:46:45 Array 'CORTEX_TERABYTE' rebuilding started. 
 2006/12/15 1:19:50 Array 'CORTEX_TERABYTE' rebuilding completed. 
 2006/12/16 17:36:24 Array 'CORTEX_TERABYTE' verifying started.

WinXP Event Viewer:
Code:
Event Type:	Error
Event Source:	rr1740
Event Category:	None
Event ID:	9
Date:		14/12/2006
Time:		5:10:31 PM
User:		N/A
Computer:	CORTEX
Description:
The device, \Device\Scsi\rr17401, did not respond within the timeout period.

For more information, see Help and Support Center at [url]http://go.microsoft.com/fwlink/events.asp[/url].
Data:
0000: 00 00 10 00 01 00 66 00   ......f.
0008: 00 00 00 00 09 00 04 c0   .......À
0010: 01 01 00 50 00 00 00 00   ...P....
0018: 00 00 00 00 00 00 00 00   ........
0020: 00 00 00 00 00 00 00 00   ........
0028: 00 00 00 00 00 00 00 00   ........
0030: 00 00 00 00 07 00 00 00   ........

Does this mean one of my 7200.10's is failing or is this something that Western Digital's TLER would alleviate if I had WD drives?
 
I do not know if the seagate hdds have the TLER problem, but I have not heard anything about it before. If I were you, I would run the Seagate diagnostics on that drive, maybe taking the array offline in the meantime or using a spare if you have one.
 
SilverMK3 said:
Does this mean one of my 7200.10's is failing or is this something that Western Digital's TLER would alleviate if I had WD drives?


I think you have that backwards...the TLER on WD drives cause drive drop offs even though there is nothing wrong with them. WD's raid edition (RE) drives do not suffer from this problem.
 
hardwarephreak said:
I think you have that backwards...the TLER on WD drives cause drive drop offs even though there is nothing wrong with them. WD's raid edition (RE) drives do not suffer from this problem.


Actually, I think you have it backwards. TLER = Time Limited Error Recovery. If the drive can't read a sector within a given time it will give up and call that sector bad. Non-TLER will keep trying to recover for a few minutes, causing dropouts on RAID controllers.
 
You are correct...wow...started drinking too early in the day to respond to this post. As evidenced by the actual meaning of the acronym TLER.

Nothing to see here, move along....
 
So are these errors common or are they indicative of a dying drive?
I'm downloading the SeaTools ISO as I type this. I'm assuming I'll have to pull the drive from the array to do the diagnostics...
 
SilverMK3 said:
So are these errors common or are they indicative of a dying drive?
I'm downloading the SeaTools ISO as I type this. I'm assuming I'll have to pull the drive from the array to do the diagnostics...
most likely yes.
 
So I guess my drive is toast... Horrible clicking and squealing noises accompanied my 2hr torture/diagnostic test. Here are the results:
Code:
SeaTools Desktop v3.02.04
Copyright (c) 2005 Kroll Ontrack Inc.

12/18/2006 @ 8:15 PM

The following information has been generated by SeaTools Desktop.  Use
this information to help you recognize and resolve potential data access
problems.


System Information:
BIOS Date                 10/27/05
Conventional Memory size   626 K
Extended Memory size      58532 K
IO Channel type            PCI



Drive Information:
SIZE         MODEL
---------    ---------------------
320 GB       ST3320620AS                             


Serial Number = 5QF0XC68


Diagnostic Results:

Seagate DiagATA Quick Test Result:  Failed
    Recommendation:
    The "Quick Test" is adequate for most situations.
    Consider running the "Full Test" which
    verifies each sector on the drive if you need to run a more
    comprehensive diagnostic.



Results from Seagate's DiagATA/SCSI:
-----------------------------------------------------------------

              DIAGATA.EXE Version 3.08.50629ML
Copyright (c) 2002-2005 by Seagate Technology LLC.  All rights reserved.

-----------------------------------------------------------------
Timer Resolution: 0.000122
Short Test Begin: 18-Dec-2006 19:57:47
Cable Test - 0 Errors
Buffer Test - 0 Errors
Identify Data
   Model Number: ST3320620AS
   Serial Number: 5QF0XC68
   Firmware Revision: 3.AAC
   Default CHS: 16383-16-63
   Current CHS: 16383-16-63
   Current Capacity: 16514064 Sectors
   Total Capacity: 625142448 Sectors
   ID Method: Unknown
SMART Check: Passed
DST Poll Time = 60 seconds
DST - Errors - Status: 07
SMART Check: Passed
Short Test Failed: 18-Dec-2006 19:58:49


-----------------------------------------------------------------
End results from Seagate's DiagATA/SCSI

ATA Full Test Result:  Failed



Results from Seagate's DiagATA/SCSI:

-----------------------------------------------------------------

              DIAGATA.EXE Version 3.08.50629ML
Copyright (c) 2002-2005 by Seagate Technology LLC.  All rights reserved.

-----------------------------------------------------------------
Timer Resolution: 0.000122
Long Test Begin: 18-Dec-2006 20:00:32
Cable Test - 0 Errors
Buffer Test - 0 Errors
Identify Data
   Model Number: ST3320620AS
   Serial Number: 5QF0XC68
   Firmware Revision: 3.AAC
   Default CHS: 16383-16-63
   Current CHS: 16383-16-63
   Current Capacity: 16514064 Sectors
   Total Capacity: 625142448 Sectors
   ID Method: Unknown
SMART Check: Passed
Full Scan (0 to 4580810) - Errors
   ---------------------------
   |LBA       |Error Register|
   ---------------------------
   |200858    |10h           |
   |205375    |10h           |
   |208359    |10h           |

  <<<<< ---- Clipped ---- >>>>>>

   |4580809   |10h           |
   ---------------------------
Long Test Failed: 18-Dec-2006 20:15:38


-----------------------------------------------------------------
End results from Seagate's DiagATA/SCSI



******************************************


Recommendation:
If you are not experiencing data loss and SeaTools reports File 
System Structure errors, they may be caused by a lock-up or 
failure to shutdown Windows correctly. Many times, these errors
may be repaired through normal system maintenance which
includes using the Windows provided "Defrag" and
"Scandisk / Chkdsk / Error Checking" utilities.

If you are experiencing a hardware error, you should isolate
the cause and replace the failing component. If you are unsure how
to proceed with repairs, contact a computer professional. After
completing any maintenance tasks, run SeaTools again to
verify that all errors have been repaired. If errors continue to
occur, the system may not be stable. Again, contact a computer 
professional.

If you have experienced a data loss, cease drive operation 
immediately.  Professional data recovery service is the best 
option to recover your data.


========================================================


What's the best way to deal with this? RMA the drive and just slap the replacement back into the array? It'll re-build itself with no issues?
 
SilverMK3 said:
What's the best way to deal with this? RMA the drive and just slap the replacement back into the array? It'll re-build itself with no issues?
Depends, is uptime important to you? If you need uptime protection, then you should consider getting a replacement drive from a vendor while you wait for the RMA. You could then use the RMA-replacement as a hot or cold spare, provided that you 'trust' it.
 
Back
Top