Im currently building my Adaptec 5805 with RAID 5 on 8xWD 2TB RE4 disks on linux. I have previously had it running in RAID 5 on 6x2TB.. but I had to rebuild the whole array because of a bad striping report.
Suddenly Adaptec StorMan is starting to report 10-15 of thoes in the last 1-2 hours (from the same physical drive):
07 October 2011 22:43:07 CEST INF laffy Sense data: Medium error (UNRECOVERED READ ERROR). Controller 1, channel 0, SCSI device ID 3, LUN 0, cdb [28 00 27 9a 6b 00 00 01 00 00 00 00], data [70 00 03 00 00 00 00 00 00 00 00 00 11 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00]
07 October 2011 22:43:08 CEST WRN laffy Medium error: controller 1, channel 0, SCSI device ID 3, LUN 0, start LBA 279a6b00, end LBA 279a6bff, bad block recovery possible
07 October 2011 22:43:08 CEST WRN laffy Bad block recovery: controller 1, channel 0, SCSI device ID 3, LUN 0, bad block recovery completing
07 October 2011 22:43:08 CEST WRN 418:A01C-S--L-- laffy Bad Block discovered: controller 1 (279a6b00).
07 October 2011 22:47:15 CEST INF laffy Running: RAID 5 scrub - 0%. 0 different sectors. Controller 1, logical device 0
07 October 2011 22:47:15 CEST INF 307:A01C-S--L00 laffy Building/Verifying: controller 1, logical device 0 ("device 0").
However when I check the smart data on the disk using smartctl I got the following:
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
1 Raw_Read_Error_Rate 0x002f 200 200 051 Pre-fail Always - 0
3 Spin_Up_Time 0x0027 253 253 021 Pre-fail Always - 8991
4 Start_Stop_Count 0x0032 100 100 000 Old_age Always - 671
5 Reallocated_Sector_Ct 0x0033 200 200 140 Pre-fail Always - 0
7 Seek_Error_Rate 0x002e 200 200 000 Old_age Always - 0
9 Power_On_Hours 0x0032 091 091 000 Old_age Always - 7046
10 Spin_Retry_Count 0x0032 100 100 000 Old_age Always - 0
11 Calibration_Retry_Count 0x0032 100 100 000 Old_age Always - 0
12 Power_Cycle_Count 0x0032 100 100 000 Old_age Always - 129
192 Power-Off_Retract_Count 0x0032 200 200 000 Old_age Always - 126
193 Load_Cycle_Count 0x0032 200 200 000 Old_age Always - 544
194 Temperature_Celsius 0x0022 119 111 000 Old_age Always - 33
196 Reallocated_Event_Count 0x0032 200 200 000 Old_age Always - 0
197 Current_Pending_Sector 0x0032 200 200 000 Old_age Always - 2
198 Offline_Uncorrectable 0x0030 200 200 000 Old_age Offline - 0
199 UDMA_CRC_Error_Count 0x0032 200 200 000 Old_age Always - 0
200 Multi_Zone_Error_Rate 0x0008 200 200 000 Old_age Offline - 0
Current_Pending_Sector has been on the value 4 but was reduced to 2 (im running a smartctl -t offline on the disk as well).
Should I be very worried? like shutting down the server, put the harddrive into my desktop and check with WD diagnostic tool? Or should I be okay as long smart dosent report increasing Current_Pending_Sector / Offline_Uncorrectable?
Suddenly Adaptec StorMan is starting to report 10-15 of thoes in the last 1-2 hours (from the same physical drive):
07 October 2011 22:43:07 CEST INF laffy Sense data: Medium error (UNRECOVERED READ ERROR). Controller 1, channel 0, SCSI device ID 3, LUN 0, cdb [28 00 27 9a 6b 00 00 01 00 00 00 00], data [70 00 03 00 00 00 00 00 00 00 00 00 11 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00]
07 October 2011 22:43:08 CEST WRN laffy Medium error: controller 1, channel 0, SCSI device ID 3, LUN 0, start LBA 279a6b00, end LBA 279a6bff, bad block recovery possible
07 October 2011 22:43:08 CEST WRN laffy Bad block recovery: controller 1, channel 0, SCSI device ID 3, LUN 0, bad block recovery completing
07 October 2011 22:43:08 CEST WRN 418:A01C-S--L-- laffy Bad Block discovered: controller 1 (279a6b00).
07 October 2011 22:47:15 CEST INF laffy Running: RAID 5 scrub - 0%. 0 different sectors. Controller 1, logical device 0
07 October 2011 22:47:15 CEST INF 307:A01C-S--L00 laffy Building/Verifying: controller 1, logical device 0 ("device 0").
However when I check the smart data on the disk using smartctl I got the following:
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
1 Raw_Read_Error_Rate 0x002f 200 200 051 Pre-fail Always - 0
3 Spin_Up_Time 0x0027 253 253 021 Pre-fail Always - 8991
4 Start_Stop_Count 0x0032 100 100 000 Old_age Always - 671
5 Reallocated_Sector_Ct 0x0033 200 200 140 Pre-fail Always - 0
7 Seek_Error_Rate 0x002e 200 200 000 Old_age Always - 0
9 Power_On_Hours 0x0032 091 091 000 Old_age Always - 7046
10 Spin_Retry_Count 0x0032 100 100 000 Old_age Always - 0
11 Calibration_Retry_Count 0x0032 100 100 000 Old_age Always - 0
12 Power_Cycle_Count 0x0032 100 100 000 Old_age Always - 129
192 Power-Off_Retract_Count 0x0032 200 200 000 Old_age Always - 126
193 Load_Cycle_Count 0x0032 200 200 000 Old_age Always - 544
194 Temperature_Celsius 0x0022 119 111 000 Old_age Always - 33
196 Reallocated_Event_Count 0x0032 200 200 000 Old_age Always - 0
197 Current_Pending_Sector 0x0032 200 200 000 Old_age Always - 2
198 Offline_Uncorrectable 0x0030 200 200 000 Old_age Offline - 0
199 UDMA_CRC_Error_Count 0x0032 200 200 000 Old_age Always - 0
200 Multi_Zone_Error_Rate 0x0008 200 200 000 Old_age Offline - 0
Current_Pending_Sector has been on the value 4 but was reduced to 2 (im running a smartctl -t offline on the disk as well).
Should I be very worried? like shutting down the server, put the harddrive into my desktop and check with WD diagnostic tool? Or should I be okay as long smart dosent report increasing Current_Pending_Sector / Offline_Uncorrectable?