SATA drive failing?

danswartz

2[H]4U
Joined
Feb 25, 2011
Messages
3,715
I have a 4-disk raidz for non-critical storage on my zfs server. 640GB WD blue drives. The host OS (centos7) was warning me about smart issues on one drive. I ran a long smart test on it and it says:

# 2 Extended offline Completed: read failure 20% 20915 1216029173

Various articles tell me to write a block of zeros to that sector and it will remap it. It doesn't. Here are the (I believe) germane smart fields:

Reallocated_Sector_Ct 0
Current_Pending_Sector 21
Offline_Uncorrectable 1

The drive shows as 'healthy'. As I alluded to, if I write to it, it doesn't remap. The dd fails with 'input/output error', and the above counts do not change. This doesn't sound good?
 
3rd-ed. I just had a similar issue with a Hitachi 2TB drive. It was third-tier storage so no RAID (JBOD) but it would keep dropping out of Windows Explorer and subsequently Disk Management.

Replaced it the next day (Amazon Prime to the rescue) and haven't had an issue since. Don't risk it unless you don't care about your data. I'm not an extremist who is like "OMG every second counts. Get it in a bag in the freezer right now" (or whatever new storage voodoo is in vogue right now) - BUT I did wait an extra day once when a drive started flaking out and I ended up losing about 60GB of unrecoverable (and irreplaceable) media of family and friends. This was before Crashplan and also before backups were commonplace because the drives were so expensive.

Short version: Replace it ASAP, like jrweis said
 
I agree. Replace drive, once sectors go bad more will follow.

I (a person who has done 75+ RMAs) say not always. There may be an isolated media defect or a mechanical problem. I always test to verify that the drive is not usable before I RMA. One reason is you do not always get a better drive back than the one you sent in..
 
i appreciate the answers. i would like to know why the recommended method for clearing up the pending sectors doesn't seem to work? honestly, that bothers me more than anything else (it implies the drive is not doing something it should be, no?)
 
just to clarify, there absolutely IS at least one unwriteable sector. if i try writing all zeros to the device, it bombs out on the sector i showed. and the drive will NOT remap it apparently. most likely the only reason i haven't noticed is this is a brand-new pool and hence contains no real data yet...
 
The dd fails with 'input/output error', and the above counts do not change.

Is this a read or a write? I would expect an input error on the 21 sectors that the drive can not read.

The drive shows as 'healthy'

By experience it will show up as healthy almost to the point that it becomes totally unreadable meaning the full drive SMART pass / fail is nearly useless.

if i try writing all zeros to the device, it bombs out on the sector i showed.

I have not seen that too often. Most of the time with the drives that I have RMA'd I could write but the drive had new pending sectors each full disk write followed by a read. In my case this was a badblocks 4 pass test. If a drive can not pass that it is useless to me even in 3 parity raid.
 
Last edited:
The howto I saw said to remap the sector by doing 'dd if=/dev/zero of=/dev/sdXXX seek=N count=1'. I do that, and it hangs for several seconds and then fails with 'input/output error' (EIO?) And the counts from smart are unchanged. e.g. 'no remap for you!'
 
i appreciate the answers. i would like to know why the recommended method for clearing up the pending sectors doesn't seem to work? honestly, that bothers me more than anything else (it implies the drive is not doing something it should be, no?)

I use badblocks and let run in 1-2 days.
when I see bad blocks are displayed, I know that those pending would be move to relocate or back to normal

badblocks -p100 -w -v

2 weeks ago, my centos was complaining on a failing HD on my raidz2
I Swapoed with a "new" drive, and ran badblock on a failing HD for ~1 1/2 days, without errors. and recheck smart status.. all are clear...
I Swapped back the drive to raidz2..
 
cantalup, that doesn't explain anything though. as far as i know, sata drives are supposed to remap a bad sector when you write to it. this one doesn't. that sounds broken, no?
 
cantalup, that doesn't explain anything though. as far as i know, sata drives are supposed to remap a bad sector when you write to it. this one doesn't. that sounds broken, no?

you need to force the drive read and write a little bit low.
the firmware(in drive) will remap or make it good (after reading succefully).

badblocks will write/read by sector

try badblocsk first with random write within 1-2 days...( I use -p100 -w -vv)
when you see badblocks see some bads, those would be remap (relocating) to spare space

I did 2 drive , one was clear out no relocation, one was shown bad block( relocating was increased beyond threshold level, aka going to trash/basura)

good luck!
 
The howto I saw said to remap the sector by doing 'dd if=/dev/zero of=/dev/sdXXX seek=N count=1'. I do that, and it hangs for several seconds and then fails with 'input/output error' (EIO?) And the counts from smart are unchanged. e.g. 'no remap for you!'
Yes, that is (precisely) the correct method. [drescherjm: Since Dan did not specify a bs=N (blocksize), dd uses its default of 512 (bytes), which is appropriate for this command. (equal to LBA size)]

[Hypothesis:] The drive's firmware is flawed--it is not anticipating the case where the "manufacturing-level" formatting for that LBA is glitched such that the firmware can not even verify the LBA# (and that case is [mistakenly!] not included in this firmware's "it's bad, I'll remap" logic). That step occurs before any attempt to actually write (or read) data.

I wonder if the other 20 pending-sectors are the same story. You would need to do a dd read with a skip=N (or dd write with seek=N) -- assuming that this LBA in question is the lowest of the 21.

Does a SMART report show an Error Log ? Some models of WDC drives are notorious for not even logging errors.

--UhClem
 
It shows no errors other than the ones from trying to remap it, but then again, that drive has been in spares for several years, and I highly doubt I ever used up to that 80% (or maybe those sectors even went bad after I shelved the drive.) Not being able to remap that one sector (even if the others can) makes this trash as far as I am concerned. Thanks for the sanity check!
 
Back
Top