• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Stress test new HDDs?

Snowdog

[H]F Junkie
Joined
Apr 22, 2006
Messages
11,257
So far I seem to be running 50% failures within a month of buying a new big (2TB+) HDD.

Last year it was 2TB WD Green that failed after 20 days.

Now it is a Seagate 3TB after about a week.

Are we back to the days where you need to stress test new HDDs before you can trust them?
 
I've been told the best way to check a drive before use is to run badblocks. I have not done this myself, but it was mentioned in almost every thread about drive verification I read when looking to check my own 3tb drives.

Personally, I use Western Digital's Data Life Guard tool to write 0's to the disc, full verify, and then full format before use. This has served me well on the last four 3tb drives I put into my htpc.
 
So far I seem to be running 50% failures within a month of buying a new big (2TB+) HDD.

Bad luck or poor shipping..


Are we back to the days where you need to stress test new HDDs before you can trust them?

I believe so and I have around 200 drives here at work spinning 24/7/365. For the last 4 or so years I have RMA'd 10 to 20 drives each year because they failed and then failed to pass my rigorous testing with the linux utility called badblocks.

For users with some understanding of how linux devices work and can type commands from the shell / command line I recommend a 4 pass badblocks read / write test then look at the SMART attributes using smartctl. For windows users that are not comfortable with linux or do not understand the shell/command line I recommend a few full formats followed by looking at the SMART data with a program like CrystalDiskInfo.
 
Bad luck or poor shipping..




I believe so and I have around 200 drives here at work spinning 24/7/365. For the last 4 or so years I have RMA'd 10 to 20 drives each year because they failed and then failed to pass my rigorous testing with the linux utility called badblocks.

For users with some understanding of how linux devices work and can type commands from the shell / command line I recommend a 4 pass badblocks read / write test then look at the SMART attributes using smartctl. For windows users that are not comfortable with linux or do not understand the shell/command line I recommend a few full formats followed by looking at the SMART data with a program like CrystalDiskInfo.

Same deal here. I fire up PartedMagic and test with badblocks. No need to use smartctl (directly at least) as PartedMagic has a nice GUI interface for viewing SMART data. It also has pretty much any other data storage tool bundled as well (Clonezilla, Gparted, Photorec, hdparm, etc)
 
Bad luck or poor shipping..

Seems to be a lot of bad luck going around on 2TB+ drives. I never mail order HDDs, these were picked up at a local shop who gets them in bulk, likely with better shipping than individual shipping to your house.


I believe so and I have around 200 drives here at work spinning 24/7/365. For the last 4 or so years I have RMA'd 10 to 20 drives each year because they failed and then failed to pass my rigorous testing with the linux utility called badblocks.

How many of those are >2TB consumer drives. I have noticed a drop in reliability as the density has gone up.

For users with some understanding of how linux devices work and can type commands from the shell / command line I recommend a 4 pass badblocks read / write test then look at the SMART attributes using smartctl. For windows users that are not comfortable with linux or do not understand the shell/command line I recommend a few full formats followed by looking at the SMART data with a program like CrystalDiskInfo.

I have no issues using Linux (been using it since Slackware in the early 1990's), but it isn't my main OS and this is my only computer. So taking 2 days to just run badblocks via liveCD, is two days without my computer, which is also my HTPC/internet.

So a windows stress test would be more handy.

Also even though Windows is throwing up a failure message every 15 mins and Seagate self test, and smart test report a failure, a SMART read with CrystalDiskInfo show no issues:

I am seeing zeros on Reallocated_Sector_Ct, Current_Pending_Sector, Uncorrectable_sector_Ct, Reported_Uncorrectable_errors and UDMA_CRC_Error_Count.
 
Last edited:
How many of those are >2TB consumer drives. I have noticed a drop in reliability as the density has gone up.

The RMAs on 2TB drives have been low. Although only 20 to 30 are 2TB+. I think we had 1 DOA and that is it. Most of the 2TB drives are hitachi and now a few toshiba enterprise drives. While the bulk of the failures were with Seagate 7200.10, 7200.11 and 7200.12 drives. That is not to say all of the failures were Seagate drives, we have had failures from every manufacturer except toshiba..
 
Last edited:
Anyway the Seagate is gone, replaced by a WD 3TB GP. They had no Seagates left so that made up my mind.

So Tosh/Hitachi Enterprise drives are good. That doesn't do much for the average consumer buy consumer WD/Seagates with 1 or 2 year warranties.

It seems like QA/reliability/warranties are all dropping on these drives.
 
It seems like QA/reliability/warranties are all dropping on these drives.

For desktop drives (home or work) I will not purchase any drive with less than a 3 year warranty.
 
For desktop drives (home or work) I will not purchase any drive with less than a 3 year warranty.

Where does that leave people? There are no 3TB WD Black drives. Green has 2 year warranty. Red is for NAS. The Seagate 3TB had a 1 year warranty.
 
I usually run dd. I'll do a full write of zeros, then a full read to /dev/null. I'll monitor dmesg for any io errors and check the smart status.

But yeah seems drive reliability has really taken a dump in the past years. :(
 
I have no issues using Linux (been using it since Slackware in the early 1990's), but it isn't my main OS and this is my only computer. So taking 2 days to just run badblocks via liveCD, is two days without my computer, which is also my HTPC/internet.

You can just do a single write-read pass of badblocks if you cannot wait for multiple passes to complete.
 
Prime95 or Orthos for a recommended 12 hours or more at least, make sure there is no massive vdroop otherwise you will get instability.
 
I read somewhere that 4 pass badblocks takes 70 hours on a 1 TB.

That is 15 hours/pass/TB so it would still be 45 hours on 3TB drive for a single pass.
 
I read somewhere that 4 pass badblocks takes 70 hours on a 1 TB.

That is 15 hours/pass/TB so it would still be 45 hours on 3TB drive for a single pass.

It takes around 25 hours for a 4 pass badblocks on a 7200 RPM 2TB drive. When I get a new batch I run multiple instances of badblocks 1 instance per drive simultaneously.
 
it took 86 hours for a 5K4000 on my system
this was with the default # of passes

http://i.imgur.com/nvGg1.jpg

I would look into that. It sounds like a problem with the drive or interface if it is running that slow. iotop should show the current data rate for the drive while badblocks is run. I look at this for my drives where they start out at 155MB / s and end somewhere in the 80MB/s to 90MB/s.
 
I just ran the full 4 tests (badblocks) on 5 x 3TB toshiba drives, all tests completed with 0 errors, and took right about 200 hours to run... these were all hoooked up via usb (as they were external drives) this would be faster on an internal bus - but i didnt want to void the warranty by ripping the enclosures open just in case.

btw had these hooked up to three laptops... 1 drive on one, 2 drives on two all going at the same time.

peace of mind, before putting them into my raid 6... ill be doing this from now on...
 
For all you badblocks fan-boys ... you are definitely stressing your disks (not a bad thing) but you are not testing them.

See this (currently evolving) thread [link] for the (smelly) details. Quick summary (as exposed here [link]): badblocks doesn't report squat :(. (except, maybe, the infinitessimally rare occurrence of a drive's ECC mechanism not reporting an error [due to ECC "hash collision"])


============================
CORRECTION: [28Nov12]

badblocks does report (very) bad blocks, and, for that, it can not be faulted. The test scenario upon which I based my [(now) misplaced] criticism, involved a drive that was throwing UNCorrectable errors, and corresponding nullifying increases/decreases to SMART's Current_Pending_Sector count, BUT none of those (15+) flaky sectors were sufficiently persistent in their flakiness to cross the AHCI driver's RETRY threshold, and, hence, did not result in any error returns to the calling program's (ie, badblocks) read() requests. Thus, badblocks had nothing to report.

[It is unfortunate that there is no mechanism available for a (privileged) user program to be informed of drive errors (below the RETRY threshold), provoked by its own read()s. I'll mention this to the author of badblocks, who is also heavily involved with kernel/filesystem development.]

This highlights the importance of simultaneously monitoring the tested drive's (mis-) behavior by other means.

=======================================


--UhClem "Everything is fine ... until it isn't."

> Hey, it's nice out.
<< Yeah, I think you oughta leave it out.
 
Last edited:
For all you badblocks fan-boys ... you are definitely stressing your disks (not a bad thing) but you are not testing them.

See this (currently evolving) thread [link] for the (smelly) details. Quick summary (as exposed here [link]): badblocks doesn't report squat :(. (except, maybe, the infinitessimally rare occurrence of a drive's ECC mechanism not reporting an error [due to ECC "hash collision"])

--UhClem "Everything is fine ... until it isn't."

> Hey, it's nice out.
<< Yeah, I think you oughta leave it out.
Do you have a recommendation for an alternative?
 
In the end I just went with a full format, some HD Tune drive tests and now copying a > 1TB of Video files, which will be followed by zone defrag to move video archive to the end of the disk.

This will probably keep the drive active for 48 hours or so.... Then I will start putting new data on it. But I will increase my backup frequency for a few weeks.
 
Do you have a recommendation for an alternative?

His post is misleading. The only salient point is that one should always check the SMART attributes after running badblocks.

The usual procedure is to run badblocks and then check the SMART attributes for reallocated sectors, etc. (see the list of what to check as posted by drescherjm). Actually, best is to record the SMART attributes before and after badblocks.
 
see the list of what to check as posted by drescherjm

Hmm. I remember writing about that yesterday but today I do not see it in this thread. Time to search..

Edit: Anyways here are the attributes that I am concerned at most:
"Reallocated_Sector_Ct" "Current_Pending_Sector" "Offline_Uncorrectable" "UDMA_CRC_Error_Count" "Hardware_ECC_Recovered"

This is taken from a script at:
https://github.com/drescherjm/jmdgentoooverlay/blob/master/Other/shell-scripts/examine_mdraid.sh

and also the script

https://github.com/drescherjm/jmdgentoooverlay/blob/master/Other/shell-scripts/examine_smart.sh
 
Last edited:
Anyway the Seagate is gone, replaced by a WD 3TB GP. They had no Seagates left so that made up my mind.

So Tosh/Hitachi Enterprise drives are good. That doesn't do much for the average consumer buy consumer WD/Seagates with 1 or 2 year warranties.

It seems like QA/reliability/warranties are all dropping on these drives.

Yep you can tell they do not even trust their own drives since the warranties have dropped, some as low as 1 year.
 
Hmm. I remember writing about that yesterday but today I do not see it in this thread. Time to search..

Edit: Anyways here are the attributes that I am concerned at most:
"Reallocated_Sector_Ct" "Current_Pending_Sector" "Offline_Uncorrectable" "UDMA_CRC_Error_Count" "Hardware_ECC_Recovered"


The problem with my recently failed Seagate. All of those were ZERO. I couldn't detect anything amiss with a third party SMART util.

Yet windows was complaining every 15 minutes that I had driver errors, and I needed to backup the drive ASAP. Seagate Tools failed it's own SMART test (which doesn't show any actual values).

But still, nothing on third party SMART utils.
 
His post is misleading.
Not in the least. Especially if your read the referenced post (which gave more details). Of course, to get the full benefit of both posts, you need to have a good understanding of modern disk drives, and an appreciation for the importance of sound software engineering principles. (Maybe you're misleading yourself :))
The only salient point is that one should always check the SMART attributes after running badblocks.
(More ignorance ...)
Firstly, nothing excuses a program whose stated purpose is search a devcie for bad blocks [badblocks/man/8] for not taking the appropriate action, most important being to inform the user, upon getting an error return from a read() system-call. Second, regardless of any SMART tests you perform before and after the badblocks run, depending on what tests you invoked, you could very well have a very flaky and demonstrably untrustworthy drive and not get any indication from those two SMART reports. But being apprised of the read() failures would at least have been a red flag for the astute user.

[My initial posting (here), and this follow-up, are intended as a (prudent) warning. I take no responsibility for educating .]

--UhClem
 
Second, regardless of any SMART tests you perform before and after the badblocks run, depending on what tests you invoked, you could very well have a very flaky and demonstrably untrustworthy drive and not get any indication from those two SMART reports.

And again this guy is misleading people. There is no need to run SMART tests, which are not very useful (as was already discussed).

The procedure, as I already explained, is to record the SMART attribute values, run badblocks, then look at the SMART attribute values again, paying particular attention to the ones drescherjm listed.
 
And again this guy is misleading people. There is no need to run SMART tests, which are not very useful (as was already discussed).

The procedure, as I already explained, is to record the SMART attribute values, run badblocks, then look at the SMART attribute values again, paying particular attention to the ones drescherjm listed.

As I pointed out above, you can have a failing drive that doesn't show any non-zero values in those SMART attributes.
 
As I pointed out above, you can have a failing drive that doesn't show any non-zero values in those SMART attributes.

Is the procedure I mentioned so difficult to understand? Three simple steps. But you skipped the most important one, and you think that proves the procedure is faulty? :confused:
 
I had posted:
Second, regardless of any SMART tests you perform before and after the badblocks run, depending on what tests you invoked, you could very well have a very flaky and demonstrably untrustworthy drive and not get any indication from those two SMART reports.
And again this guy is misleading people. There is no need to run SMART tests, which are not very useful (as was already discussed).
Yes, using the phrase "SMART tests you perform" was a typo/thinko; and I'm puzzled why I would have uttered it, because I've never run a SMART test (but use "smartctl -a" frequently). My intention was "SMART reports you get". (Note the trailing phrase "those two SMART reports".)

I will re-state my position (with more care):
Firstly, nothing excuses a program whose stated purpose is search a device for bad blocks [badblocks/man/8] when, upon getting an error return from a read() system-call, it does not take the appropriate action, most important being to inform the user. Second, regardless of any SMART reports you get before and after the badblocks run, depending on what badblocks options you invoked, you could very well have a flaky and demonstrably untrustworthy drive and not get any indication from those two SMART reports (or a comparative analysis of them). In fact, from a SMART report(s) perspective, badblocks is as likely to destroy evidence of a questionable drive, as it is to expose it. But being apprised of the read() failures would at least be a clear, and valuable, red flag for the astute user.
The procedure, as I already explained, is to record the SMART attribute values, run badblocks, then look at the SMART attribute values again, paying particular attention to the ones drescherjm listed.
The only SMART attributes with any possible meaning, in this discussion, are Reallocated_Sector_Ct and Current_Pending_Sector; but you need to know, and understand, the significance of those two, and the dynamic relationship between them[**]. When you put a badblocks invocation between two SMART reports, you need to add to your knowledgebase, the low-level operations performed by badblocks (on the drive), the sequence of those operations, and how those operations might cause changes to those two attributes.

Else, you are very possibly lulling yourself into a false sense of security (that your drive is healthy). And wasting a lot of time doing it.

[**] If you frequently get sick, but always seem to recover, do you honestly think you're a healthy individual?

--UhClem
 
Second, regardless of any SMART reports you get before and after the badblocks run, depending on what badblocks options you invoked, you could very well have a flaky and demonstrably untrustworthy drive and not get any indication from those two SMART reports (or a comparative analysis of them).

In testing HDDs, almost anything is possible. But the situation you describe is not likely, and is hardly worth worrying about in the context being discussed. Hence, you are again misleading people.
 
In testing HDDs, almost anything is possible. But the situation you describe is not likely, and is hardly worth worrying about in the context being discussed. Hence, you are again misleading people.
Nope, but we'll have to agree to disagree :)

Thanks for reminding me that professionals (even retired ones) should not engage in (highly-)technical discussions with hobbyists.

Remember: Everything is fine ... until it isn't.
If you're going to do something, do it right.
Good enough ... isn't.
 
In testing HDDs, almost anything is possible. But the situation you describe is not likely, and is hardly worth worrying about in the context being discussed. Hence, you are again misleading people.

I have never ever seen that happen. However I have only had a few hundred drives (we keep between 100 to 200 spinning 24/7/365 for the last 15+ years) with clearly over 100 total failures over the years and not 10s of thousands or 100s of thousands.

it does not take the appropriate action, most important being to inform the user.

It does inform the user when what is read is not what is expected (the pattern it wrote in the write phase) however. I see this happen all the time with failing drives.
 
Last edited:
Thanks for reminding me that professionals (even retired ones) should not engage in (highly-)technical discussions with hobbyists.

Hey, thanks for even trying to help out all us dipshits. Shame we're too stupid to understand. I suppose you could educate, but then again, you stated it wasn't your role to educate. I guess that's why you haven't responded to requests for alternatives.

If engaging in discussions here is really bothering you, I wouldn't mind if you don't.
 
As I pointed out above, you can have a failing drive that doesn't show any non-zero values in those SMART attributes.

A clean SMART report doesn't mean the drive won't fail, I think people understand that. Certain SMART indicators are, however, correlated with drive failure probability. The point of running badblocks isn't the badblocks report itself its seeing if it trips any SMART indicators.
 
The point of running badblocks isn't the badblocks report itself its seeing if it trips any SMART indicators.

I see you point and its kind of what I say for windows users.

In my own testing I also take the badblocks report into consideration. The reason is a modern drive should not have any OS visible bad blocks after writing every single sector for a few passes. If there were weak sectors they should have been remapped so the OS does not see the sectors. When I run badblocks I sometimes allow for that with the -p option (which from memory is number of passes to complete when badblocks still exist).
 
BTW: here is some example output from a drive that was kicked out of a 6 to 7TB raid6 array consisting of 9 1TB WDC black drives (8 + 1 spare).

Code:
fileserver1 image_root # fdisk -l /dev/sde

Disk /dev/sde: 1000.2 GB, 1000204886016 bytes, 1953525168 sectors
Units = sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 512 bytes
I/O size (minimum/optimal): 512 bytes / 512 bytes
Disk identifier: 0x00000000

   Device Boot      Start         End      Blocks   Id  System
/dev/sde1            2048     2099199     1048576   fd  Linux raid autodetect
/dev/sde2         2099200     4755239     1328020   82  Linux swap / Solaris
/dev/sde3         4755240    29945159    12594960   fd  Linux raid autodetect
/dev/sde4        29945160  1953520064   961787452+  fd  Linux raid autodetect
fileserver1 image_root # badblocks -wsv /dev/sde
Checking for bad blocks in read-write mode
From block 0 to 976762583
Testing with pattern 0xaa: done
Reading and comparing: done
Testing with pattern 0x55: ^[[Adone
Reading and comparing: done
Testing with pattern 0xff: done
Reading and comparing: done
Testing with pattern 0x00: done
Reading and comparing: done
Pass completed, 0 bad blocks found. (0/0/0 errors)
fileserver1 image_root # smartctl --all /dev/sde
smartctl 6.0 2012-10-10 r3643 [x86_64-linux-3.6.6-gentoo-fileserver1-raid1] (local build)
Copyright (C) 2002-12, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Western Digital Caviar Black
Device Model:     WDC WD1001FALS-00J7B0
Serial Number:    WD-WMATV6692998
LU WWN Device Id: 5 0014ee 600096ffb
Firmware Version: 05.00K05
User Capacity:    1,000,204,886,016 bytes [1.00 TB]
Sector Size:      512 bytes logical/physical
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA8-ACS (minor revision not indicated)
SATA Version is:  SATA 2.5, 3.0 Gb/s
Local Time is:    Wed Nov 21 11:12:50 2012 EST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x84) Offline data collection activity
                                        was suspended by an interrupting command from host.
                                        Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (18600) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        ( 214) minutes.
Conveyance self-test routine
recommended polling time:        (   5) minutes.
SCT capabilities:              (0x3037) SCT Status supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x002f   200   200   051    Pre-fail  Always       -       0
  3 Spin_Up_Time            0x0027   249   231   021    Pre-fail  Always       -       7508
  4 Start_Stop_Count        0x0032   100   100   000    Old_age   Always       -       118
  5 Reallocated_Sector_Ct   0x0033   192   192   140    Pre-fail  Always       -       61
  7 Seek_Error_Rate         0x002e   200   194   000    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   073   073   000    Old_age   Always       -       19743
 10 Spin_Retry_Count        0x0032   100   100   000    Old_age   Always       -       0
 11 Calibration_Retry_Count 0x0032   100   100   000    Old_age   Always       -       0
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       116
192 Power-Off_Retract_Count 0x0032   200   200   000    Old_age   Always       -       115
193 Load_Cycle_Count        0x0032   200   200   000    Old_age   Always       -       118
194 Temperature_Celsius     0x0022   110   096   000    Old_age   Always       -       40
196 Reallocated_Event_Count 0x0032   199   199   000    Old_age   Always       -       1
197 Current_Pending_Sector  0x0032   200   200   000    Old_age   Always       -       0
198 Offline_Uncorrectable   0x0030   200   200   000    Old_age   Offline      -       0
199 UDMA_CRC_Error_Count    0x0032   200   200   000    Old_age   Always       -       0
200 Multi_Zone_Error_Rate   0x0008   200   200   000    Old_age   Offline      -       0

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
No self-tests have been logged.  [To run self-tests, use: smartctl -t]


SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

fileserver1 image_root # smartctl --test long /dev/sde
smartctl 6.0 2012-10-10 r3643 [x86_64-linux-3.6.6-gentoo-fileserver1-raid1] (local build)
Copyright (C) 2002-12, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF OFFLINE IMMEDIATE AND SELF-TEST SECTION ===
Sending command: "Execute SMART Extended self-test routine immediately in off-line mode".
Drive command "Execute SMART Extended self-test routine immediately in off-line mode" successful.
Testing has begun.
Please wait 214 minutes for test to complete.
Test will complete after Wed Nov 21 14:47:10 2012

Use smartctl -X to abort test.
fileserver1 image_root #

I am sorry I do not have the before SMART that showed 7 current pending sectors.
I am going to call this one a localized media defect if it passes the SMART long test followed by a little more testing which will probably include executing the command to tell the raid to replace a drive with this drive without degrading the array (this ability is a recent addition to mdadm that I really like).\


Edit: BTW, "No Errors Logged" is interesting to me. This is the first time I have seen a drive have a failure like this and not log the error. Remember I have had 75+ RMAs and many more out of warranty failures since 1997 when I started my paid career. Again not 10s of thousands so statistically not very significant..
 
Last edited:
Back
Top