• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Seagate Green 2TB ST2000DL003 sample variation

MetaGenie

Limp Gawd
Joined
Nov 6, 2009
Messages
246
I just bought five ST2000DL003 with the intention of using them in a RAID array. Because I'm using an Areca ARC-1261ML which can do sustained sequential transfers up to 840 MB/s (maybe even faster — I haven't yet had the opportunity to max it out), the slowest drive among these is going to govern how fast the entire array can be (in sustained transfers). So how fast each one is important to me.

I ran a full test in HD Tune Pro 4.61 (full version) on each drive (the full test took just under 5 hours on each drive, but generates a very nice neat graph). The results were very interesting — and then when I viewed all five results at the end, the results were stunning. It turns out that sorting the drives by average transfer rate is identical to sorting them by serial number!

I think it would be unwise to share the full serial numbers, but the first one is 5YD31xxx and the last is 5YD36xxx. I was so stunned by the result that I double and triple and quadruple checked; they are definitely in order by their average sustained sequential transfer rate. Note that since Seagate ST2000DL003 serial numbers are alphanumeric, the sorting is done with 0-9 first followed by A-Z; the slowest drive is the first in the sort with the earliest (lowest) serial number, and the fastest drive has the latest (highest) serial number. The chance of this happening randomly should be 1/5! = 1/120... pretty slim odds of this happening by chance.

There are some minor sources of error, for when I was multitasking and reading some files on other hard drive volumes it interfered a little bit with the test (slowed down transfer rate a little in a particular point of the graph &#8212; notably towards the end of the test of the first drive and third drive) but this should affect the average by <0.1 MB/s.

All five drives are divided into 17 transfer rate zones, and the dividing points between zones are the same on all five drives.

I have ordered seven more of this same model drive from a different place. (The plan is to put 12 of them in a RAID 6.) When they arrive it'll be most interesting to see if they fit into the pattern. At the moment I'm not entirely convinced this didn't happen by random chance, but if all 12 drives fit into the pattern there will be no disputing it.

BTW sample variation is not specific to Seagates... I've seen it in Western Digitals as well; among the same model and firmware, the zoning pattern itself is subject to variation, as well as the average transfer rate.

Here are all five, sorted by serial number (from lowest to highest, top to bottom):
rXnMsQQ.png

(ignore the CPU usage &#8212; that seems to vary based on if I'm multitasking at the time the test finishes)
 
Last edited:
I just bought five ST2000DL003 with the intention of using them in a RAID array.
... the slowest drive among these is going to govern how fast the entire array can be (in sustained transfers). So how fast each one is important to me.
[Yes, in a striped array, a data transfer isn't over till the *last* fat-lady sings.]
I ran a full test in HD Tune Pro 4.61 (full version) ... The results were very interesting — ...
Thanks for posting them. However, I would not try to find any correlation to serial #. A variation (between "identical" drives) in transfer rate performance is unavoidable, but I would prefer not to see it as large as the ~8% you measured.

Because this variation is due to a single mechanical component, it will manifest itself uniformly across the entire drive. Hence, the maximum transfer rate measurement is a better indicator (than average), since average will be affected by various other factors (mapped defects, system "distractions", etc.). As a result, your drives #3 & #4 should be swapped in the order.

A good question might be, "Why do all but your #5 underperform the stated specification for "Sustained data transfer rate OD 144 MB/s"? [Note: Seagate's "MB" is almost certainly 10^6; so the 144 is really 138 MB/s in HDTune's 2^20 terms]
[Ref: Page 11 of http://www.seagate.com/staticfiles/support/disc/manuals/desktop/Barracuda Green/100649225b.pdf ]

I appreciate that Seagate does publish such specifications (try to find comparable specs for Samsung HDs), but they should be more reliable than your experience indicates.

I have ordered seven more of this same model drive from a different place. ...

Please do follow-up with your test results.

-- UhClem
 
I would prefer not to see it as large as the ~8% you measured.
I appreciate that Seagate does publish such specifications (try to find comparable specs for Samsung HDs), but they should be more reliable than your experience indicates.
I really wouldn't be surprised if all manufacturers' green drives were like this. I've seen more variation among Western Digital WD15EADS, and among WD20EADS, than among these Seagates.

As you said, I would like there to be less variation. It seems to me that the faster drives might get stressed more by having to pause while the slower drives finish, because this would require more seeking. But I'm used to mass-produced things having lots of individual variation. Camera lenses and binoculars for example, have tons of individual variation in optical quality in my experience (among identical models).

Because this variation is due to a single mechanical component, it will manifest itself uniformly across the entire drive.
Please make it clear whether this is speculation. My hypothesis is that RPM variations are the root cause of these transfer rate variations, which I'm guessing yours is too, but I wouldn't call it a foregone conclusion without further data from other types of measurements.

Hence, the maximum transfer rate measurement is a better indicator (than average), since average will be affected by various other factors (mapped defects, system "distractions", etc.). As a result, your drives #3 & #4 should be swapped in the order.
I strongly disagree on this point. Just cut out the #3 and #4 images and flip between them, and you'll see that the very first zone is just an exception &#8212; in all the remaining 16 zones, #4 is faster than or equal in speed to #3. My hypothesis here is that the adaptive mapping of drive #4 (done during manufacturing) decreased the average data density within the first zone because the underlying medium couldn't handle the maximum data density very well.
Qmnth.gif


Also note that within each zone, the transfer rate fluctuates quite a bit. HDTune only takes 200 samples, so these are highly averaged and don't show the fluctuations.

So it certainly wouldn't make sense to rank them by their maximum zone-averaged transfer rate, unless I'm using them short-stroked.

A good question might be, "Why do all but your #5 underperform the stated specification for "Sustained data transfer rate OD 144 MB/s"? [Note: Seagate's "MB" is almost certainly 10^6; so the 144 is really 138 MB/s in HDTune's 2^20 terms]
[Ref: Page 11 of http://www.seagate.com/staticfiles/support/disc/manuals/desktop/Barracuda Green/100649225b.pdf ]
I wouldn't jump to the conclusion that Seagate is using 10^6 in their MB/s. On a finer-grained scale, the fastest drive actually exceeds 144 MiB/s. Observe what happens when HDTune's is short-stroked on drive #5 (this is still a "full" test, but just on the first 10 GB):
1QWGs.png


The pattern here suggests to me that adaptive density mapping is being done, and it's only on average that the zones cleanly stair-step downwards. Different platter surfaces, and different local areas on the same platter, probably have varying abilities in their maximum reliable data density. The periodicity could come from both the platter rotation and the interleaving of platter surfaces (offhand I'm not sure which effect is stronger here... it would require estimating the size of a track in MB).

On an even finer-grained scale (short-stroked to 1GB) it even reaches 149.5 MiB/s:
Fdijj.png


I intend to do a full short-stroked comparison, but right now I don't have all the drives plugged in so it'll have to wait.


P.S. Welcome to HardForum!
 
Last edited:
It seems to me that the faster drives might get stressed more by having to pause while the slower drives finish, because this would require more seeking.
Worry not. There would be no excess seeks, just an occasional un[der]-productive revolution.
Please make it clear whether this is speculation. My hypothesis is that RPM variations are the root cause of these transfer rate variations, which I'm guessing yours is too, but I wouldn't call it a foregone conclusion without further data from other types of measurements.
Of course it's speculation [I don't write the HD firmware for Seagate:)] Long ago, I did write a couple of disk drivers for Unix, but that was well before the word "firmware" existed. So, my speculation is guided by a pretty thorough (albeit dated) understanding of disk drive basics. I believe the variations are due to .... [short answer:] cylinder skew.
Also note that within each zone, the transfer rate fluctuates quite a bit. HDTune only takes 200 samples, so these are highly averaged and don't show the fluctuations.
HDTune's plots offer a concise representation of "performance", but I prefer "dd" [*nix] for zooming in, and seeing/measuring exactly what I want. [Note that, on Windows, the Cygwin environment (with "dd" etc.) is very useful.]


The pattern here suggests to me that adaptive density mapping is being done, and it's only on average that the zones cleanly stair-step downwards. Different platter surfaces, and different local areas on the same platter, probably have varying abilities in their maximum reliable data density.
Take care not to over-analyze things. While such variability (in surface magnetics) is real, adapting to it (even exploiting it) is impractical (they're not cutting 50 carat diamonds here).
The periodicity could come from both the platter rotation and the interleaving of platter surfaces (offhand I'm not sure which effect is stronger here... it would require estimating the size of a track in MB).
"Interleaving of platter surfaces" ?? .... Huh!!
On an even finer-grained scale (short-stroked to 1GB) it even reaches 149.5 MiB/s:
Yup, but that's why the spec states "Sustained".
I intend to do a full short-stroked comparison, but right now I don't have all the drives plugged in so it'll have to wait.
I think that a 200MB focus on drive #1 and #5 will support my "hypothesis". Do several runs of each, and discard the "outliers". Thanks.

-- UhClem
 
Long ago, I did write a couple of disk drivers for Unix, but that was well before the word "firmware" existed. So, my speculation is guided by a pretty thorough (albeit dated) understanding of disk drive basics.
Cool, nice to know I'm speaking to a fellow programmer. Do you mean Linux, or some other flavor of Unix?

As for me, my research into hard drive internals began when an IBM "Deathstar" I bought in 2001 developed a huge batch of bad sectors 1 year later. I wrote various small programs to assess and summarize the damage, recover as much as I could, and analyze the failure in an attempt to understand it. The project was very successful. I had no way of reverse-engineering the firmware, but I "clean-room" reverse-engineered the drive's ECC and RLL algorithms (which allowed me to recover many of the bad sectors in the files most important to me) and its list of Primary (P-list) defects by timing the sector-to-sector seeks across the whole drive (which allowed me to make a physical map of the grown defects).

So I think I have a good understanding of how drives worked back in 2000, but a lot has changed since then. I'm pretty sure drives are much more adaptive now (i.e. adapted to quirks in their recording surfaces during manufacturing) and more complex in other ways.

I believe the variations are due to .... [short answer:] cylinder skew.
Ah yes, I had to compensate for cylinder skew and head skew when converting LBA to physical CHS in my "Deathstar" project. Figuring this out was very satifying and made the graph look much better. (I also learned that on that drive, every 512 tracks, one track was reserved for grown defects; and the last step in getting the graph perfect was learning that radially adjacent sectors are reserved for embedded servo.)

I don't see how cylinder skew could explain instance-to-instance sustained transfer rate variations in drives of the same model... all it does is make it so that by the time the drive has seeked to the next cylinder, it is positioned at sector 0 (well actually, a little bit before it, to allow for variances).

Please be more specific?

HDTune's plots offer a concise representation of "performance", but I prefer "dd" [*nix] for zooming in, and seeing/measuring exactly what I want. [Note that, on Windows, the Cygwin environment (with "dd" etc.) is very useful.]
I did not know that "dd" could generate a fine-grained report of transfer rate. Unless... you wrote a shell script calling "dd" for a string of very small transfers?

I've been meaning to write a program to do exactly the tests I want, but in the meantime I decided to go with HDTune.

Take care not to over-analyze things. While such variability (in surface magnetics) is real, adapting to it (even exploiting it) is impractical (they're not cutting 50 carat diamonds here).
I really do think that in modern drives, they have to do this to squeeze out every bit of storage space they can. If hard drives have to be given a P-list during manufacturing (which is adaptive to individual flaws on the magnetic surfaces), then why not density mapping too?

On the other hand, the P-list might explain the whole thing (and not "density mapping"). Maybe swaths of physically adjacent sectors are cut out from adjacent tracks, making a whole string of consecutive tracks slower due to containing less data. Because of platter surface interleaving, it would turn into a gradually morphing repeating pattern of transfer rates.

"Interleaving of platter surfaces" ?? .... Huh!!
Er, I just mean that the data immediately following a cylinder/head is on the same cylinder, on the next head. Once the last head is finished, it goes to the next cylinder on the first head. I'm using the term "platter surface" because this corresponds directly to a head. I can't just say "platter" because every platter has two heads (except on a drive with an odd number of heads).

Uh, well that's how drives did it back in 2000. Now they might go ahead for a certain number of cylinders on the same head before switching to the next head. I remember reading about that new development somewhere. Probably has to do with cylinder seek being quicker than head seek.

I think that a 200MB focus on drive #1 and #5 will support my "hypothesis". Do several runs of each, and discard the "outliers". Thanks.
Can't do that with silly HDTune. It only short-strokes in "gB". So... I'll get back to you on that.
 
Last edited:
Oops, looks like I forgot to follow-up on this by posting the test results of all twelve drives together. As it turned out the next seven drives didn't follow the pattern of average transfer rating sorting the same as serial number (it would have been enormously surprising if they had).

Anyway, sorry for the thread necro, but here's the full list:
9qJuvee.png
 
Back
Top