• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Crushing resilver times

packetboy

Limp Gawd
Joined
Aug 2, 2009
Messages
288
Just lost another drive on this array...same thing happened about two weeks ago and resilver time was about 50 hours...not it's 250 hours! WTF.

Anyone have any ideas?

Code:
-bash-3.00$ /sbin/zpool status zulu02
  pool: zulu02
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
 scrub: resilver in progress for 17h26m, 6.78% done, 239h56m to go
config:

        NAME                         STATE     READ WRITE CKSUM
        zulu02                       DEGRADED     0     0     0
          raidz2                     DEGRADED     0     0     0
            c0t5000CCA369C40CEBd0    ONLINE       0     0     0
            c0t5000CCA221DFD971d0    ONLINE       0     0     0
            c0t5000CCA221D332FCd0    ONLINE       0     0     0
            c0t5000CCA221D420A3d0    ONLINE       0     0     0
            c0t5000CCA369C447B9d0    ONLINE       0     0     0
            c0t5000CCA221D87B16d0    ONLINE       0     0     0
            c0t5000CCA369C0D329d0    ONLINE       0     0     0
            c0t5000CCA221D4212Ed0    ONLINE       0     0     0
            replacing                DEGRADED     0     0 23.7M
              c0t5000CCA221D33315d0  FAULTED      5     9     0  too many errors
              c0t5000CCA221D87517d0  ONLINE       0     0     0  117G resilvered
            c0t5000CCA221D420FBd0    ONLINE       0     0     0
          raidz2                     ONLINE       0     0     0
            c0t5000CCA221DFD300d0    ONLINE       0     0     0
            c0t5000CCA221DFDA97d0    ONLINE       0     0     0
            c0t5000CCA221DA9542d0    ONLINE       0     0     0
            c0t5000CCA221DABAA2d0    ONLINE       0     0     0
            c0t5000CCA221DAD356d0    ONLINE       0     0     0
            c0t5000CCA221DF7FBBd0    ONLINE       0     0     0
            c0t5000CCA221DFA2EBd0    ONLINE       0     0     0
            c0t5000CCA221DFD311d0    ONLINE       0     0     0
            c0t5000CCA221DFD398d0    ONLINE       0     0     0
            c0t5000CCA221DFDA8Fd0    ONLINE       0     0     0
        cache
          c0t5000000009990015d0      OFFLINE      0     0     0
 
Dicky backplane?
Slowing the whole process down?

Or a couple of dicky ports on the backplane.

Only way to tell would be test each drive / port individually... eg a DD write test to a drive (eg make a vdev of a single drive).... then the next drive on and on.... see if one is significantly slower than the rest.

if so, swap the drive to another port and retest.... if the drive picks up speed and or throws less errors you know it's a dicky port... on the backplane. (and or cable to the backplane / expander etc)

Then RMA the backplane if it's at fault.

.
 
Depends...

ZFS being so sensitive to errors would be the best tool to "catch" any errors... As you have found.;)
COW is good for catching problems like this. (even slight ones like a fault cable etc or some form of interference along the cable route.)

You can double check the simple things 1st I guess.

Cables seated properly (nothing has vibrated slightly loose etc)
Same goes for drives in their caddies.... properly screwed into caddie....seated and latched in the bay solidly... and stick a finger on each and every caddie to check for excessive vibrations etc.
No unsheilded SAS cables running near power feeds etc.

Things like that are fun to track down:(

.
 
Total estimate time has finally started to come down, but resilver time still measured in DAYS:


Code:
  pool: zulu02
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
 scrub: resilver in progress for 39h56m, 36.86% done, 68h25m to go
config:

        NAME                         STATE     READ WRITE CKSUM
        zulu02                       DEGRADED     0     0     0
          raidz2                     DEGRADED     0     0     0
            c0t5000CCA369C40CEBd0    ONLINE       0     0     0
            c0t5000CCA221DFD971d0    ONLINE       0     0     0
            c0t5000CCA221D332FCd0    ONLINE       0     0     0
            c0t5000CCA221D420A3d0    ONLINE       0     0     0
            c0t5000CCA369C447B9d0    ONLINE       0     0     0
            c0t5000CCA221D87B16d0    ONLINE       0     0     0
            c0t5000CCA369C0D329d0    ONLINE       0     0     0
            c0t5000CCA221D4212Ed0    ONLINE       0     0     0
            replacing                DEGRADED     0     0 58.0M
              c0t5000CCA221D33315d0  FAULTED      5     9     0  too many errors
              c0t5000CCA221D87517d0  ONLINE       0     0     0  588G resilvered
            c0t5000CCA221D420FBd0    ONLINE       0     0     0
          raidz2                     ONLINE       0     0     0
            c0t5000CCA221DFD300d0    ONLINE       0     0     0
            c0t5000CCA221DFDA97d0    ONLINE       0     0     0
            c0t5000CCA221DA9542d0    ONLINE       0     0     0
            c0t5000CCA221DABAA2d0    ONLINE       0     0     0
            c0t5000CCA221DAD356d0    ONLINE       0     0     0
            c0t5000CCA221DF7FBBd0    ONLINE       0     0     0
            c0t5000CCA221DFA2EBd0    ONLINE       0     0     0
            c0t5000CCA221DFD311d0    ONLINE       0     0     0
            c0t5000CCA221DFD398d0    ONLINE       0     0     0
            c0t5000CCA221DFDA8Fd0    ONLINE       0     0     0
        cache
          c0t5000000009990015d0      OFFLINE      0     0     0

What have other people seen with resilver times for 2TB drives?
...granted...these are nearly FULL 2TB drives:

Code:
# df -h | grep zulu02
zulu02                  28T   820G   1.1T    43%    /zulu02
 
What affects resilver time, is IOPS. Resilver is similar to random read/write patterns. The more IOPS, the faster resilver. Mirrors are fastest, becuase they give highest IOPS. Raidz2 is not that fast. But still, it should not take days for an idle zpool. Something is weird here.

Of course, if you use the zpool heavily, while resilvering, then it will take longer time.
 
compression?
dedupe?
lots of snapshots?

Edit*

Actually just looking, and your time to complete seems to have dropped considerably?

was
17h26m, 6.78% done, 239h56m to go
now
39h56m, 36.86% done, 68h25m to go

so something has picked up speed.

Doubled the time running.... but estimation of finish time has dropped down 75%

.
 
Yeah, it definitely seems related to doing lots of small I/Os. I am replacing two 600GB sata drives with 1TB sas drives in a mirrored vdev. Did one side yesterday and doing the other now. Check this out:

pool: tank
state: ONLINE
status: One or more devices is currently being resilvered. The pool will
continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
scan: resilver in progress since Tue Jul 10 08:40:45 2012
44.6G scanned out of 441G at 8.93M/s, 12h37m to go
14.0G resilvered, 10.11% done
config:

NAME STATE READ WRITE CKSUM
tank ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
c2t5000C50041AB0C47d0 ONLINE 0 0 0
c2t5000C50041BD3E87d0 ONLINE 0 0 0
mirror-1 ONLINE 0 0 0
c2t50014EE206CED4ECd0 ONLINE 0 0 0
c2t50014EE25C240034d0 ONLINE 0 0 0
mirror-2 ONLINE 0 0 0
c2t5000C50041BD703Fd0 ONLINE 0 0 0
replacing-1 ONLINE 0 0 0
c2t50014EE2AEDF7498d0 ONLINE 0 0 0
c2t5000C50041BC61B7d0 ONLINE 0 0 0 (resilvering)

9MB/sec seems awfully slow, but looking at iostat:

r/s w/s kr/s kw/s wait actv wsvc_t asvc_t %w %b device
117.1 11.4 1854.5 65.3 0.0 0.1 0.0 0.8 0 3 c2t50014EE2AEDF7498d0
108.5 11.1 1834.3 65.3 0.0 0.1 0.0 0.7 0 3 c2t5000C50041BD703Fd0
0.0 66.3 0.0 3773.8 0.0 8.8 0.0 132.9 0 99 c2t5000C50041BC61B7

You can see it is reading from both sides of the mirror to get the data to write to the new drive. What is surprising is that the two read drives are showing virtually no busy%. I'm betting there is just some nuance to the iostat output I don't yet grok. This is definitely an area zfs could improve...
So the target
 
I replaced my 2TB Hitachi 7K3000 with about 400G data on disk, takes about 3 hours.
 
I've got a couple of fairly large ZFS boxes that I use for backups - 70x 500gb 2.5" hitachi drives (laptop, 5400rpm) in 7x 10-disk RaidZ2's. One of them has no ZIL, it loses about 2 drives/month or so. Sometimes it takes <48hrs to resilver 2 drives, sometimes it takes 250+ hours to resilver 1. What I've noticed is two fold -

This one is more obvious; If you have a lot of IO on the array, resilvering sucks. It seems that even a modest load on the server tanks the resilver pretty bad.

The less obvious factor I've found is that if any other drives are acting up AT ALL the resilver takes forever. So far I've never seen any errors show up in a zpool status unless it was a hard-failure or pulled drive (IE the errors aren't making it to ZFS), but if the drive is throwing any errors that are being corrected before it gets to ZFS it massively impacts the resilver. If you use Napp-It, you can see disk-errors on the disk menu. When the drives are happy and I/O is light, resilvers are fast.

An unrelated note - the system with a ZIL sees 100% less failure of drives (so far) - these drives are piss-poor (they were free to me, new, from a different department), but the ZIL smooths out the IO massively making them last much longer in the same role.
 
I made the tuning changes in that article and my resilver speed jumped by about 50%.
 
I've got a couple of fairly large ZFS boxes that I use for backups - 70x 500gb 2.5" hitachi drives (laptop, 5400rpm) in 7x 10-disk RaidZ2's. One of them has no ZIL, it loses about 2 drives/month or so. Sometimes it takes <48hrs to resilver 2 drives, sometimes it takes 250+ hours to resilver 1. What I've noticed is two fold -

This one is more obvious; If you have a lot of IO on the array, resilvering sucks. It seems that even a modest load on the server tanks the resilver pretty bad.

The less obvious factor I've found is that if any other drives are acting up AT ALL the resilver takes forever. So far I've never seen any errors show up in a zpool status unless it was a hard-failure or pulled drive (IE the errors aren't making it to ZFS), but if the drive is throwing any errors that are being corrected before it gets to ZFS it massively impacts the resilver. If you use Napp-It, you can see disk-errors on the disk menu. When the drives are happy and I/O is light, resilvers are fast.

An unrelated note - the system with a ZIL sees 100% less failure of drives (so far) - these drives are piss-poor (they were free to me, new, from a different department), but the ZIL smooths out the IO massively making them last much longer in the same role.

Until I read the end of your message I was wondering why you would have such a strange build, 70 ports means a lot of money, for only 500GB slow drives, and unreliable at that. But I guess if they're free that's better, although you must have a healthy stack of spares with that failure rate !
 
Until I read the end of your message I was wondering why you would have such a strange build, 70 ports means a lot of money, for only 500GB slow drives, and unreliable at that. But I guess if they're free that's better, although you must have a healthy stack of spares with that failure rate !

Each array will store around 36TB with the compression rates we see on our backups, with a total server cost around $6,500/ea. They will smoke most of our other SAN/NAS servers (especially the one with a ZIL) for most workloads, except when drives start to fail.

I've still got a few hundred of the drives in boxes, so I'm good on spares for a while.
 
So the homemade ZFS solution is doing as good, or better, than your other SAN/NAS servers? Cool! Your boss must be happy, saving money.
 
It's neat, and it would cost a fortune for similar sized servers from the big players. We still need a new primary VMware SAN though. We had a mid-range (IE expensive) "fully redundant" SAN in place that was a royal piece of shit - now that I've proven ZFS, we're evaluating and HA Nexenta or TrueNAS for that spot.

I'd love to go with TrueNAS as I like the guys at iX and their price is far better, but they're mid-VMware HCL approval.
 
iX is a good group of guys as are the BSD folks in general, however, i wouldn't trust their HA at this point.

Nexenta's HA is solid and well tested. truthfully it isn't even their HA they use rsf-1 from highavailability.com.

Something to consider nexenta is a few months from releasing 4.0 which should have all the support for the sandy bridge E5 xeons (currently 3.1.3 has limited support ie no c606 sas controller support) as well as the LSI 9207. i know for a fact both of those are currently deployed in production environments and work but nexenta doesn't 'officially' support them yet.

the reason i bring that point up is 4.0 and the new xeons plus pci-e 3.0 offer a MASSIVE performance potential above and beyond 1366 based xeons (which are very capable). i rolled out a monster nexenta HA setup two weeks ago because there was a requirement that couldn't wait. given the option to wait though i would have waited for the official support instead of having to plan for a rather stressful change window to swap out motherboards, controllers, and CPUs later in the year. technically i should be able to do the upgrade with no down time but it is still going to be a rather stressful change.

that all said though in terms of raw performance without obscene costs associated, emc, netapp, dell ... none of them can touch what you can do with nexenta. DDN is really the only hardware vendor that has proprietary hardware that is faster than 'off the shelf' components but they're extremely expensive.
 
Eventually..it finished:

Code:
  pool: zulu02
 state: ONLINE
 scrub: resilver completed after 92h4m with 0 errors on Thu Jul 12 10:40:33 2012
config:

        NAME                       STATE     READ WRITE CKSUM
        zulu02                     ONLINE       0     0     0
          raidz2                   ONLINE       0     0     0
            c0t5000CCA369C40CEBd0  ONLINE       0     0     0
            c0t5000CCA221DFD971d0  ONLINE       0     0     0
            c0t5000CCA221D332FCd0  ONLINE       0     0     0
            c0t5000CCA221D420A3d0  ONLINE       0     0     0
            c0t5000CCA369C447B9d0  ONLINE       0     0     0
            c0t5000CCA221D87B16d0  ONLINE       0     0     0
            c0t5000CCA369C0D329d0  ONLINE       0     0     0
            c0t5000CCA221D4212Ed0  ONLINE       0     0     0
            c0t5000CCA221D87517d0  ONLINE       0     0     0  1.69T resilvered
            c0t5000CCA221D420FBd0  ONLINE       0     0     0
          raidz2                   ONLINE       0     0     0
            c0t5000CCA221DFD300d0  ONLINE       0     0     0
            c0t5000CCA221DFDA97d0  ONLINE       0     0     0
            c0t5000CCA221DA9542d0  ONLINE       0     0     0
            c0t5000CCA221DABAA2d0  ONLINE       0     0     0
            c0t5000CCA221DAD356d0  ONLINE       0     0     0
            c0t5000CCA221DF7FBBd0  ONLINE       0     0     0
            c0t5000CCA221DFA2EBd0  ONLINE       0     0     0
            c0t5000CCA221DFD311d0  ONLINE       0     0     0
            c0t5000CCA221DFD398d0  ONLINE       0     0     0
            c0t5000CCA221DFDA8Fd0  ONLINE       0     0     0
        cache
          c0t5000000009990015d0    OFFLINE      0     0     0

While the first one was recovering, lost another drive on a different server...that server is nearly identical to the other server, yet reslver was "just" one day:

Code:
  pool: zulu03
 state: ONLINE
 scrub: resilver completed after 25h27m with 0 errors on Thu Jul 12 01:39:29 2012
config:

        NAME                       STATE     READ WRITE CKSUM
        zulu03                     ONLINE       0     0     0
          raidz2-0                 ONLINE       0     0     0
            c0t5000CCA221D87C67d0  ONLINE       0     0     0
            c0t5000CCA221D3121Cd0  ONLINE       0     0     0  1.78T resilvered
            c0t5000CCA221D8681Cd0  ONLINE       0     0     0
            c0t5000CCA221D42054d0  ONLINE       0     0     0
            c0t5000CCA221DFF08Cd0  ONLINE       0     0     0
            c0t5000CCA221DFF27Cd0  ONLINE       0     0     0
            c0t5000CCA221E7A014d0  ONLINE       0     0     0
            c0t5000CCA221E68A4Bd0  ONLINE       0     0     0
            c0t5000CCA221E76C0Dd0  ONLINE       0     0     0
            c0t5000CCA221F15BDAd0  ONLINE       0     0     0
          raidz2-1                 ONLINE       0     0     0
            c0t5000CCA228C03202d0  ONLINE       0     0     0
            c0t5000CCA228C0A32Ad0  ONLINE       0     0     0
            c0t5000CCA228C01CC0d0  ONLINE       0     0     0
            c0t5000CCA228C01EBBd0  ONLINE       0     0     0
            c0t5000CCA228C056DAd0  ONLINE       0     0     0
            c0t5000CCA228C095B5d0  ONLINE       0     0     0
            c0t5000CCA228C095C9d0  ONLINE       0     0     0
            c0t5000CCA228C096ABd0  ONLINE       0     0     0
            c0t5000CCA228C09518d0  ONLINE       0     0     0
            c0t5000CCA228C0967Fd0  ONLINE       0     0     0
 
Back
Top