• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Looking for guidance with SMART testing on a server

Joined
Sep 19, 2012
Messages
24
Hello all,

I am running a 2U Supermicro dual Xeon E5620's, 60 gig OS (Ubuntu 14.04 Server) with back up drive, dual 256 gig SSD ZIL/LOG drives, and 6 WD Red 4TB drives in a RaidZ2 array. This machine has been running successfully for a little over a year now, and I am still SLOWLY making my way through setting everything up. It is successfully running as a NAS, Plex Media Server, and running a couple Minecraft servers.

I just installed smartmontools earlier today, and am setting up the /etc/smartd.conf to run tests on the drives. Looking around various places, it was recommended to have short tests run often, on the order of daily, and longer test weekly. That being said, is that a good plan for the setup I have? Also, if that is done, would be a VERY BAD idea to have the Long tests run concurrently, or better to split it up and have each of the 4TB drives run the long test in a dedicated window every morning early in the morning?

For the MOST part, the server is idling between midnight and roughly noon to 5 PM, depending on the day, so I would have any maintenance done during that time. I just complete bash scripts to back up my minecraft servers daily around 3 am, and the plex server once a week about 4 am.

I am still VERY VERY green with Ubuntu, and haven't ever done ANYTHING with SMART data before. Also, I came across zpool scrubbing, which I will look into and possibly set up as well.



Here are the drives I have in the system, leaving out two 500 gig drives that are in there that aren't in use and I don't even remember inserting... oh well. Plan is to add this all to the /etc/smartd.conf so that I can keep the drives straight...


Code:
Main OS drive, and backup drive (will be upgraded soon, hopefully)

/dev/sda
#Device Model: OCZ-AGILITY3           Serial Number: -   Firmware Version: 2.15
#Test Durations: Short: 1 min, Long: 48 min, Conveyance: 2 min

/dev/sdb
#Device Model: OCZ-AGILITY3           Serial Number: -   Firmware Version: 2.15
#Test Durations: Short: 1 min, Long: 48 min, Conveyance: 2 min

ZIL/LOG drives...

/dev/sdc
#Device Model:TOSHIBA THNSNH256GBST   Serial Number:-           Firmware Version: HTRAN101
#Test Durations: Short: 2 min, Long: 14 min

/dev/sdd
#Device Model:TOSHIBA THNSNH256GBST   Serial Number: -           Firmware Version: HTRAN101
#Test Durations: Short: 2 min, Long: 14 min


RaidZ2 pool...

/dev/sde
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 80.00A80
#Test Durations: Short: 2 min, Long: 523 min, Conveyance: 5 min

/dev/sdf
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 80.00A80
#Test Durations: Short: 2 min, Long: 520 min, Conveyance: 5 min

/dev/sdg
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 80.00A80

#Test Durations: Short: 2 min, Long: 511 min, Conveyance: 5 min
/dev/sdh
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 80.00A80
#Test Durations: Short: 2 min, Long: 539 min, Conveyance: 5 min

/dev/sdi
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 82.00A82
#Test Durations: Short: 2 min, Long: 517 min, Conveyance: 5 min

/dev/sdj
#Device Model:WDC WD40EFRX-68WT0N0    Serial Number: -        Firmware Version: 80.00A80
#Test Durations: Short: 2 min, Long: 524 min, Conveyance: 5 min
 
Are you running zfs scrubs regularly? If so, I wouldn't bother with regular smart tests, instead just monitor the smart parameters. Backblaze has an article about it, here Hard Drive SMART Stats but the tldr is look for changes in these:

  • SMART 5 – Reallocated_Sector_Count.
  • SMART 187 – Reported_Uncorrectable_Errors.
  • SMART 188 – Command_Timeout.
  • SMART 197 – Current_Pending_Sector_Count.
  • SMART 198 – Offline_Uncorrectable.
For the SSD, good luck, I've had some number of SSD failures in production, but none had any warning signs I could tell. (Well, one was acting dumb for a long time before it gave up, that one probably had warning signs, but it was before I started doing smart stats on our spinny disks; the vast majority of the SSDs that fail on me just decide not to talk on the sata bus anymore)
 
I am actually running my first ever (that I recall) zpool scrub now. Roughly 1TB per hour, 500 Meg done, 6 hours to go.

I appreciate the link you provided, I will need to look into it.

Thank you!!
 
Ok, so it is STILL running, about 30% done, with 15 hours to go. No erros found, so that is a plus.

Do these scans address fragmentation as well? Zpool was 28% fragmented last night before the scan... shame on ME.
 
zpool scrubs don't touch fragmentation, as far as I am aware. I wouldn't worry about it too much, or at least I don't. There's not a good way to solve it on the system I use though.

I also run smartmontools, and I wouldn't bother with all that SMART scanning of your system. Periodic pool scrubbing is a good idea, and my system does it once a week. I do run smartd, and you might look more into some of the reporting your system does. I run FreeBSD, and it sends out a daily email. With smartd enabled, it includes a quick check of the disks in that email. If any SMART issues occur, it sends an email out right away. For example, I use external USB drives for backup currently. I've had a couple of the 3.5" Seagates overheat because they have no ventilation. smartd sends an email when it gets an overheat warning, or if a drive dies, drops out, etc. For that, it is a very handy feature set.

I'm not sure where the recommendation comes from for all that SMART testing. I've never seen it before, other than mention of smartd and ZFS systems, basically. I think someone thought "I can write a script and run this all the time!!!" and started spreading the recommendations. How often do you proactively do a SMART test on any other system? Not often by my experience, unless there's a perceived/real problem, or you're rebuilding a system or something.

The ZFS scrub, however, is IMO important to do periodically.
 
I sure hope this isnt idicative of a problem, but it is only at 30% scanned, 48 hours to go, lol. No errors still.... could it be that I am at 68% of array in use? It was at like 28.95% done earlier...
 
With default settings, the scrub runs with very low priority. If you're using your array, because it's day time and you're awake, that's going to slow down the scrub. The scrub reads all of your files, so the more space you've used, the longer it takes. A smart surface scan is also going to run at low priority, and take forever.
 
Back
Top