• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Will I benefit from ZFS?

mda

2[H]4U
Joined
Mar 23, 2011
Messages
2,212
Hello all,

I'm running a smallish (by [H] standards) ext3-based Linux RAID machine as our company file and mail server.

Specs:
Intel i3 3xxx
Asus B75
4GB DDR3 1333
6x WD Black 2TB arranged in 3 volumes of RAID1 (to provide hard limits on users)
CENTOS 6.4 x64

I'm currently in the process of setting up a machine to backup the contents of this file server with the ff specs:
Intel i3 4xxx
Asus H87
8GB DDR 1600
<no HDDs yet, but I'm thinking 6x WD Red 2TB also in 3 volumes of RAID1)
Thinking of running CENTOS 6.4 x64 OR CENTOS 7

I was reading up about ZFS and thought this would be nice as it prevents silent data corruption.

My concern is that if I only run ZFS on the backup and don't migrate the main fileserver to ZFS, this will not really help especially in cases where files on the main server gets corrupted, and having a perfect copy of this corrupt file on the backup server won't help.

My 2/3 questions at this point is:
1. Will I benefit from ZFS? Is ZFS really worth migrating the old server to? (My default answer will be yes)
2. How do I go about migrating the old fileserver while minimizing downtime?
-- I considered making the new machine as the main fileserver/mail server as this computer has superior specs, but upon checking, the old system is hardly being pushed by its current functions anyway.

** Note: I have very very limited experience with Linux and my command line knowledge sucks. =/

Thanks!
 
Last edited:
If you were really afraid of silent data corruption you should have got ECC RAM for both of your machines. Of course you can run ZFS on non-ECC RAM and it will be working fine, but ZFS was developed for server hardware and servers always use ECC. In case of a bad memory cell ZFS can corrupt your data significantly more than any other filesystem. By the way, a normal RAID can protect you quite well against harddisk read errors. And mdadm definitely works reliable, I have seen it correcting quite a few unreadable sectors over the year.

If you want to migrate to ZFS do it on both machines. That way you can make proper use of ZFS send/receive. Since the new machine has better hardware (and will probably be more power-efficient) you can use this as the new primary server, transfer all data over (it is only 6 TB) and then make your old server the backup machine. You should think about using RAIDZ2 instead of three mirrors, that way you can have a single pool and gain one disk. Space limits (quotas) can be enforced on a dataset basis by ZFS. I would use rsync to transfer the data over, but if you need that last bit of speed a tar-mbuffer-untar chain will also work. If you have some limited experience with Linux you can stay with that. I run ZFS on Ubuntu 14.04.

Is it worth it? I would say yes, for the snapshots alone. For me that was the main reason to migrate my data along with send/receive. The data protection was just a bonus. But you should be aware of the fact that ZFS pools are going to get fragmented with no means to defragment them at the moment short of transfering all data to a new pool. This is uncritical for a media library, but a pool that is constantly randomly written to with multiple writers (not the typical home scenario, maybe except bittorrent clients) will slow down significantly over time. I'm currently in the process of transfering my main "scratch" pool with VM images and databases to SSDs because of that.

Why buy only 2 TB drives? As the current sweet spot I would get 4 TB Reds.
 
Last edited:
If you were really afraid of silent data corruption you should have got ECC RAM for both of your machines. Of course you can run ZFS on non-ECC RAM and it will be working fine, but ZFS was developed for server hardware and servers always use ECC. In case of a bad memory cell ZFS can corrupt your data significantly more than any other filesystem. By the way, a normal RAID can protect you quite well against harddisk read errors. And mdadm definitely works reliable, I have seen it correcting quite a few unreadable sectors over the year.

Thanks for this.

Very good point on the ECC memory, although it may be too late to get proper hardware for ECC at this point, as it will mean I will have to replace the board, CPU and RAM for at least one machine.

Yes, we are using the standard Linux RAID mdadm but I just inherited the setup and have little knowledge of how it works. I should dig up the config files/settings and see how I can prevent this.

If you want to migrate to ZFS do it on both machines. That way you can make proper use of ZFS send/receive. Since the new machine has better hardware (and will probably be more power-efficient) you can use this as the new primary server, transfer all data over (it is only 6 TB) and then make your old server the backup machine. You should think about using RAIDZ2 instead of three mirrors, that way you can have a single pool and gain one disk. Space limits (quotas) can be enforced on a dataset basis by ZFS. I would use rsync to transfer the data over, but if you need that last bit of speed a tar-mbuffer-untar chain will also work. If you have some limited experience with Linux you can stay with that. I run ZFS on Ubuntu 14.04.

Is it worth it? I would say yes, for the snapshots alone. For me that was the main reason to migrate my data along with send/receive. The data protection was just a bonus. But you should be aware of the fact that ZFS pools are going to get fragmented with no means to defragment them at the moment short of transfering all data to a new pool. This is uncritical for a media library, but a pool that is constantly randomly written to with multiple writers (not the typical home scenario, maybe except bittorrent clients) will slow down significantly over time. I'm currently in the process of transfering my main "scratch" pool with VM images and databases to SSDs because of that.

Why buy only 2 TB drives? As the current sweet spot I would get 4 TB Reds.

I'd like the ability to do snapshots but I don't think have the spare space to do that at this time.
**I have about 2.5TB free at this point in the existing 6TB setup so I don't have room to create one snapshot of the whole thing.

Am reading about ext3 now as it appears there is also another option of ext4 as well.

My main reasons of getting 2TB drives are due to local costs of the hardware (cost / GB cheaper on the 2TB vs 3/4TB right now for me), and that all our clone servers (all standard desktop-level machines) use 2TB drives in RAID1 - we will have to keep less HDD spares and less possible complications.

Also, WRT to RAIDZ2, I am concerned about drives dying during rebuilding. Is this a big deal? This was one of our concerns when we moved our Oracle machine from a 4 Disk RAID5 to an 8 Disk RAID10 in our latest server upgrade.

Once again, thanks for the reply. Very helpful =)
 
Ext4 is an evolution of the ext3 filesystem. It is the de facto standard filesystem in Linux today and ext3 can be converted to ext4. It is considered very stable and fast, but there are no significant new features.

A snapshot consumes no additional space when you create it and only grows by the amount of changed or deleted data. So in your case you would have to change 2.5 TB of data (remember, new data would consume space anyway) before the disk is full. And if you have some data that frequently changes you can store it on a separate dataset and leave it out from snapshots.

With RAIDZ2 you can lose two disks before the pool fails, with mirrors you can only lose one. RAID5/Z1 is not recommended anyway for disks of that size. There is a chance that the other disk dies during the rebuild of mirror, too! I would only recommend mirrors if you need more IOPS than a single vdev can provide or if you want to add disks on the fly.
 
Thanks. Will do a little more research.

Leaning towards ZFS again now, but wary of the migration that needs to be done.

Having the mail server go down for a while isn't something I look forward to :D
 
If your server is so important that a downtime is a problem,
you must think of a new machine where you can test the setup and sync the data.
Switch over can be done then within minutes.

If your server is that important you have many reasons to switch to zfs
- always data you can trust (checksums)
- always consistent filesystem (copy on write, no fschk needed)
- snapshots without delay and initial space consumption
- intelligent arc read cache that can be extended with SSDs
- dedicated ZIL devices for fast secure writes
- storage pools that can grow without limit on the fly
- async replication where open files are not a problem
- no write hole problems on raid like with hardware raid 1/5/6
 
If data corruption is the only thing you care about check out btrfs as
a) It's mainlined in the linux kernel
b) Takes care of silent data corruption
c) Has an inplace upgrade from ext3/4
That said, it's slower (currently) than ZFS and is nowhere near feature parity with ZFS.
Conversion doc https://btrfs.wiki.kernel.org/index.php/Conversion_from_Ext3

For ZFS I'd recommend one of the BSDs as it's a native filesystem and not a kernel module.
 
If you go the freebsd route, try trueos - it has boot environment support. e.g. you can have snapshots of your running system, so that upgrades/installs/etc can be backed out if you get hosed.
 
Back
Top