• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

XFS vs ZFS

You could say that of md-RAID and btrfs, however the RAID code in btrfs came directly from md-RAID. So...what's your point?
.

I was responding to a question as to why md-raid and lvm were not in one package. Don't see what was unclear about that, or why you think I was deserving of attitude...
 
I was responding to a question as to why md-raid and lvm were not in one package. Don't see what was unclear about that, or why you think I was deserving of attitude...

Very obviously they were developed by different people at different times, same applies to btrfs, ext4 and every other storage mechanism in the Linux kernel - that didn't stop the RAID code from md-RAID making its way into btrfs to give it RAID capability.

My point was that hitherto (and going forward), there's no credible reason why LVM and md-RAID could not be integrated into a single mechanism and toolset, especially seeing as both were around long before btrfs.


EDIT :

Admittedly, I took the tone of your post as sarcasm. :)
 
Sigh, not intended that way, I assure you. Problem on a forum is that you can often not tell how knowledgeable folks are about specific issues.
 
Huh? What is wrong with mdadm? I'd choose that over hardware RAID in many cases, since mdadm is more flexible.

This is why I like MDADM software RAID, it has great flexibility with drives and if using RAID 0 or 1, even with encryption it hardly eats any CPU cycles, even with an older processor that is single core.

RAID 5 and 6 take a bit more CPU power if encryption is used, but ZFS has all of that and error correction, making it far superior when it comes to raw storage, security, and safety of the data contained on it.
 
That is hardly a reason to switch to hardware RAID. It is a minor issue, especially since LVM is not needed in many cases.

Interesting thought. I've never had a need for mdadm since every Linux box I support has hardware RAID. Meanwhile LVM is a basic necessity on every Linux box we support.

AIX, Solaris, HP-UX, Irix, and if memory serves, Tru64/Digital Unix, Dynix/Ptx, and probably a slew of other UNIX OSes support raid within their LVMs. Linux can if you get Veritas Volume Manager, but since LVM is decent, it's not usually needed. Yet LVM doesn't have this functionality.

I can see someone that's playing around at home with Linux, but on a production server, it really should be using LVM. Much easier to grow a file system when needed, etc.
 
Hardware RAID is probably the highest performance you will get next to a high-end ZFS system, but MDADM software RAID has great flexibility and OS compatibility.
 
I can see someone that's playing around at home with Linux, but on a production server, it really should be using LVM. M

That is a sweeping statement that has many examples to the contrary. LVM is not used, and not needed, in many a production server.
 
That is a sweeping statement that has many examples to the contrary. LVM is not used, and not needed, in many a production server.

Not used, I can see. That's not to say that it's good or bad. Just that it's being done. Not needed is a more rare thing in my mind. Though maybe that just boils down to a preference thing. Coming from an AIX and Solaris background, I'm used to working in LVMs so a Linux without LVM doesn't seem like a server to me.

Still, I'm curious. When would you run a Linux box, other than home use, without LVM?
 
Not used, I can see. That's not to say that it's good or bad. Just that it's being done. Not needed is a more rare thing in my mind. Though maybe that just boils down to a preference thing. Coming from an AIX and Solaris background, I'm used to working in LVMs so a Linux without LVM doesn't seem like a server to me.

Still, I'm curious. When would you run a Linux box, other than home use, without LVM?

When would you run a Linux box without LVM? Well, you'd do this for pretty much any server workload with predictable, well known usage...as in any workload where LVM is just additional overhead rather than benefit. Also, in any environment where growth is so fast that its more effective just to build the next server than reconfigure the last one. Suggesting that "its not a server without LVM" is just silly talk.
 
When would you run a Linux box without LVM? Well, you'd do this for pretty much any server workload with predictable, well known usage...as in any workload where LVM is just additional overhead rather than benefit. Also, in any environment where growth is so fast that its more effective just to build the next server than reconfigure the last one. Suggesting that "its not a server without LVM" is just silly talk.

I suppose this comes back to preference more than anything. The overhead of LVM on a system likely to be 90% idle or higher is insignificant. If you really are so worried about performance that you don't want the overhead of LVM (HPC?) then buy a bigger box ;) Storage growth? In the case of some appliactions, you can't necessarily build the "next one" as the application doesn't function like that. Web? Sure, but even in web they can get into the necessity of file system growth.

Now I could see it in something more like embedded systems where there's potentially zero growth. Even then I'd go with LVM for the simplicity of creating the volumes through LVM rather than creating partitions.

Silly talk? No, we're back to personal preference. Those that prefer not to use LVM will find ways to do so. Those that prefer LVM will most likely use it everywhere, as is the case with the servers I support.
 
Like all technologies, LVM has its use cases. For managing large data volumes (big file server, loads of user accounts), LVM is extremely useful. For managing volume growth in a virtual environment (either local volumes or iSCSI target) LVM would be extremely useful.

I can't really see its value in a web server cluster, where the nodes would likely mount from a central fileserver via NFS (although the fileserver itself might use it), and it would almost certainly be useless in a commodity box running Vyatta or Untangle.

Either way, I would rather see md-RAID and LVM subsumed into each other than not.
 
Like all technologies, LVM has its use cases. For managing large data volumes (big file server, loads of user accounts), LVM is extremely useful. For managing volume growth in a virtual environment (either local volumes or iSCSI target) LVM would be extremely useful.

I can't really see its value in a web server cluster, where the nodes would likely mount from a central fileserver via NFS (although the fileserver itself might use it), and it would almost certainly be useless in a commodity box running Vyatta or Untangle.

Either way, I would rather see md-RAID and LVM subsumed into each other than not.

The value is in standardization and automation. A standard kickstart for us would put a base OS along with our necessary basics, including LVM. Every system gets this then, meaning one kickstart template that Cobbler just modifies for host specifics. Then we provision additional software from there based on server purpose. Standardizing on LVM just means simplicity down the road for the admins as you're dealing with Linux boxes the same way each time. It was always annoying when we had a few Sun boxes on Disk Suite with the rest on Veritas Volume Manager. You always hoped you didn't have to deal with the Disk Suite systems with any disk/volume problems.

Vyatta and Untangle are more black box it seems, in which case I'd just treat them as such. Most likely I'd never touch those boxes as they're function-specific and would be put to use by network and/or security teams. Everything's Red Hat Enterprise Linux.
 
I need to write very big files (video editing) : about 500 Go par file (~ 85 GB per hour during 6 hours)

I was thinking about XFS wich is an extra File System for managing thes kind of file.

All tests we read about ZFS and big file vs another file system are not correct !!
=> when you install ZFS, by defaut compression is ON.

First of all, video files are allready compressed by a codec and trying to compress a video in .rar or .zip or ... means nothing !!

Could it be right to say ZFS speed R/W without compreesion is superieor to XFS ??

Cheers.
 
You certainly can disable compression per ZFS filesystem (I do not believe it is the default.) and ZFS was designed to support large files and such, Give it a try...
 
Indeed, for these type of file :

1) Limit the cache size
2) Disallow file and/or metadata caching on a per fs basis
3) Use an SSD to make a huge cache.​

1) limited ARC (zfs page cache) to 4GB via this line
Code:
/etc/modprobe.d/zfs.conf
options zfs zfs_arc_max=4294967296 zfs_vdev_cache_size=536870912

2) disallowed the video files from being stored in the page cache
Code:
zfs set primarycache=metadata santanas/video
zfs set secondarycache=metadata santanas/video

3) configured a small 32GB SSD as an L2ARC (secondary page cache)
Code:
ZIL : zpool add datas log name of the SSD|partition
L2ARC : zpool add datas cache name of the SSD|partition
.... Benchmark ZFS & L2ARC
=> http://blogs.oracle.com/brendan/entry/l2arc_screenshots
=> http://blogs.oracle.com/brendan/entry/test
=> http://serverfault.com/questions/16...sd-slower-for-random-seeks-than-without-l2arc
=> http://www.anandtech.com/Show/Index...=2&slug=zfs-building-testing-and-benchmarking

4) disable "time update"
Code:
zfs set atime=off santanas

5. Disable compression, deduplication and snapshot
Code:
zfs set compression=off santanas/video
zfs set dedup=of santanas/video
zfs set com.sun:auto-snapshot=false rpool/santanas/video

6. Disable checksumming
Code:
 zfs set checksum=off santanas/video
	zfs set sharenfs=on santanas/video

++
 
Last edited:
6. Disable checksumming
Code:
 zfs set checksum=off santanas/video
	zfs set sharenfs=on santanas/video

Why in gods name would you do that?


And I don't understand what exactly you are doing here - you ask a question, then you post a bunch of random tweaks. What is your goal here? If you are trying to tweak stuff for performance do you have any before/after numbers? I suspect quite a bit of the things you are doing are totally pointless -

ZFS is intelligent about what it keeps in the ARC/L2ARC - it's typically not going to keep a 30GB video file in there - unless that's the only thing you are ever accessing, and then, well, might as well keep it in there.

Compression and deduplication are off by default - you don't need to turn them off unless you turned them on previously.

Why are you limiting the ARC to 4GB? Why not let ZFS manage it itself based on OS memory pressure? (Which is what it does)

But I'm still confused as there is no logical flow from your first post to your second? And generally with ZFS you are better off not tuning - you are ranging from doing harmful things to doing pointless things.
 
The default checksum algorithm is fletcher4 which is very fast. I agree 1000% about not turning that off.
 
And I don't understand what exactly you are doing here
I upgrade my shared stroage from a 8 TB XFS Raid 5 to 32 TB ZFS. (20x 2 To RaidZ2 + 2xSSD mirrored ZIL + 2x SSD mirrored L2ARC) with 10 GbE double port, LSI lsi 3082e-r + expander HP.

I need to handle at the same time 8 video streams (HD) recording each at 185 Mb/s (24 MB/s) + 8 video Low Resolution streams recording each at 2 Mb/s (250 kB/s) during 4 to 8 hours
.... and at the same time, I have to be able to read and edit (several play / stop / fast forward / rewind, plays agin and again ...) simultaneously 1 to 3 of these HD streams on 12 stations.

Code:
zfs set checksum=off filesystem
... to reduce CPU usage.
Code:
zfs set checksum=off santanas/video
zfs set sharenfs=on santanas/video
This "folder" contains videos wich are NFS shared ; once edited in less than 4 hours and brodcasted less than 6 hours from starting to record ... nobody really care about this stuff
=> as soon as videos are brodacsted, they are copied on another storage to be archived.
=> 4 days after, they are deleted from this ZFS storage.

Cheers :cool:
 
Last edited:
There is also a lot more overhead with ZFS (in disk-space) and I don't feel its quite ready yet for primetime linux. I use ZFS on opensolaris on my backup machine though.
you're crazy but w/e. btw if you had read the docs on zfs you would see this 'overhead' is there for a reason. it prevents corruption. if you use 100% of available space on a CoW FS it dies a horrible death.

this minimal space reservation is easily made up by way of dedup and if you really need it compression. features which jfs does not have.
 
I upgrade my shared stroage from a 8 TB XFS Raid 5 to 32 TB ZFS. (20x 2 To RaidZ2 + 2xSSD mirrored ZIL + 2x SSD mirrored L2ARC) with 10 GbE double port, LSI lsi 3082e-r + expander HP.

Sounds reasonable so far - couple of points:
*Mirroring the L2ARC drives isn't typical - *if* an L2ARC drive dies you don't every have any data loss, but you do have a performance loss (whatever extra IOPS that drive was providing). Mirroring would protect you from this. But consider, if you ran both SSD's as L2ARC drives you would have twice as much L2ARC space, and if one dies you would still have the same amount of L2ARC space as you would have had in the mirrored setup setup - so in the failure mode you are in the same boat - but in the non failure mode you only have half as much cache available.

*How much system memory do you have - L2ARC takes up system memory (L2ARC pointers have to reside in the ARC) - if you are still limiting your ARC to 4GB and have a large L2ARC it's possible you are not getting any benefit from your ARC. I would max out your system memory and let the ARC manage itself.

*32TB with 20 2TB drives means you have 2 RaidZ-2 VDEV's, right? What that means is you essentially only have 2 2TB drives worth of "raw" iops (though you get more effective iops from you zil and l2arc of course). Reconfiguring your disks for more VDEV's could give you a significant performance boost - especially as you seem to be having a high random (multiple stream) demand. 20TB (10 2 disk mirrors) is going to give you your highest performance setup - you will essentially have the raw read iops of 20 disks and raw write iops of 10 disks. But you loose 12TB of storage also. If this is just a "scratch" drive then going to 4 5-disk raid-z1 vdev's would give you 32TB workable still, but 4 twice the raw iops. If this is the primary storage then you might not want to drop your protection level down though. It really

It depends on how much space you ultimately need - if you can get by with 20TB then the 10 sets of mirrors would be a very performant setup.

Code:
zfs set checksum=off filesystem
... to reduce CPU usage.

Are you really having an issue with CPU usage? You are essentially giving up the majority of ZFS's data integrity features with this, and really the cost isn't that high. Have you actually measured a before/after difference, or is this just premature optimization?
 
fletcher4 has virtually insignificant CPU cost. And to save that, you are throwing away one of the prime reasons for using ZFS. Sheesh...
 
*Mirroring the L2ARC drives isn't typical - *if* an L2ARC drive dies you don't every have any data loss, but you do have a performance loss (whatever extra IOPS that drive was providing). Mirroring would protect you from this. But consider, if you ran both SSD's as L2ARC drives you would have twice as much L2ARC space, and if one dies you would still have the same amount of L2ARC space as you would have had in the mirrored setup setup - so in the failure mode you are in the same boat - but in the non failure mode you only have half as much cache available.
Yess. 2 cases
=> Somtimes, the storage unit is set on a location without technician
=> Sometimes times the storage could be with a technicien but in a country where spare unit takes time to have (or not ... ie Italy)
*How much system memory do you have - L2ARC takes up system memory (L2ARC pointers have to reside in the ARC) - if you are still limiting your ARC to 4GB and have a large L2ARC it's possible you are not getting any benefit from your ARC. I would max out your system memory and let the ARC manage itself.
SSD L2ARC & ZIL : 4x CORSAIR FORCE 40 GO S-ATA II
RAM : 48 GB (12 4 GB 1066 MHz ECC Registred)
*32TB with 20 2TB drives means you have 2 RaidZ-2 VDEV's, right?
2x 10-disk 2 TB on RAID-Z2 16 TB usable on each pool = 32 TB with and so 4x 2 TB for security "in case of"
Are you really having an issue with CPU usage? You are essentially giving up the majority of ZFS's data integrity features with this, and really the cost isn't that high. Have you actually measured a before/after difference, or is this just premature optimization?
CPU : 2x Xeon E5640
Samba (Windows PC which read the LowResolution) is quite CPU consomming...
We need to test more... but right now, we only have 10 Hard Drives and with the HDD pricelist, now way to go forward.
fletcher4 has virtually insignificant CPU cost. And to save that, you are throwing away one of the prime reasons for using ZFS. Sheesh...
Ok

Regards.
 
For me, it is like car tuning
You can optimize wheels and spoilers but if you really want speed, you need raw power.

In your case, you will not get the needed speed unless you optimize Raid-levels
Your main tuning option is a faster raid-config. For your needs, a Raid Z2/Raid-6 is not suitable
regardless of the filesystem.

Use a pure Raid-0/ ZFS Pool build from single disk vdevs or (i would)
use mirrors with redundancy and the double read performance with half capacity

In your case you do not have sync writes or you can disable sync write for performance.
So use all available SSD's as unmirrored Read cache. Disabling atime is another good setting.
I would not touch any other setting.
 
Last edited:
For me, it is like car tuning
Use a pure Raid-0/ ZFS Pool build from single disk vdevs or (i would)
use mirrors with redundancy and the double read performance with half capacity

Reading : ZFS Performance – RAIDZ vs RAID 1
=> http://www.stringliterals.com/?p=161
seems to say : Raid 10 is ok for small files...
In my case, files are between 80 GB to 500 GB per stream

Let’s first do our standard 100gb write test.

Code:
sa@quasar:~$ time (mkfile 100g /tank/foo)

real    2m29.438s
user    0m0.336s
sys     0m40.903s

This yields a write performance of 669 MB / sec; faster than the RAID10 result of 458 MB / sec. This can be attributed to spreading the workload over more devices, as only 6.25 GB was written to each drive. The stripe of mirrors required that 10 GB be written to each drive. The limiting factor for throughput is still the PCI-X bus, which wrote 836 MB / sec of total information (data + parity) in order to support the payload of 669 MB / sec of data. (669 x 1.25 due to a 4:1 data to parity ratio.)

The following output compares the filebench results from yesterday’s RAID 10 configuration with today’s 5×4 raidz configuration:



Code:
raidz     5x4     rand-read1                273ops/s   4.3mb/s     14.6ms/op      114us/op-cpu
raid10   10x2     rand-read1                412ops/s   6.4mb/s      9.7ms/op       65us/op-cpu

Operations per second dropped from 412 to 273, a drop of nearly 34% in performance for small random reads. This is because more devices had to participate in each individual read operation, reducing the speedup possible through parallel reads, as is possible with mirrors.


In your case you do not have sync writes or you can disable sync write for performance.
So use all available SSD's as unmirrored Read cache. Disabling atime is another good setting.
I would not touch any other setting.
Ok

Regards
 
you have multiple concurrent read write streams on a fragmented filesystem.
you should care about these "best for small files" because thats your real workload.

not to forget:
you would not care about tuning, if your zfs2 ist fast enough
your main tuning option is amount of vdevs
 
Last edited:
People have different priorities but traditionally, RAS (Reliability, Availability, S...?) is most difficult and the most expensive property to achieve. Performance is easy and cheap. Big IBM Mainframes are quite slow, cpu wise. However, they have very good RAS. RAS requires special custom made chipsets, double(triple) redundant cpus and buses and everything. No single-point-of-failure. And the software needs to be rewritten to reroute or redo calculations which are corrupt. Hardware needs to detect corruption. etc. RAS is very expensive and highly sought after in Enterprise.

If you omit all safety checks and just focus on performance, you can easily get that. What do you prefer, a 6GHz cpu where 0.001% of all calculations are wrong - or a 3GHz cpu where all calculations are correct? I know what I prefer.



Same with XFS and ZFS. The only single reason I use ZFS, is it protects against data corruption. XFS does not, as research shows. I discuss this at length in:
http://hardforum.com/showpost.php?p=1037889274&postcount=16
Here is research that shows filesystems XFS, JFS, etc are very bad at detecting data corruption and does not cut it:
http://www.zdnet.com/blog/storage/ho...;siu-container

". . . ad hoc failure handling and a great deal of illogical inconsistency in failure policy . . . such inconsistency leads to substantially different detection and recovery strategies under similar fault scenarios, resulting in unpredictable and often undesirable fault-handling strategies.. . We observe little tolerance to transient failures; . . . . none of the file systems can recover from partial disk failures, due to a lack of in-disk redundancy."


As the OP writes about data corruption
http://hardforum.com/showthread.php?p=1037889274#post1037889274
I just found out that one folder that I copied over manually a few years ago to my server has half viewable images and a few .avi’s that VLC had some issues playing. My question to you is what do you do to verify that the backup? I know that I can open each folder and check. The issue is that I have about 190Gb of media, photos and videos, 300Gb of virtual machines, and about 10Gb of documents that I would like to keep safe.
Is it worth to get a few MB/sec more, and risk this scenario? Fast and unsafe (XFS) - or slow and safe (ZFS)? My data is far too precious for me to risk it.




It is well known that Linux developers cut corners and cheat, just to win benchmarks. If you never do any safety checks, you can have very a fast solution.
Also fsck on my 36 TB array takes around 20 minutes where as the xfs fsck is *much* slower and from what I heard takes 1 GB of ram per 1 TB your volume is.
Assume the read speed of your raid is 300MB/sec. If XFS does a fsck of 36TB in 20 minutes, it means it reads 30.000MB/sec - far more than your read speed. The conclusion is that XFS does not check all data, it only checks some parts. If XFS fsck says "yes, I checked your raid and everything is good, you have no data corruption" - would you trust XFS?

And yes, it is well known that "fsck" only checks metadata. The data is not checked, thus you can still have data corruption after a successful check. And you need to take the raid offline and wait until fsck is done.

ZFS does its "scrub" (fsck) on a live, mounted filesystem. And ZFS scrub takes longer time, but it checks everything: metadata and data on the disk. And you can use your raid while checking.



JFS is very fast for file-deletion. To me it really came down to what I can trust my data too. JFS was developed by IBM and is rock-stable IMHO....Pretty much all my large file-systems use JFS now (ranging from 9->36 TB) ever since I was bitten by running XFS.
As computer science researchers show above, both JFS and XFS might corrupt data. A question, how fast does JFS do a fsck? If JFS does a fsck of 36TB data in 20 minutes too, then it is an indication that JFS fsck is not to be trusted.




Here are other explanations of Linux devs cheating just to win benchmarks, so the solutions are unsafe:
http://milek.blogspot.com/2010/12/linux-osync-and-write-barriers.html
This is really scary. I wonder how many developers knew about it especially when coding for Linux when data safety was paramount. Sometimes it feels that some Linux developers are coding to win benchmarks and do not necessarily care about data safety, correctness and standards like POSIX.

Likewise, Ted Tso, the creator of Linux ext4 agrees:
http://phoronix.com/forums/showthre...38-File-System-Comparison&p=181904#post181904
In the case of reiserfs, Chris Mason submitted a patch 4 years ago to turn on barriers by default, but Hans Reiser vetoed it. Apparently, to Hans, winning the benchmark demolition derby was more important than his user's data. (It's a sad fact that sometimes the desire to win benchmark competition will cause developers to cheat, sometimes at the expense of their users.)

We tried to get the default changed in ext3, but it was overruled by Andrew Morton, on the grounds that it would represent a big performance loss, and he didn't think the corruption happened all that often --- despite the fact that Chris Mason had developed a python program that would reliably corrupt an ext3 file system if you ran it and then pulled the power plug

So again, I would not trust on Linux storage solutions. And as I have said, research in a link above, proves that XFS, JFS, ReiserFS, etc are all unsafe. And researchers have proved that Hardware raid is also unsafe, and might corrupt your data.




Yeah, I never understood why md and LVM were never integrated. It really would make sense - call it LV-RAID. :)
Linux developers thinks it is a really bad idea to let ZFS control everything, a single monolithic piece of code that integrates everything from RAM down to raid to filesystem. Linux devs calls ZFS "rampant layering violation" and thinks you should have several differnt layers instead. That is superior, they say. One layer with raid, one filesystem layer, etc.

However, because ZFS has control of everything, ZFS can in fact check if the data in RAM is the identical to what was stored on disk - ZFS controls the entire chain. Linux filesystems that have different layers can not do that - and that is the reason Linux solutions are unsafe and might corrupt your data. I discuss this at length in the link above. BTW, ZFS is not a monolithic piece of code, it has layers, but differently from normal standard solutions.

ZFS does great things because it violates layers. The ZFS dev team realized the need to violate layers, and discusses this issue in several interviews and documents. Any attempt to clone ZFS will need to violate layers, too. Just like BTRFS - which also violates layers.




I need to write very big files (video editing) : about 500 Go par file (~ 85 GB per hour during 6 hours)

I was thinking about XFS wich is an extra File System for managing thes kind of file.
Say your raw data file is 500 GB big. Then you do a small edit in the middle of the file, the edit is only 10MB big, and you want to save the file again.

With XFS you need to save the file again, and you will waste another 500GB of disk space. Say you do 10 edits and save the file every time. Then you have 10 edits x 500GB = 5TB storage.

I would instead use ZFS for this. Because ZFS uses COW architecture, ZFS has some advantages to this type of work. Say you do a small edit in the middle of the file and you save the entire file, but what happens is that ZFS only saves the edit. The entire file will not be saved again. This is transparent to the application you are using, and it will not know anything. Say you do 10 edits, with ZFS you get 10 edits each 10MB big = 100MB. And then the original file, so in total you use 500GB + 100 MB storage. (For this to work, you must use a feature called ZFS Snapshot. I have done this myself, but with Virtual Machines. I have a single VM as a Master template, and every other VM will only have the changes saved to disk. Say that I have 200 VMs, then I only have one master VM saved to disk, and all the rest of the 200 VMs will only save the changes. There will not be 200 copies of the Master VM. But instead there will be one Master VM, and 200 changes saved to file. This is with ZFS Snapshots.)



These Snapshots, is a real killer feature. Say I install Linux, and then do an upgrade. And the upgrade breaks the kernel. What do I do then? I either fix the problem by booting for a liveCD and issue some commands. Or I reinstall everything. If I get a virus or a hacker root kit, I should reinstall everything anyway.

If instead use Solaris 11 and ZFS, I do like this instead. Before I do an upgrade, I take a Snapshot. If the upgrade breaks anything, I just reboot and in GRUB I see every snapshot. I choose one of the snapshots and boot into it. Then I destroy the upgrade which messed up my kernel. I have undoed the upgrade. This takes a few seconds to do.

Every snapshot writes on a new part of the disk, old data are never touched and left intact. Thus I can back in time, by a reboot. Snapshots are like CVS or SVN of the system disk. I can boot into any snapshot. Say I have a root kit, then I just reboot into an earlier well known functioning snapshot and destroy all latest snapshots. It takes a few seconds. Then all root kits are gone, no more traces of the hacker.

ZFS snapshots are a real killer. XFS does not have this function either.





Regarding ZFS speed.
Phew, it does seem like ZFS is the clear winner between the two, at this point, but let's examine now performance and OS support. XFS beats ZFS in read/write/mixed MB/s throughput benchmarks as well as random IOs/s. Large file creation is faster in XFS too. Random read/writes are higher performing in XFS, especially XFS writes. Standard deviations for random read/writes are close but ZFS does win this category. When it comes to multithreaded read/write/mixed IOs/s, XFS wins by a large margin. In dealing with huge file multithreaded mixed IOs/s, the results are closer but XFS still takes the cake. You'll find untarring to be slightly quicker in XFS but when it comes to tarring, ZFS wins by an enormous margin. Clearly, in overall performance, XFS wins. [data here]
Yes, it is true that ZFS often is slower than other filesystems on a single disk setup. That is because ZFS does lot of checksums and controls. If you dont spend cpu nor ram on safety, then you can get a faster system. And unsafe system. Is it worth it?

However, regarding ZFS performance. ZFS is built by Sun, Enterprise storage with large servers. Thus, ZFS is built for large servers. It scales excellent. Thus, performance is not a problem. If you have a single disk, ZFS might be slowest. But as you put in more and more disks, ZFS continues to scale linearly. Other solutions stop scaling, they are not built for large Enterprise scaling solutions.

For instance, Ted Tso creator ext4, said that he has not had access to large servers as 48 cores earlier.
http://thunk.org/tytso/blog/2010/11/01/i-have-the-money-shot-for-my-lca-presentation/
and for a long time, 48 cores/CPU’s and large RAID arrays were in the category of “exotic, expensive hardware”, and indeed, for much of the ext2/3 development time, most of the ext2/3 developers didn’t even have access to such hardware. One of the main reasons why I am working on scalability to 32-64 nodes is because such 32 cores/socket will become available Real Soon Now

Thus, Linux filesystems might be fast on single disk, or few disks. But when you add many disks, ordinary filesystems lag behind ZFS. Also, you can add SSD to ZFS and increase performance hugely, to 100.000 of IOPS and many GB/sec.

Thus, ZFS is the fastest filesystem out there (because it scales best) when you go into big storage servers. If you use a few disks only, there are other filesystems that are faster (because they cheat and cut corners as Linux developers explained above). Here we see benchmarks of 8 SSD disks BTRFS vs ZFS:
http://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg05647.html


PS. ZFS deduplication is broken. Never use it until it is fixed. Beware.
 
This is a Big Huge review ; thank you very much !!

I'm just a bit septic about the thing you said :
Say your raw data file is 500 GB big. Then you do a small edit in the middle of the file, the edit is only 10MB big, and you want to save the file again.

With XFS you need to save the file again, and you will waste another 500GB of disk space. Say you do 10 edits and save the file every time. Then you have 10 edits x 500GB = 5TB storage.

I would instead use ZFS for this. Because ZFS uses COW architecture, ZFS has some advantages to this type of work. Say you do a small edit in the middle of the file and you save the entire file, but what happens is that ZFS only saves the edit. The entire file will not be saved again. This is transparent to the application you are using, and it will not know anything. Say you do 10 edits, with ZFS you get 10 edits each 10MB big = 100MB. And then the original file, so in total you use 500GB + 100 MB storage. (For this to work, you must use a feature called ZFS Snapshot. I have done this myself, but with Virtual Machines. I have a single VM as a Master template, and every other VM will only have the changes saved to disk. Say that I have 200 VMs, then I only have one master VM saved to disk, and all the rest of the 200 VMs will only save the changes. There will not be 200 copies of the Master VM. But instead there will be one Master VM, and 200 changes saved to file. This is with ZFS Snapshots.)
Really ? ... I used to think in a Video Editing mode, that's not possible !!
Tell me if I'm wrong in demonstrating by this way :
=> if you edit a 500 MB file and extract, let's better saying, 48 MB (corresponding to 2s of video HD broadcast) = in video editing, this is not an extract ;
=>> a new file wich has been processed by DECODEC and then CODEC is created. As we said, the file is 1 new generation (by decompressing & compressing).

// all video editing sowftware can undo up to 32. They link back to their rendered media at every cancel.

++
 
This is a Big Huge review ; thank you very much !!

I'm just a bit septic about the thing you said :...
++
ZFS snapshots works like this:
Whenever you do a snapshot, a new place on the disk is used for all saves. Old data is left intact, and never altered. If you do a new snapshot, a new place on the disk is used for all saves. Say you have a file, "letter.doc" which contains this text "AAA AAA". Then you can do like this with ZFS:

snapshot 1:
letter.doc: AAA AAA

now you create a snapshot 2, and change the contents of "letter.doc" to AAB
snapshot 2:
letter.doc: ...........B

And again:
snapshot 3:
letter.doc: ...........C

Now you have three snapshots on disk. In each snapshot, only the change is saved. In snapshot 1, "letter.doc" contains 6 characters, thus it will be 6 bytes big.

In snapshot 2, "letter.doc" has only changed in one character, thus, letter.doc will be 1 byte big. When I load "letter.doc", ZFS will first load what is needed from snapshot 1, and then load only the change from snapshot 2. Thus, the application will see the file "AAA AAB"

Thus, In my word processing application, I can go into snapshot 1 or 2 or 3, and load whichever version of "letter.doc" I need. All snapshots are active at the same time.

It works exactly like Apple OS X "Time Machine". The difference is that Time Machine have a copy of the file for each edit and wastes huge space. In ZFS everything is taken care of automatically and the application will never know what is going on in the background. ZFS snapshots is all transparent. But you must manually create a snapshot. You do that by typing "zfs snapshot mybigfile@newSnapshotOnMondayTralalalala". This takes one second and you can then work with any snapshot. The original file, or the new snapshot.

Thus, if you have huge files that you need to work with, say 500GB files, then ZFS is the way to go. Because only the changes will be saved. No need to save the entire file again. Every snapshot is located in its own filesystem. Every time you create a snapshot, you create a new filesystem. And those filesystems only contains the changes. But everything is read from the original file.

Was this a better explanation?




UPATE:
// all video editing sowftware can undo up to 32. They link back to their rendered media at every cancel.
Yes, ZFS allows any number of undo. As long as you create a snapshot before a change, you can rollback. One guy had 64.000 snapshots on his system disk, it consumed 3GB of space - because he did very little changes. He snapshotted every 5th second. Just for fun. And can now rollback to which state he wants to.

The difference is that ZFS can have a tree of undo. You can go back to snapshot 23 and work there, but all other snapshots are still left intact. They are not affected. Then you can go to snapshot number 36 and work there, and create several new snapshots. Then go back to 23. etc. The software need not to have this functionality, everything is taken care of by ZFS. Automatic.

In your video editing, if you do a undo back to number 23, then all undo are lost? You need to only work at level 23? You can not have a branch of edits?
 
Last edited:
  • Like
Reactions: oiiwj
like this
ZFS snapshots works like this:
Whenever you do a snapshot, a new place on the disk is used for all saves. Old data is left intact, and never altered. If you do a new snapshot, a new place on the disk is used for all saves. Say you have a file, "letter.doc" which contains this text "AAA AAA". Then you can do like this with ZFS:

snapshot 1:
letter.doc: AAA AAA

now you create a snapshot 2, and change the contents of "letter.doc" to AAB
snapshot 2:
letter.doc: ...........B

And again:
snapshot 3:
letter.doc: ...........C

Now you have three snapshots on disk. In each snapshot, only the change is saved. In snapshot 1, "letter.doc" contains 6 characters, thus it will be 6 bytes big.

In snapshot 2, "letter.doc" has only changed in one character, thus, letter.doc will be 1 byte big. When I load "letter.doc", ZFS will first load what is needed from snapshot 1, and then load only the change from snapshot 2. Thus, the application will see the file "AAA AAB"

Thus, In my word processing application, I can go into snapshot 1 or 2 or 3, and load whichever version of "letter.doc" I need. All snapshots are active at the same time.

It works exactly like Apple OS X "Time Machine". The difference is that Time Machine have a copy of the file for each edit and wastes huge space. In ZFS everything is taken care of automatically and the application will never know what is going on in the background. ZFS snapshots is all transparent. But you must manually create a snapshot. You do that by typing "zfs snapshot mybigfile@newSnapshotOnMondayTralalalala". This takes one second and you can then work with any snapshot. The original file, or the new snapshot.

Thus, if you have huge files that you need to work with, say 500GB files, then ZFS is the way to go. Because only the changes will be saved. No need to save the entire file again. Every snapshot is located in its own filesystem. Every time you create a snapshot, you create a new filesystem. And those filesystems only contains the changes. But everything is read from the original file.

Was this a better explanation?
Yes, thanks a lot.
Snapshot seems to work in the modification of the same files.

If I have a picture xyz.jpg, I change the color and save it under its own name xyz.jpg, ZFS only save the modification of the color, keeping the original picture with thier original tone ?

In video, if I modify the color, I have to render ; so the process of DECODEC and CODEC comes on the road ... to make a new file : the software doesn't modify the color on a part of the original 500 GB file, it creates another one which has the the size of the duration of the colored part ! ;)

That's why I stay seceptic about using snapshot in video editing...

In your video editing, if you do a undo back to number 23, then all undo are lost? You need to only work at level 23? You can not have a branch of edits?
If I do a undo back to number 23, and I save from the 23, I can still undo back 9 times but I can't redo anymore : rendered files are deleted.

++
 
Yes, thanks a lot.
Snapshot seems to work in the modification of the same files.

If I have a picture xyz.jpg, I change the color and save it under its own name xyz.jpg, ZFS only save the modification of the color, keeping the original picture with thier original tone ?

In video, if I modify the color, I have to render ; so the process of DECODEC and CODEC comes on the road ... to make a new file : the software doesn't modify the color on a part of the original 500 GB file, it creates another one which has the the size of the duration of the colored part ! ;)

That's why I stay seceptic about using snapshot in video editing...
If you do an edit that affects every frame, then you will get a full copy of the file. If you do an edit that only affects a few frames, then only those frames need to be saved.



If I do a undo back to number 23, and I save from the 23, I can still undo back 9 times but I can't redo anymore : rendered files are deleted.

++
With ZFS, all are still there on the disk. You can go into every edit no matter how. If you are inside snapshot 23, you can create lot of new snapshots and keep nr 23 as master.

Anyway, easiest would be if you tried it out yourself. Install OpenSolaris / Solaris 11 / FreeBSD or any other ZFS OS inside VirtualBox which is free. Then you can try snapshots. So you keep the raw data file on ZFS, and share that directory via CIFS, NFS or SAMBA - to Windows application.
 
Great !!

Tank you for having taken your time in these explications.

Indeed more tests will be followed. ;)

Anyway, easiest would be if you tried it out yourself. Install OpenSolaris / Solaris 11 / FreeBSD or any other ZFS OS inside VirtualBox which is free. Then you can try snapshots. So you keep the raw data file on ZFS, and share that directory via CIFS, NFS or SAMBA - to Windows application.
What do you thing about ZFS on Linux : http://zfsonlinux.org/
... Onnv Build v147 / pool v28 / FS v5 ??
:confused:

Regards.
 
ZFS under linux is not stable. I would avoid it.

Use any of these OSes where ZFS is stable:
OpenIndiana (the open source version of OpenSolaris), FreeBSD, Solaris 11 Express, OpenSolaris (old, not supported anymore)
 
Back
Top