• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Faster, Higher, Stronger : powerful storage

Here: http://hardforum.com/showpost.php?p=1036642118&postcount=156

With

6x 2TB Samsung F4s, a x3440 with 8G of RAM on a Supermicro X8SI6-F.

User devilmouse posted his results

Code:
Now testing RAIDZ2 configuration with 6 disks: cWmRd@cWmRd@
READ:	353 MiB/sec	352 MiB/sec	= 353 MiB/sec avg
WRITE:	372 MiB/sec	335 MiB/sec	= 353 MiB/sec avg
=> RaidZ2 not enough fast for me.

Do you think I can outperform the speed with Constellation ES.2 or WD RE4 ?

Cheers.

St3F
 
Not by much, that is 4 disks of usable data, at approx 100MB/sec per disk, so that is basically maxing out that setup.

If it had more disks, that raidz2 could go faster, like 10 disks, should be getting a good 700MB/sec, if the cpu can keep up.
 
Not by much, that is 4 disks of usable data, at approx 100MB/sec per disk, so that is basically maxing out that setup.

If it had more disks, that raidz2 could go faster, like 10 disks, should be getting a good 700MB/sec, if the cpu can keep up.
Alright ; to bad for me.

So, I think I have no other choice to go with
  • 2x 10 mirror or 2x 12 mirror (with system in another rack)
  • 3 TB Seagate constellation ES.2
  • 4x spare disk or 8x spare disk (with system in another rack)
  • 2x controller
... in the case of 2 disks failure in 1 mirrored vDev at the same time,
=> 2 spare drives will automatically replace these 2 fault disk ?
=> how long does it take to rebuild with this storage capacity ?

Cheers.

St3F
 
Alright ; to bad for me.

If your hardware is fast enough, you can add another Raid-Z2 to the
pool to nearly double performance (then similar to Raid 60), or add a third, fourth..
Raid Z2 for more.

So, I think I have no other choice to go with
  • 2x 10 mirror or 2x 12 mirror (with system in another rack)
  • 3 TB Seagate constellation ES.2
  • 4x spare disk or 8x spare disk (with system in another rack)
  • 2x controller
... in the case of 2 disks failure in 1 mirrored vDev at the same time,
=> 2 spare drives will automatically replace these 2 fault disk ?
=> how long does it take to rebuild with this storage capacity ?

Cheers.

St3F

If you have a double failure in the same raid-1 vdev, your pool is lost.
If the double failure is in different vdevs and you have two hotspares,
both will be replaced automatically.

Rebuild time is between 30min and a few hours in case of mirrors
 
If your hardware is fast enough, you can add another Raid-Z2 to the
pool to nearly double performance (then similar to Raid 60), or add a third, fourth..
Raid Z2 for more.
So, if I follow he said :
Not by much, that is 4 disks of usable data, at approx 100MB/sec per disk, so that is basically maxing out that setup..
@ 100MB/sec per disk, 6 disk RaidZ2 = 4 disks of usable data = 400 MB/s

With a 24x 3,5" rack storage, I can setup 4x vDev 6 disks RaidZ2... as Raid60.

In this kind of configuration, could I reach 1600 MB/s ?
Does the couple Supermicro X8DTH-6F + Xeon E5606 reach this rate ?

IIf you have a double failure in the same raid-1 vdev, your pool is lost.
If the double failure is in different vdevs and you have two hotspares,
both will be replaced automatically.

Rebuild time is between 30min and a few hours in case of mirrors
Ok ; understood.

Cheers.

St3F
 
Last edited:
@ 100MB/sec per disk, 6 disk RaidZ2 = 4 disks of usable data = 400 MB/s

With a 24x 3,5" rack storage, I can setup 4x vDev 6 disks RaidZ2... as Raid60.

In this kind of configuration, could I reach 1600 MB/s ?
Does the couple Supermicro X8DTH-6F + Xeon E5606 reach this rate ?

If you are using really fast disks (non 4k disks), you may reach such values
in a single user benchmark where you read or write sector per sector/ track by track.

But if another user need data from another track, these values are useless.
If you look in extreme to multiuser sync writes (ex ESXi on NFS), where each write
must be confirmed, write values can go below 10% of these sequential values.

For multi-user, you should not look only to sequential values. I would concern of I/O
beeing not too bad and I would try to increase RAM to a level where most reads on
current files can be served from RAM to keep sequential read and writes values
from disk as high as possible.

But without testing a special config under your workload, nobody can say, ok good enough.
So you need to max it out or you need to test a medium config that can be improved.
 
That should give you somewhere in the 1400-1600MB/sec range.

You will have 4x2 disks worth of overhead too. So it would have to all be a nice 6g expander, or just direct to disks to make it so the overhead parity disks don't kill your transfer speed.

X8SI6-F would work fine, but is limited to 32gigs of ram, I think for your specs, you will be shooting yourself you can't expand memory more. This also affects how much l2arc you could use if you wanted to add any in. It also seems to be very limited in pcie slots.
 
But if another user need data from another track, these values are useless.
If you look in extreme to multiuser sync writes (ex ESXi on NFS), where each write
must be confirmed, write values can go below 10% of these sequential values.
Yes, keep in in mind : files will not be under 20 MB and will reach about 350 GB of 20 video streams (120 Mb/s each) simultaneously

From : http://www.zdnet.com/blog/storage/chunks-the-hidden-key-to-raid-performance/130
Do you do video editing or a lot of Photoshop work? Then your average request size will be large and your performance will be dominated by how long it takes to get the data to or from the disks. So you want a lot of bandwidth to move data quickly. To get a lot of bandwidth you want each disk to shoulder part of the load, so you want a small chunk size. What is small? Anywhere from 512 bytes (one block) to 8 KB.

From : http://www.solarisinternals.com/wik...2.2C_RAIDZ-3.2C_or_a_Mirrored_Storage_Pool.3F
A RAIDZ configuration maximizes disk space and generally performs well when data is written and read in large chunks (128K or more).

For best performance I will use NFS wich is iSCSi with sharing capability.
For these video streaming feeds, I will tweak the 10 GbE network card for large buffer chunk size (at least 512 kb)

I read, NFS requires sync writes
.. http://hardforum.com/showthread.php?t=1684910
=> for NFS you said "- you do not need a l2arc with large files, think about more RAM like 192 GB-256 GB"
=> for ZIL you advice using DDRdrive or ZeusRam disc witch perform Random IOPS (4k Block): up to 100,000.
In Europe, these product are hard to find and very expensive : STEC ZeusRAM Z4 8GB SAS = ~ 3200 € ex VAT !!! :eek:

What about
- OCZ Vertex 3 Max IOPS SATA III 120 GB wich perform Random IOPS (4k Block): up to 85 000 and R/W @ up to R/W 500MB/s ? ( ~158 € ex VAT) ? ... they are MLC, but if I buy 4, in Raid 10 could be ok ?
- OCZ RevoDrive 3 X2 Max IOPS PCI-Express SSD wich perform Random IOPS (4k Block) up to 220 000 @ R/W 1 500 MB/s (~668 €ex VAT) ? ... they are MLC, but if I buy 2, in Raid 1 could be ok ?
- OCZ Deneva 2 C Sync 60 Go wich perform Random IOPS (4k Block) up to 65 000 @ R/W 500 MB/s (~ 153 € ex VAT) ... they are MLC, but if I buy 4, in Raid 10 could be ok ?
- Plextor M3 Pro 128 Go wich perform Random IOPS (4k Block) up to 75 000 @ R/W 500 MB/s (~ 145 € ex VAT) ? ... they are MLC, but if I buy 4 in Raid 10 could be ok ?
- including the command vfs.zfs.cache_flush_disable=1 in /boot/loder.conf ?
... the guy say "If it's on the ZIL, why do we need to flush it to the drive? A crash at this point will still have the transactions recorded on the ZIL, so we're not losing anything."
Some guys tested successfully this command, even with power failure : http://forums.freebsd.org/showthread.php?t=30856

For multi-user, you should not look only to sequential values. I would concern of I/O
beeing not too bad and I would try to increase RAM to a level where most reads on
current files can be served from RAM to keep sequential read and writes values
from disk as high as possible.
From my draft, I'm modifying the configuration :

Main :
  • 3U with 16 x 3,5"
  • Motherboard X8DTH-6F-O
  • CPU : 2x 5640
  • Ram : 192 GB (4x KIT Crucial 48 Go (3 x 16 Go) DDR3 1066 MHz CL7 ECC Registered QR X4)
  • System Drive : 2x Intel Solid-State Drive 520 Series 120 Go (5 years of warranty) Raid 1 / mirrored
  • ZIL / LOG : ??
  • L2ARC : none ... will be using only NFS
  • Spare drives for storage : 8x WD RE4 2 To
  • Controller : LSI SAS 9207-4i4e
  • Ethernet : Intel Ethernet Server Adapter X520-SR2 : Dual-port 10 GbE SFP+ optic SR
  • PSU : Dual 620w redundant
Storage JBOD

That should give you somewhere in the 1400-1600MB/sec range.
Yummy !! ♥

X8SI6-F would work fine, but is limited to 32gigs of ram, I think for your specs, you will be shooting yourself you can't expand memory more. This also affects how much l2arc you could use if you wanted to add any in. It also seems to be very limited in pcie slots.
X8DTH-6F-O not X8SI6-F ;)
 
Last edited:
What about
- OCZ Vertex 3 Max IOPS SATA III 120 GB which perform Random IOPS (4k Block): up to 85 000 and R/W @ up to R/W 500MB/s ? ( ~158 € ex VAT) ? ... they are MLC, but if I buy 4, in Raid 10 could be ok ?
- OCZ RevoDrive 3 X2 Max IOPS PCI-Express SSD which perform Random IOPS (4k Block) up to 220 000 @ R/W 1 500 MB/s (~668 €ex VAT) ? ... they are MLC, but if I buy 2, in Raid 1 could be ok ?
- OCZ Deneva 2 C Sync 60 Go which perform Random IOPS (4k Block) up to 65 000 @ R/W 500 MB/s (~ 153 € ex VAT) ... they are MLC, but if I buy 4, in Raid 10 could be ok ?
- Plextor M3 Pro 128 Go which perform Random IOPS (4k Block) up to 75 000 @ R/W 500 MB/s (~ 145 € ex VAT) ? ... they are MLC, but if I buy 4 in Raid 10 could be ok ?
I reply to me by myself :D

- OCZ RevoDrive 3 X2 Max IOPS PCI-Express SSD : need drivers too be recognized // only for Windows :(
- OCZ Vertex 3 Max IOPS SATA III 120 GB only has 3 years of warranty
- OCZ Deneva 2 C Sync 60 Go is a bit old and only has 3 years of warranty + no capacitor // Go to Deneva 2 R series D2RSTK251M11--0200
- Plextor M3 Pro 128 Go is a good product for L2ARC
+ from : Sebulon => http://forums.freebsd.org/showpost.php?p=180668&postcount=116
I have however bought an ACARD 5.25" SATA-II SSD RAM DDR2 - ANS-9010BA, and it worked horribly, haha It was really crappy. After starting a write to the device with dd or something, messages was flooded with write DMA errors and then it vanished from the OS.

.... but for ZIL wich need better write operation & data rate, Plextor is behind Intel 520 "Cherryville" - 120 Go which can perform Random IOPS (4k Block) up to 80 000 @ 500 MB/sec

So, for my use : files between 20 MB to 350 GB, transfer by NFS, no SSD for L2ARC but 96 Go Ram, is Intel 520 "Cherryville" - 120 in Raid 10 good for ZIL ?

Cheers.

St3F
 
Last edited:
I reply to me by myself :D

- OCZ RevoDrive 3 X2 Max IOPS PCI-Express SSD : need drivers too be recognized // only for Windows :(
- OCZ Vertex 3 Max IOPS SATA III 120 GB only has 3 years of warranty
- OCZ Deneva 2 C Sync 60 Go is a bit old and only has 3 years of warranty
- Plextor M3 Pro 128 Go is a good product for L2ARC
.... but for ZIL wich need better write operation & data rate, Plextor is behind Intel 520 "Cherryville" - 120 Go which can perform Random IOPS (4k Block) up to 80 000 @ 500 MB/sec

So, for my use : files between 20 MB to 350 GB, transfer by NFS, no SSD for L2ARC but 96 Go Ram, is Intel 520 "Cherryville" - 120 in Raid 10 good for ZIL ?
Going on my research, I keep on sharing here my find.

For ZiL : max IOPS + Max write datarate
=> SSD SLC
=> SSD MLC with supercap, as these with Intel G3, Marvell C400 or Sandforce SF2000 controller.​

Intel 520 is not supercap capable
OCZ Deneva 2 C Sync 60 Go is not supercap capable
OCZ Deneva 2 R Serie has capacitance, so is supercap
OCZ Talos 2 R 200 Go 2,5'" TL2RSAK2G2M1X-0200 has capacitance, so is supercap + SAS ... but 3 years of warranty

SSD with Intel G3, Marvell C400 or Sandforce SF2000 controller seems owning supercap.

.... to be followed :)
 
Last edited:
The storage I'm willing to build is for Video application, writing and reading files from 20 MB to 320 GB...

So, let's say, this is a Huge HTPC network storage, shared over NFS.

Storage review tests the Intel 710, which is supercap equipped, for my required application :

The first real-life test is our HTPC scenario.
In this test we include:
.... playing
1 x 720P HD movie in Media Player Classic,
1x 480P SD movie playing in VLC
3x movies downloading simultaneously through iTunes,​
... recording
and 1x 1080i HDTV stream being recorded through Windows Media Center​
=> over a 15 minute period.

Higher IOps and MB/s rates with lower latency times are preferred. In this trace we recorded 2,986MB being written to the drive and 1,924MB being read.

intel_710_200gb_storagemark2010_htpc.png

:cool:

Other SSD with the same HTPC test.
Orher SSD with Supercap : Left green for Intel 320, right blue for Corsair Performance Pro SSD

intel_320_300gb_storagemark2010_htpc.png
corsair_performance_pro_256gb_repeatingrandom_storagemark2010_htpc.png

Other SSD from 2012 marker

ocz_vertex4_512gb_storagemark2010_htpc.png

Here a WONDER comparison between Intel 320, Intel 710 and DDRDrive x1
=> http://www.slideshare.net/OpenStora...r-drive-zil-accelerator-by-christopher-george (from page 29)

Cheers.

St3F
 
Last edited:
Zeus vs Intel

Zeus IOPS 16GB
Code:
  Local writes:
  raw            64 MB/s

  raw bs=128k    133 MB/s

  Score as ZIL:
  raw            55 MB/s

Intel X25-E 32GB

Code:
  Local writes:
  raw            77 MB/s

  raw bs=128k    197 MB/s

  Score as ZIL:
  raw            60 MB/s

=> http://forums.freebsd.org/showpost.php?p=137905&postcount=37

NFS Mirrored ZIL

Code:
Ordinary HW
Intel 320    40GB  = 30MB/s
OCZ Vertex 2 60GB  = 32MB/s
OCZ Vertex 2 120GB = 36MB/s
Intel 320    120GB = 52MB/s
Zeus IOPS    16GB  = 55MB/s
Intel X25-E  32GB  = 60MB/s
OCZ Deneva 2 200GB = 67MB/s
OCZ Vertex 3 240GB = 70MB/s

=> http://forums.freebsd.org/showthread.php?t=23566
 
Last edited:
I'm thinking ZFS could not be the right way to share BIG DATA through NFS !!!!

I'm reconsidering the question of the NFS use !! ... Shall I have to go on CIFS / SMB ?!?
.. or AFP (workstaiton connected will be Mac) => http://mattconnolly.wordpress.com/2010/06/13/zfs-performance-networked-from-a-mac/

From : http://www.hob-techtalk.com/2009/03/09/nfs-vs-cifs-aka-smb
As our tests are now finished, we can say the winner is NFS: If you are reading from a server share, the results are slightly better than for CIFS. When writing to a server share, CIFS is clearly faster for all writing benchmarks.

If you are interested in the results with charts, please download this pdf
 
Last edited:
I'm thinking ZFS could not be the right way to share BIG DATA through NFS !!!!

I'm reconsidering the question of the NFS use !! ... Shall I have to go on CIFS / SMB ?!?
.. or AFP (workstaiton connected will be Mac) => http://mattconnolly.wordpress.com/2010/06/13/zfs-performance-networked-from-a-mac/

From : http://www.hob-techtalk.com/2009/03/09/nfs-vs-cifs-aka-smb


If you are interested in the results with charts, please download this pdf

Andrew Galloway's blog wrotes on 03/01/2012 "The Case Of The Mysterious ZIL Performance"
=> http://nex7.com/node/12

We were testing with a single write stream, from a single dd process(dd if=/dev/zero of=/volumes/test-pool/test-folder bs=8K count=150000, with sync=always and compress=off). It turns out that what we perceived to be honestly pretty terrible ZIL IOPS potential wasn't.. entirely accurate. You see, this logic that does a 'for' loop through each log vdev, writing and syncing and then going on to the next, will in fact quiesce any and all writes that have come in between the time he began his last commit to disk and his next one into that next one. So you see, you can only do 3000 IOPS on this particular device, but with a single stream dd it is 3000 8K IOPS. If, instead, you do say 16 threads of dd at 8K, now it is doing still 3000 IOPS, but they're 128 KB per IOP! This 'bucket' approach is what we just didn't realize happened at first, by way of our poor initial testing methodology of a single dd. A single dd writing sync to something is waiting for that acknowledge back for each 8K IOP, so it can never put more than one 8K block into each IOP alone.. you need to run multiple threads simultaneously to achieve that.

Now to be clear, 16 threads of dd at 8K is actually over 300 MB/s, which means it can't actually sustain that either, you've more than maxxed out the 3G SAS limit (which this drive is) by then, but the point is there. The mystery is solved. The ZFS ZIL can and will utilize multiple top-level vdevs defined as log devices, but it will do so in a round-robin capacity, and it will wait until it has completed a transaction against one top-level vdev before talking to the next (and while waiting, will queue up to 128K of data to write in the that next operation). If the vdevs are not very low-latency, that latency will become their Achilles' Heel, and seriously impact the ZIL (and thus also the latency of writes from all your clients). This is why utilizing devices like a STEC ZeusRAM or ZeusIOPS are so critical. Their average write latency is significantly lower than most other disks, including other SSD's (especially in the case of the ZeusRAM, since it is in actuality RAM), both in part to their great design and also because they effectively instantly answer (ignore) the sync() request ZFS does, as they have their own battery backup and can safely do that, further reducing the write latency ZFS perceives.


In summary -- this is all yet further proof that latency can be a serious killer, and that investing in some good, very low-latency log devices for workloads that have a log of ZIL traffic is absolutely critical to achieves success with ZFS, instead of introducing a bottleneck into your pool configuration. It is very fair to say that even if your chosen log device has reportedly extremely high IOPS, we'll never notice it with how we write to it (send down, cache flush, and only upon completing send down more -- as opposed to a write cache utilizing sequential write workload) if it does not ALSO have a very, very low average write latency (we're talking on the low end of microseconds, here).
 
Last edited:
LSI Nytro could be fine ... but prices ... oO

Code:
Nytro WarpDrive SLC-based solutions

    Nytro NWD-WLP4-200:              $6,596                                            
    Nytro NWD-WLP4-400(2):          $12,195

Nytro WarpDrive eMLC-based solutions            

    Nytro NWD-BLP 4-400:              $5,945                                   
    Nytro NWD-BLP 4-800:              $10,895         
    Nytro NWD-BLP 4-1600(2):        $20,795         

Nytro XD caching bundle                                      

    Nytro NXD-BLP 4-400:               $10,695
    Nytro NXD-BLP 4-800(2):           $15,645
 
I'm thinking ZFS could not be the right way to share BIG DATA through NFS !!!!

why? zfs does big data well. zfs and nfs work well together. you're over thinking this which is good and bad. it appears you're doing everything you can to read and educate yourself, this is a good thing. it also appears that youre taking all this information all into account all at the same time ... this is a bad thing.

what is your workload?

large files and huge files.

so lets examine performance IO.

there are two main areas of IO performance and 2 (main/broad) types of workloads

1) latency

2) bandwidth

latency dictates how quickly you can respond to or deliver requests.

bandwidth is how much data you can move per second.

Performance profile 1:
small 4k-16k random IO

here, the number one governing factor for performance is latency. how quickly you can respond to these small requets dictates your performance. in this realm, bandwidth only servers as a measure of how many of these small IO can be handled at once.

within the ZFS world you manage this profile in 3 main ways.
1) lots of ram
2) and or lots of l2arc
3) log device (slog/zil)

or
4) pools made entirely of SSD (can use a slog here too for even better performance)

Performance profile 2:
large file sequential transfer

here the governing factor is bandwidth where bandwidth means the individual link speed and the bandwidth to the HBA itself. this profile is the reverse of the small random IO profile where latency is governed completely by bandwidth.

say you have a 100MB/s file and your bandwidth is 1MB/s. your latency to service that request is 10s. if your bandwidth is 1000MB/s your latency is .1s

within zfs if you're building specifically for this performance profile you tackle it like so
1) lots of spindles in raidz2 sets (where spindle = many sas/sata connections)
2) multiple HBAs in x8 pci-e 2 slots with vdevs created across HBAs
3) l2arc sized for current workload.

lets look at the first point. in this thread folks have suggested mirrors and z2 to you. neither is incorrect however as you've stated usable space is a priority. given this fact, z2 is the route to take. yes, z2 sucks for random IO but random IO is NOT your workload. yes, you'll have X number of people accessing different things at different times or simultaneously but that is still an incredibly small 'random' load that i will cover on point 3.

multiple HBAs, lets get extreme here. so you want to create 4 6 disk raidz2 sets. cool, you can do that with 3 HBAs or a single HBA depending on the chassis/expander. using a fairly standard LSI 2 port card you have a theoretical maximum of 48gbps or 6GBps. Looks good right? but wait, you aren't accounting for the PCI-e bus which if it is a pci-e 2.0 x8 slot then you're max transfer is 4GB/s. if that were a pci-e 3.0 x8 slot with a pci-e 3.0 HBA youre going to shift that bottleneck to the HBA where 48gbps/6GBps is 2GBps below what the slot can transfer.

so what if you had 6 HBAs, each with 6 disks directly connected and you create your raidz2 vdevs using a single disk on each hba?

presuming you have pci-e 2.0 x8 slots for all of these (unlikely for 1366 boards, standard for socket r/2011 sandy boards) well then you have 24GB/s of motherboard slot bandwidth.

caveats, spinning disks dont push HBAs as is so there is no point to doing that, just using it as an example for how to add more bandwidth.

lastly, point 3, you seem to have a good idea for how large these files will or won't be. this makes sizing l2arc very simple. you need enough l2arc to store whatever you're working read set is likely to be. keep in mind there is no 'raid' with l2arc so everything you atach is usable. if you have 1TB of data being read at any given time, add 4 512GB SSDs or 8 256GB SSDs. 8 256 drives will allow you to service more concurrent reqeusts (your 20 users) but you may not hae ports/slots for those and frankly 20 users aren't going to push 4 SSDs that hard.

once those working sets are cached in l2arc, and they will get cached, performance will be stellar.


dont over complicate things. figure out your workload, build for it, go have cocktails.
 
Thank you for this review madrebel.

Things are missing :
  1. expandable capability
  2. sharing protocol
  3. I need to write 1 feed at least @ 1600 MB/s or write 4 different feeds at the same time @ 15 MB/s during the read of 24 other different feeds @ 15 MB/s

So, after reading more and more, I've been thinking to operate like this :

1. Expandable capability

Node

JBOD

=> A. 1x LSI SAS 9212-4i4e external connected to 1x Expander Intel res2sv240 to deserve 2 vDevs :
- 6x 2 To vDev 1
- 6x 2 To vDev 2
- connector OUT to future JBOD

=> B. 1x LSI SAS 9212-4i4e external connected to 1x Expander Intel res2sv240 to deserve 2 vDevs :
- 6x 2 To vDev 3
- 6x 2 To vDev 4
- connector OUT to future JBOD

2. Sharing protocol

SMB/CIFS : (to avoid slow NFS sync writes to ZIL)
- for 1 workstation on MacPro which only writes up to 4x feeds at the same time @ 15 MB/s each feeds
- 2 of these feeds will write a 350 GB file each ; 2 other feeds will write a 10 GB file each ... all at the same time

NFS :
- for the 8 other workstations which read up to 4 feeds at the same time @ 15 MB/s each or write only 1 feed @ 15 MB/s

3. Performance

Seems more sequential than random.
When reading : 2s latency is ok
When writing, less than 1s
No pool made only of SSD (budget limited to 25 k$ for 32 TB usable)
Other system are 16x 2 TB Raid 6 as Supermicro VTrak or Sonnet and seems to be ok for this kind of application.

So, what do you think ? Am I making a movie, a sweet dream or a very bad nightmare ?

Cheers.

St3F
 
  1. expandable capability
  2. sharing protocol
  3. I need to write 1 feed at least @ 1600 MB/s or write 4 different feeds at the same time @ 15 MB/s during the read of 24 other different feeds @ 15 MB/s
expansion is a facility of adding HBAs (direct connect to jbods) or sas switches (1 or more HBAs in, 1 to 8 or more jbods out) or daisy chaining jbods which requires a controller in the jbod that allows that (off the top, cant think of a jbod that doesnt).

unless you have a religious aversion to using windows NFS client i would use NFS since you appear to be standardizing on 10gig for the filer at least. although .. actually if you absolutely have toe write at 1.6GB/s you can't do that with 10gig on iscsi or NFS (writes are inherently single path even with mpio iscsi or 802.3ad NFS). so if 1.6GB/s is an actual metric you have to hit you're only choice is 40gig infiniband either running rdma or IPoIB.

*snip*
[*]HotSpare drives : 4x 2 TB (WD RE4 or Constellation ES 2) on
overkill. i'm not saying don't use hotspare but 4 hotpsares is a lot of wasted space IMO.
SMB/CIFS : (to avoid slow NFS sync writes to ZIL)
there is ALWAYS a ZIL, regardless of write, regardless of local to the system or from network. ZIL is always resident in RAM. nfs, iscsi, and frankly any network protocol that enforces o_sync or f_sync will suck from a small random IO standpoint if you do not have a second log device (slog, zil drive).

i can tell you nfs with the zeusRam screams on 10gig. ;atency and IOPs are incredible.
Seems more sequential than random.
your load as described is almost 100% sequential regardless of concurrency.
When reading : 2s latency is ok
2s is terrible unless you're talking full transfers of large files.
When writing, less than 1s
see above.
No pool made only of SSD (budget limited to 25 k$ for 32 TB usable)
understandable, they are spendy.
So, what do you think ? Am I making a movie, a sweet dream or a very bad nightmare ?
honestly, if your transfer requirements are legitimate ... 10gig network cannot do 1.6GB/s for a single transfer. if you have a dual port card and 1.6GB is system wide peak then you can break that up across the 2 ports fine however ...

i don't think you have enough disk to push 1.6GB/s. you may at first but once you get past the outer 40% of tracks your throughput will go down below that number.

that budgetary number is pretty slim considering the performance requirements. you're close to meeting the numbers theoretically but in production, with overhead ... you're expecting a lot from not much hardware.
 
expansion is a facility of adding HBAs (direct connect to jbods) or sas switches (1 or more HBAs in, 1 to 8 or more jbods out) or daisy chaining jbods which requires a controller in the jbod that allows that (off the top, cant think of a jbod that doesnt).
LSI SAS 9212-4i4e HBA has external connector.
I'd like to chain JBODS ... you said it required a controller in the jbod : expander in the JBOD is not ok ?
.. with 2 HBA in the NODE and 2 expanders in the 1st JBOD, I used to think I can chain to other JBODS. (as I read)

unless you have a religious aversion to using windows NFS client i would use NFS since you appear to be standardizing on 10gig for the filer at least. although .. actually if you absolutely have toe write at 1.6GB/s you can't do that with 10gig on iscsi or NFS (writes are inherently single path even with mpio iscsi or 802.3ad NFS). so if 1.6GB/s is an actual metric you have to hit you're only choice is 40gig infiniband either running rdma or IPoIB.
All clients are Mac !
The workstation which only writes feeds is a Mac too directly connected to the 10 GbE
The ZFS system owns a dual 10 GbE
The other 10 GbE port is connected to a switch.
If 1,2 GB/s is possible, that's could be fin.

overkill. i'm not saying don't use hotspare but 4 hotpsares is a lot of wasted space IMO.
There are 8 Hotspare in the Node (3U) : 4 on the HBA A and 4 on the HBA B
.... the room where the storage will be is in a building which is very difficult to come in.
So, less intervention is done, better I'm free. :)
Moreover, it's located about 120 miles away from our location.

i can tell you nfs with the zeusRam screams on 10gig. ;atency and IOPs are incredible.
Can you develop ?
2s is terrible unless you're talking full transfers of large files.
Yes indeed : if I press "Play" button, no problem to wait 2s to what to the stream

that budgetary number is pretty slim considering the performance requirements. you're close to meeting the numbers theoretically but in production, with overhead ... you're expecting a lot from not much hardware.
As said, other equipments at the same price, in Raid 6 with 16 hard drive, do the job... so I expect to ZFS to give all its potential ! ;)

Cheers.

St3F
 
Last edited:
As said, other equipments at the same price, in Raid 6 with 16 hard drive, do the job... so I expect to ZFS to give all its potential ! ;)

look ZFS is really powerful and can be really fast, no doubt. however, you're kind of designing for best case, again, i'm presuming your numbers (IO requirements) are accurate.

16 drives in raid 0 are going to do 1.6GB/s read and write. however, 16 drives in raid10 are not going to write at 1.6GB/s. whatever you have in production now cannot under any circumstances sustain 1.6GB/s a second write unless you're doing raid0. It just isn't possible. with raid0 each drive has to sustain 100MB write speed to hit that number. in raid10 (mirroring) each write drive (yes all drives are writing but only half will contribute to overall throughput) would need to sustain 200MB/s to hit 1.6GB/s. No spinning drive can sustain that for any length of time. even the brand new 1TB raptors from what I saw can only hit those transfer speeds for the outer tracks.

Also keep in mind when you're reading benchmarks, many times those are ideal conditions. In production there will be overhead.

If you're needs/numbers are what you actually need then what you're planning will fall short. If you're planning to build something better than what is already in production, you will likely exceed it but you won't hit 1.6GB/s write speeds. it just isn't feasible.
 
That test is using SVM, software raid without ZFS.
Ok, here is another test with ZFS. It uses 16 SSD disks and delivers 2.5 GB/sec
http://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg05689.html

The more spindles, the faster it gets. If you need to hit 1.6 GB/sec with disks, just add more disks.

Here is a ZFS server which hits 2 GB/sec NFS in best case. Maybe you can mimic this server, and build a copy of it? Or build a beefier clone?
https://blogs.oracle.com/brendan/entry/up_to_2_gbytes_sec

And it does 150.000 IOPS in best case.
https://blogs.oracle.com/brendan/entry/a_quarter_million_nfs_iops

Here are some new ZFS servers beating the competition. Maybe you can clone one of these?
https://blogs.oracle.com/si/entry/7420_spec_sfs_torches_netapp
 
Last edited:
those last 2 links are for a caliber of hardware that is well beyond his budget. also, that last oracle test I am almost positive was done using striping/raid0.

while not dishonest, you don't run striping in 99.9% of production environments.
 
Yes, my point was that ZFS can give very good performance and wanted to prove that. We also have the IBM supercomputer Sequoia that uses Lustre + ZFS and handles 55 PetaBytes and achieves 1TB/sec.

Hence, there are no performance limitations in ZFS, if the OP believed that.
 
I have a modest question -why not AFP instead of CIFS/NFS ?
My ZFS storage is on FreeNAS 8 and definitely the AFP transfers are the fastest -close to the 1GbE limitation. My Windows PCs can reach not more of 70% of this bandwidth.
 
Last edited:
Back
Top