ZFS build

ldoodle

Limp Gawd
Joined
Jun 29, 2011
Messages
172
Hiya,

I am planning a 15 disk ZFS build (15 is max chassis capacity). I will be using 2TB disks and have decided on the Hitachi 5900rpm SATA III disks (HDS5C3020ALA632), to be used with the new LSI Logic 9201-16i 16 port PCIe HBA. Nothing has been purchased yet and I have in the back of my mind what's said here about 4K disks so i'm still open to suggestions.

These 15 disks are dedicated for user data, as I will be using a 4-in-1 2.5" SATA hot swap caddy for the rpool disks (http://www.scan.co.uk/products/icy-dock-mb994sp-4s-4in1-sas-sata-hot-swap-backplane-525-raid-cage). The rpool disks will be connected to the motherboad SATA ports to keep them entirely seperate to the data disks.

I am however not sure what raid-z* config to go with for the user data. One of these is what i'm thinking;

5x 3 disk raid-z1 (20T usable)
3x 5 disk raid-z1 (24T usable)
3x 5 disk raid-z2 (18T usable)
2x 7 disk raid-z2 (20T usable) + hot spare

The 5 disk groupings would fit well as i'm using 5-in-3 hot swap caddies so I would look at the front of the chassis and know what belongs to what. If I ever scale up I could get another 5-in-3 caddy with 5 more disks, so this keeps it nice and neat. My concern with z1 is the single disk failure - if another drive in the vdev fails while resilvering i'm stuffed, and I guess if I have 3+ vdevs in 1 pool, and one vdev fails due to 2 drive failures, the whole pool is stuffed as well? I'm also curious as to the recommendations re # disks per different raid-z* option in conjuction with 4K disks.

Other than that, my hardware will be;

Tyan n3400b motherboard
AMD Opteron 1212 CPU
8GB ECC RAM

Oh, it's primary use will be to replace Windows Home Server, so basically a media server.

Thanks for your help
 
Depends if you want speed or capacity and or safety

Multiple vdevs = more speed as ZFS starts to then stripe across the vdev's

More safety = les vdevs but more Z's ie RaidZ2 or RaidZ3 etc

Also depends on your budget, ie you could compromise and say make 2 x vdevs of RaidZ2... so you are protected from a double drive failure on each vdev... and you should get up to 2 x a single drives amount of speed across the whole shebang (because of the striping of the vdevs)

if you are still paranoid about losing some thing, then lessen your drives in each vdev... and have 2 drive as spares;)

If you are really paranoid... when placing your order for the 15 drives.... bump it up to 16 drives and you have a spare drive on the shelf.

Only YOU can tell what value to place on your DATA and what it will/would mean to lose it all.

Don't forget the larger the vdevs and the larger the pool.... the longer the scrub time.

I would also suggest more RAM if you can stretch the budget. Maybe google recommended GB of Ram per Terrabyte of Data.

And also remember the only fill the pool to 80% recommendations.. to give ZFS some space to play.

So... work out how many devices will be attacking the system at once (media centres etc) most people find 1 or 2 systems pulling files of of the system with say gigabit connections... a single vdev or 5 drives will keep up fine.

It's the writing to ZFS that can slow dramatically with underspec Ram amounts etc... but it sounds like you are just going to load it up with STUFF then pull from that ... in a Media centre enviroment.. So you won't have much writing going on, except initially to load the box up eh?:D
 
Thanks!

My motherboard is RAM limited to 8GB unfortunately.

I know that more vdevs = more speed as the IOPS is rated at the speed of a single disk for each vdev, so a single 15 disk z* would give IOPS of 1 disk, whereas 5x 3 disk z1's would give the IOPS of 5 disks. I also know that more disks in a vdev = longer resilver times.

I will be putting personal data on there which is irreplacable (holidays pictures etc) so protection is paramount, along with films rips which is replacable as I have the media to do so, but would rather not have to re-rip 700+ films if I don't have to :D

My original plan was the 2x 7 disk z2 option with 1 hot spare but then I read on here about recommended disks per vdev being 6 or 10 for z2 and 3 or 5 for z1. If so the obvious choice is either 3x 5 disk or 5x 3 disk z1s, or 2x 6 disk z2's with 2 hot spares (one per vdev) and leave one physical disk outside of the ZFS system for simple backups < is this possible?

Hmmmmmmm
 
In terms of systems 'attacking' it at once, in terms of read it could be upto 3 or 4 (media PCs), and writing it will more than likely only be my laptop as I do everything on it (ripping, pictures etc) then bulk copy up to the media server.
 
snip....... but then I read on here about recommended disks per vdev being 6 or 10 for z2 and 3 or 5 for z1. If so the obvious choice is either 3x 5 disk or 5x 3 disk z1s, or 2x 6 disk z2's with 2 hot spares (one per vdev) and leave one physical disk outside of the ZFS system for simple backups < is this possible?

Hmmmmmmm

that would depend if you have 4k or 512b drives or not. As the number of drives CAN be a problem with the 4k drives.

Hence the reason most ZFS people either tend to go all 512b type drives not the cheap GREEN drives with the 4k headaches.

Hitachi 5K3000
2Tb
512byte sectors
SataIII/6Gbps
32Mb Cache
'CoolSpin'

Would be my recommendation...
 
8GB is fine
the Hitachi's are great
the scorpio is overkill and no value as an os disk on a server...go a "normal" 2 1/2 and run a mirrored rpool
I run two 7 disk raidz2 vdev's in a single pool in my server....and I get writes of 500-600MB in filebench...so you'll be happy.
 
I've just double checked and I think my motherboard is only PCIe 1.0 not 2.0, which is what the LSI card supports (2.0).

Is this going to have any affect on performance; is it worth upgrading the motherboard to have 2.0 support?
 
that would depend if you have 4k or 512b drives or not. As the number of drives CAN be a problem with the 4k drives.

Hence the reason most ZFS people either tend to go all 512b type drives not the cheap GREEN drives with the 4k headaches.

Hitachi 5K3000
2Tb
512byte sectors
SataIII/6Gbps
32Mb Cache
'CoolSpin'

Would be my recommendation...

My experience with ZFS is that the HD sector size doesn't have any effect on the number of drives, its ZFS's block size of 128k that matters when picking the magic number. 128k will stripe properly on both 512b and 4k drives, and larger sector sizes found on SSD's. However if you chose the wrong magic number like 4 drives in RAIDz or 5 drives in RAIDz2 you wind up with 128k / 3 = 42.666, a funky block size.

Also I'm not confident Hitachi's drives like the 5K3000 are actually 512b sector, if you look at the spec sheet it says "variable", many believe its a hybrid sector size, starting at 512b and working up to 4k.
 
Also I'm not confident Hitachi's drives like the 5K3000 are actually 512b sector, if you look at the spec sheet it says "variable", many believe its a hybrid sector size, starting at 512b and working up to 4k.

Normally variable sector size refers to the ability to allow the IO controller to create a slightly larger sector (528byte) and use the extra space for a checksum on the sector. SATA I don't think supports this (it's a SAS/SCSI feature).
reference:
http://download.intel.com/support/m...e_class_versus_desktop_class_hard_drives_.pdf

I don't think a sector size which changes on a disk would work in any reasonable manner - it would break all sorts of stuff. The firmware could abstract it, but that would be beyond bizzare. And if the firmware is abstracting the sector size it should be listed as 512e (emulated).

Also note the advanced format hitachi drives note the sector size as (512e, variable) as well, so I wouldn't think the "variable" has anything to do with advanced format or not.
reference:
http://www.hitachigst.com/internal-drives/mobile/travelstar/travelstar-5k750
 

This. Putting your irreplaceable data on the ZFS machine is fine...as long as it's not the ONLY location that data resides. Be sure you have your most precious data also backed up elsewhere and preferably a copy offsite too.
 
This. Putting your irreplaceable data on the ZFS machine is fine...as long as it's not the ONLY location that data resides. Be sure you have your most precious data also backed up elsewhere and preferably a copy offsite too.

Of course.

Any idea on the PCIe 1.0 speeds?
 
Oh does anyone know if the Marvell 9480 chipset is supported in Solaris 11 yet? I was originally looking at the Supermicro AOC-SAS2LP-MV8 SAS/SATA 6Gbps normal ATX bracket cards (http://www.supermicro.co.uk/products/accessories/addon/AOC-SAS2LP-MV8.cfm), but they were not supported in Solaris, but the article saying that is over a year old.

£130 for 2 of the Supermicros (16 ports) against £257 for the LSI (also 16 ports) is a big difference.
 
The BR10i is SAS/SATA 2.0 only isn't it, so 3Gbps rather than 6Gbps. Doesn't it also have a funny bracket, similar to the Supermicro UIO cards?

I know gigE won't ever reach the speeds of 6Gbps but i'm thinking in terms of resilvering times or disk to disk speeds, as I guess 6Gbps disks would resilver quicker than 3Gbps, or is that a completely stupid thinkg to say and only the ms seek speeds of the disk make a difference?

It would certainly be cheaper for me as the BR10i is also only PCIe 1.0 so I would not have to upgrade my motherboard, CPU and RAM in the process, as well as being cheaper for 2 cards (about £170 for 2 in UK) rather than 1 of the LSI's
 
what drives are you planning that will do over 3Gbs?

uio is fine...i have 3 uio cards..just remove the vracket or get a suitable one of an old vga card. i run them with no brackets....it's not like you are lugging the server about :)

go the br10i...a gazillion people run these in there zfs servers.
 
Does the BR10i support > 2TB drives? Reading somewhere it may not work with 3TB drives?

I was planning on using the Hitachi CoolSpin 6Gbps 5k3000 disks. No offence but I am too professional to install an add-in card without a bracket, even in a home media server.
 
5k300&#8217;s come in a variety of capacities, including 1.5TB, 2TB and 3TB...confusing naming convention huh!

you mentioned 2TB size in prev posts, so assumed you were getting the 2TB version of this model.

iirc yes, that card does not handle 3tb...iirc, ibm m1015&#8217;s with latest lsi firmware do...so could be a cheap option
 
Out of interest, what stops the BR10i from working with > 2TB disks? Is it firmware or the circuitry of the chipset?

As it's based on the LSI1068 chipset, does that mean any card based on the 1068 will also have the same problem?
 
it works, it just sees it as a 2TB drive...and yes, from what i can see, lsi have said the 1068 does not support 3TB drives...get the 2008 if you want 3 TB.
 
The different PCI-E standards are compatible with eachother. You can plug a PCI-E 2.0 card into a PCI-E 1.0 slot and the other way around without issues - you will obviously only get the speed of the slowest part, which, in case of PCI-E 1.0 is 250 MB/s bidirectional per lane, or 1 GB for a PCI-E 1.0 x4 slot.

PCI-E 2.0 is 500 MB/s bidirectional per lane.
 
Thanks spazoid.

So, if my calculations are correct I would achieve the following (without any overhead included);

Single 16 port card;

PCIe 1.0: 250MB * 8 (PCIe x8) / 16 (# of disks) = 125MB
PCIe 2.0: 500MB * 8 (PCIe x8) / 16 (# of disks) = 250MB

Dual 8 port cards;

PCIe 1.0: 250MB * 8 (PCIe x8) / 8 (# of disks) = 250MB
PCIe 2.0: 500MB * 8 (PCIe x8) / 8 (# of disks) = 500MB

As I said, that's without overhead. Based on this http://blog.zorinaq.com/?e=10, 60-70% can only be achieved in real world terms, so that means;

Single 16 port card;

PCIe 1.0: 175MB * 8 (PCIe x8) / 16 (# of disks) = 87.5MB
PCIe 2.0: 350MB * 8 (PCIe x8) / 16 (# of disks) = 175MB

Dual 8 port cards;

PCIe 1.0: 175MB * 8 (PCIe x8) / 8 (# of disks) = 175MB
PCIe 2.0: 350MB * 8 (PCIe x8) / 8 (# of disks) = 350MB

So it wouldn't be sensible to go with the single 16 port card on PCIe 1.0 any way? Are resilver times [across the same card] improved by going 6Gbps, as if not the only benefit of going 6Gbps would be with the last option (PCIe 2.0 dual 8 port cards), which I won't see anyway due to gigE so the added expense of changing the motherboard, CPU, RAM etc wouldn't be worth it. So I may as well get 2x LSI1068 based PCIe 1.0 cards and keep everything I already have.

However if resilver times are effectively halved by moving to 6Gbps I may be tempted.
 
Last edited:
However if resilver times are effectively halved by moving to 6Gbps I may be tempted.

If your array is composed exclusively of latest gen SSD's then you might see a small (10-15% decrease) in resilvering times - though I doubt it. Otherwise none of the currently available disks are really capable of overwhelming a 3Gbps interface (375MB/sec theoretical) - so the interface speed makes not a lot of difference - as only the drive cache itself can run at that speed.
 
Otherwise none of the currently available disks are really capable of overwhelming a 3Gbps interface (375MB/sec theoretical) - so the interface speed makes not a lot of difference - as only the drive cache itself can run at that speed.

Actually, SATA 3 Gbps has a theoretical max throughput of 300 MB/s. That is because SATA 3 Gbps uses 8b/10b encoding, so the computation is 3 Gbps / ( 8 bits / Byte ) x ( 8 decoded bits / 10 encoded bits ) = 0.3 GB/s = 300 MB/s
 
If your array is composed exclusively of latest gen SSD's then you might see a small (10-15% decrease) in resilvering times - though I doubt it. Otherwise none of the currently available disks are really capable of overwhelming a 3Gbps interface (375MB/sec theoretical) - so the interface speed makes not a lot of difference - as only the drive cache itself can run at that speed.

Actually, SATA 3 Gbps has a theoretical max throughput of 300 MB/s. That is because SATA 3 Gbps uses 8b/10b encoding, so the computation is 3 Gbps / ( 8 bits / Byte ) x ( 8 decoded bits / 10 encoded bits ) = 0.3 GB/s = 300 MB/s

OK, so I guess it's best to go with the cheapest option for now then (LSI1068), as there's no point in making the step up to 6Gbps as a) my motherboard doesn't have PCIe 2.0 and b) gigE won't see the difference anyway.
 
Actually, SATA 3 Gbps has a theoretical max throughput of 300 MB/s. That is because SATA 3 Gbps uses 8b/10b encoding, so the computation is 3 Gbps / ( 8 bits / Byte ) x ( 8 decoded bits / 10 encoded bits ) = 0.3 GB/s = 300 MB/s

Did not know that, thank!

Idoodle said:
OK, so I guess it's best to go with the cheapest option for now then (LSI1068), as there's no point in making the step up to 6Gbps as a) my motherboard doesn't have PCIe 2.0 and b) gigE won't see the difference anyway.

most (if not all) of the 1068 chipsets won't support > 2TB drives, so keep that in mind.
 
most (if not all) of the 1068 chipsets won't support > 2TB drives, so keep that in mind.

I meant to say SAS2008, not 1068! But it's not like i'm going to be upgrading to each iteration of hard disk size as they become available...
 
I meant to say SAS2008, not 1068! But it's not like i'm going to be upgrading to each iteration of hard disk size as they become available...

There seems to be confusion here. The SAS1068E is generally the cheapest option (about $150 for a new 8-port card), but as has been said, probably does not support drives >2.2TB. These would be cards like the IBM BR10i or the Intel SASUC8I.

LSI SAS2008 based cards are 6 Gbps and do support >2.2TB drives, but they are a bit more expensive ($235 for a new 8-port card). These would be cards like the LSI 9211-8i and the IBM ServeRAID M1015. The M1015 can sometimes be found open-box on ebay for less than $100.
 
The Br10i can be had for $50 refurbished! I have an extra one that I want to sell. PM me if you are interested.

I use a Br10i in my ZFS server and its great, got 4x 2TB Hitachi 5k3000 running on it. The rpool is on the motherboard sata controller.

-s0rce
 
Thanks for the offer S0rce, but i'm in the UK so by the time shipping's included it wouldn't be worth it. I will see what they cost new and take it from there.

Contrary to what I said above, I will go for the 1068E as it seems to be the best supported chipset for Solaris. I can't imagine I would move to 3TB drives until they are the cost of 2TB drives, which is a loooooong way away, and by that time there will probably be a new connectivity type so it will be a whole changover, not just disks/controllers.

This should fit me for the next 5 years I would imagine in terms of storage. I am even thinking of going mirrored all the way for 14TB usable....!
 
Thanks spazoid.

So, if my calculations are correct I would achieve the following (without any overhead included);

<removed for viewing pleasure>



Are you under the impression that the total bandwidth will be devided among your drives equally and that your speed will be the result of of PCI-E bandwitdth divided by #drives ?

If so, that is incorrect. The calculation of maximum speed is much more complex than that, but the drives do not share the bandwidth like that. One drive will never be able to exceed its own interface to the controller, though few drives will be limited by the SATA interface (only latest gen SSD's).

Using interfaces of different speed, you will always be limited by the lowest bandwidth interface.

What your calculations can be used for, if anything, is the speed needed by the disks on the controller in order for the disks to be limited by the PCI-E interface, and the only case where that is even remotely realistic, is with a 16 port controller in a x8 PCI-E 1.0 slot.
 
I based it on the link (http://blog.zorinaq.com/?e=10).

For reference, the maximum practical throughputs per port I assumed have been computed with these formulas:
•For PCIe gen2: 300-350MB/s (60-70% of 500MB/s) * pcie-link-width / number-of-ports
•For PCIe gen1: 150-175MB/s (60-70% of 250MB/s) * pcie-link-width / number-of-ports
•For PCI-X 64-bit 133MHz: 853MB/s (80% of 1066MB/s) / number-of-ports

I know that is not an exact science.

Coming back to the sector size, is it actually best to get 512b or 4K? I am just about ready to order what I need and I want to get the 'best' setup I can.

Thanks
 
For reference, the maximum practical throughputs per port I assumed have been computed with these formulas:
•For PCIe gen2: 300-350MB/s (60-70% of 500MB/s) * pcie-link-width / number-of-ports
•For PCIe gen1: 150-175MB/s (60-70% of 250MB/s) * pcie-link-width / number-of-ports
•For PCI-X 64-bit 133MHz: 853MB/s (80% of 1066MB/s) / number-of-ports

What I think he means with "link-width / number-of-ports" is "link-width / number of pci-e lanes" not ports on the controller.
 
Coming back to the sector size, is it actually best to get 512b or 4K? I am just about ready to order what I need and I want to get the 'best' setup I can.

As I am using low RPM drives, I guess it's between these;

Seagate Barracuda LP
Hitachi 5K300
Western Digital Cavier Green
Samsung SpinPoint F4

The WD Green drives are reported to have horrible performance in ZFS due to the sector size. The Segate looks the only one to have a true 512b sector, whereas the Hitachi is listed as 'variable' as mentioned earlier in the thread.

Buuuut, the Segate's are reported to have firmware issues ('click of death').

Who ever would have throught picking a hard drive is so hard!!!
 
Most people seem to recommend the Hitachi 5k3000's, though personally I wouldn't worry too much about it, and just buy whatever is cheaper. Where I live, I could easily buy more disks to make up for the performance deficit and get more space for the same money I would have to pay for the Hitachi's.

YMMV though, depending on pricing in your area.
 
What I think he means with "link-width / number-of-ports" is "link-width / number of pci-e lanes" not ports on the controller.

No, I think ldoodle interpreted mrb correctly. mrb was calculating the expected bandwidth per HBA port, assuming all HBA ports were occupied with drives.

So, for PCIe 2.0, which has 500 MB/s per lane, on an 8-lane PCIe bus, and a 16-port HBA, you have:

overhead_factor x 500 MB/s per lane x 8 lanes / 16 ports = overhead_factor x 250 MB/s (per port)

mrb was assuming 60% to 70% throughput after overhead, so that would be 150 - 175 MB/s per port in my example. Which is exactly what mrb computed for the 16-port SAS2116 HBA.

For reference, the maximum practical throughputs per port I assumed have been computed with these formulas:

For PCIe gen2: 300-350MB/s (60-70% of 500MB/s) * pcie-link-width / number-of-ports
For PCIe gen1: 150-175MB/s (60-70% of 250MB/s) * pcie-link-width / number-of-ports
For PCI-X 64-bit 133MHz: 853MB/s (80% of 1066MB/s) / number-of-ports
 
Most people seem to recommend the Hitachi 5k3000's, though personally I wouldn't worry too much about it, and just buy whatever is cheaper. Where I live, I could easily buy more disks to make up for the performance deficit and get more space for the same money I would have to pay for the Hitachi's.

YMMV though, depending on pricing in your area.

Funnily enough the 5K3000's a the cheapest of the 4 for me.

In terms of the controller cards I can get them for the following prices. Bear in mind I would need 2 of whatever model to give 16 ports;

Intel SASUC8I - £111.62 (SAS 3Gbps PCIe 1.0)
LSI SAS3081E-R - £133.50 (SAS 3Gbps PCIe 1.0)
IBM SerRAID BR10i - £185.56 (SAS 3Gbps (PCIe 1.0)
LSI SAS9211 - £153.00 (SAS 6Gbps PCIe 2.0)

I would prefer LSI original ones due to fimware updates and support (I guess if I flashed the Intel or IBM with stock LSI firmware they would not support them), so for the £20 difference per card it would be criminal to go with the older 3081E-R card over the 9211.

If I ever decide to upgrade my motherboard I will have PCIe 2.0 cards ready to go.
 
Back
Top