• A Great friend to the HardForum with a great kid that he is trying to get a scholorship to continue his schooling. Please give hime a vote! Only 24 hours left! Thanks.
    If you have an VOTE FOR KEENAN!

Using consumer SATA SSDs for ZFS

madilyn

n00b
Joined
Feb 20, 2014
Messages
17
I asked this elsewhere but no one figured out an answer... So I understand the issues with consumer SATA drives are:

1. Driver problems when put behind SAS interposer and expander.

2. Typically lacks supercapacitor for flushing write cache in power loss.

3. Very little overprovisioning.

4. Usually low life spans, unsuitable for 24/7 duty cycle.


Will I be fine to use consumer SATA drives on my ZFS server if I took these steps to mitigate the problems then?

1. Remove the SAS expander(s) on my server chassis and attach the SSDs directly to SAS HBAs internally (no separate head node and JBOD chassis) using SFF-8077 to 4x SATA fanout cables.

2. Use redundant power supplies and UPSes.

3. Partition the drive only to use 80% of actual storage space.

4. Get high MTBF drives, e.g. Samsung 840 Pro = 1.5 million hours MTBF.

The only problem I foresee with my setup is that it's a hassle to hot-swap and organize the drives if they are attached using fanout cables. I'll need a spacious chassis and it won't be densely populated. But I wouldn't mind doing this to avoid spending $5.5 per GB for SAS SSDs or stringing lots of 15K RPM SAS HDD three-way mirror vdevs.

Crucial M500 is obviously the most economical now, although I know the Samsung 840 Pro has a huge following - anyone using either on their ZFS servers?
 
Can you shut down the server when it is not used? Or spin down the disks?

Regarding SATA disks and SAS expanders - it is rumoured that Oracle Solaris has solved those problems in Solaris 11.1 or so. There seem to be no source on this though, only rumours.
 
Can you shut down the server when it is not used? Or spin down the disks?

Regarding SATA disks and SAS expanders - it is rumoured that Oracle Solaris has solved those problems in Solaris 11.1 or so. There seem to be no source on this though, only rumours.

Probably yes although I haven't figured out the implementation specifics of doing so. Is there a reason why I would do that? This ZFS server is acting as a storage for a cluster that has jobs queued up 24/7.

Quite a bummer that they're getting progress on versions >28, I'll probably use a open source variant, e.g. Linux or FreeBSD so I don't enjoy those fixes.
 
This is a false dichotomy. There ARE open source solaris-based ZFS implementations. AFAIK, omnios and openindiana are both based on illumos...
 
I just killed two SSDs, one of them a X-25, when experimenting with ZFS over AES-XTS. The latter amplifies the already heavy write activity coming down from ZFS and of course trim will never work through it. Pop pop went the drives. I'm poorer by two SSDs and richer in experience.

I tried killing a 256 GB Samsung 840, didn't succeed yet.
 
I just killed two SSDs, one of them a X-25, when experimenting with ZFS over AES-XTS. The latter amplifies the already heavy write activity coming down from ZFS and of course trim will never work through it. Pop pop went the drives. I'm poorer by two SSDs and richer in experience.

I tried killing a 256 GB Samsung 840, didn't succeed yet.

Mm, good, I think I'll get 840s. Only thing I don't like about them are the 512 GB capacity limits. Without a separate JBOD (since I'm not using an expander) or makeshift cabling workarounds, I'm probably limited to the number of disk bays on a single chassis per file system. I'm thinking of three-way mirrors, and the most 2.5" bays I have on a DataOn/Supermicro case is 24, which leaves me to under 4 TB of usable pool space. Maybe I'll use 2.5-to-3.5" trays on a 60 bay 4U case.

This is a false dichotomy. There ARE open source solaris-based ZFS implementations. AFAIK, omnios and openindiana are both based on illumos...

I'm sorry. I meant that I'd be using only FreeBSD or Linux - both of which are stuck at ZFS v28.

That said, if I recall correctly, the open source Solaris-based ZFS implementations forked ZFS somewhere between version 28 and 34, so I thought latest changes to ZFS in Oracle Solaris wouldn't be brought to the open source derivatives. I could be wrong.
 
What about using SSDs as L2ARC against a set of spindle hard drives? I remember reading somewhere that people ran into issues if they had too large of an SSD L2ARC or maybe it was too many SSDs that made up an L2ARC.

I remember seeing a webpage somewhere that had a SuperMicro board that had an onboard LSI 8 port SATA controller that used an all-SSD array with 8 drives. Can't find that link now, but it sounds like it'd be a viable option too.

UPDATE: Here's the page I was thinking of, Gea using 6x480GB SSDs: http://forums.servethehome.com/processors-motherboards/1020-x9srh-7tf-s2011-lsi-2308-dual-10gbe-~$500-2.html
 
What about using SSDs as L2ARC against a set of spindle hard drives? I remember reading somewhere that people ran into issues if they had too large of an SSD L2ARC or maybe it was too many SSDs that made up an L2ARC.

I remember seeing a webpage somewhere that had a SuperMicro board that had an onboard LSI 8 port SATA controller that used an all-SSD array with 8 drives. Can't find that link now, but it sounds like it'd be a viable option too.

UPDATE: Here's the page I was thinking of, Gea using 6x480GB SSDs: http://forums.servethehome.com/processors-motherboards/1020-x9srh-7tf-s2011-lsi-2308-dual-10gbe-~$500-2.html

Thanks! Very interesting, I didn't think of building a huge L2ARC. My initial thoughts were that my use case is very random IO-intensive. This means 1 problem: The L2ARC scales by striping, so I'd indeed need to stripe 6-7 disks for the L2ARC to achieve similar random IOPS performance as a 6-7 vdevs of 3-way mirrors (ignoring low-level details like cache calculations and the various bus latencies). The collective MTBF of this L2ARC would be trash, defeating the purpose of other things I've done to keep this running smoothly 24/7 (though not yet HA/no SPOF).

--> 21 SATA SSDs + no L2ARC, relative cost = $8,400
--> 21 15k RPM SAS drives + 7 SATA SSDs L2ARC + 305~ watt differential i.e. $264.18 per year, relative cost = $5,848
--> 21 15k RPM SAS drives + 7 SAS SSDs L2ARC + 305~ watt differential i.e. $264.18 per year, relative cost = $10,748

Price-wise they're all very close (considering the processors, chassis, memory etc. will add another 10k+). I'm still inclined to 21 SATA SSDs because it offers the best performance of the three, at the expense of having none of the benefits of SAS.
 
A couple of questions here: first off MTBF is not an issue here. By design L2ARC is fail-safe. Every single one of them can fail without disrupting service. Secondly, I'm not sure I believe 3 spinning disks can match the IOPs of a single (decent) SSD.
 
Last edited:
A couple of questions here: first of MTBF is not an issue here.

OK.

By design L2ARC is fail-safe. Every single one of them can fail without disrupting service.

Awesome, that's great to know. Am I correct to say this is fine even if I stripe the L2ARC disks?

Secondly, I'm not sure I believe 3 spinning disks can match the IOPs of a single (decent) SSD.

Not at all, just some Fermi estimates: Each spindle averages about 200~300 random IOPS versus a regular SSD which is about 20,000 random IOPS. Just back-of-the-envelope, I imagine I'll need to build like 100+ vdevs of three-way mirrors for similar performance... which isn't too bad if you consider that SAS 15K spindles are much cheaper than SAS SSDs.

Either ways though, all these just suggest to me that SATA SSDs have a great price point and I hope to hear more user experiences of using SATA SSDs for ZFS without SAS expanders.
 
L2ARC is a read cache, so you can literally pull them out one by one as they are being used, with no bad effects. As far as random IOPs, what kind of spinners are we talking about? For random SATA I've usually heard numbers like 100 or so random IOPs.
 
Also, it depends on your workload and working set. I have a 4TB raid10 with 8 SAS nearline drives. Most of the workload is a handful of virtual machines whose disks live on an iSCSI LUN. I get about 90% hit in the ARC (16GB of ram). About half the rest hits in the L2ARC (two 128GB SSDs), so the spinners are mainly persistent storage and don't get much activity...
 
L2ARC is a read cache, so you can literally pull them out one by one as they are being used, with no bad effects. As far as random IOPs, what kind of spinners are we talking about? For random SATA I've usually heard numbers like 100 or so random IOPs.

Also, it depends on your workload and working set. I have a 4TB raid10 with 8 SAS nearline drives. Most of the workload is a handful of virtual machines whose disks live on an iSCSI LUN. I get about 90% hit in the ARC (16GB of ram). About half the rest hits in the L2ARC (two 128GB SSDs), so the spinners are mainly persistent storage and don't get much activity...

HGST Ultrastar 15K600: http://www.tomshardware.com/reviews/enterprise-storage-sas-hdd,2612-6.html Here's its nearline sister: http://www.storagereview.com/hgst_ultrastar_7k4000_enterprise_sas_hdd_review

Thanks!

My workload is a Postgres database, about 60 tables.

30 tables -> 100 GB
1 table -> 1 GB
9 tables -> 10 MB
20 table -> 1 MB

The <=1 GB tables are always used and will fit into ARC.

My gripe is that the size of the most frequent workload is on the same order of magnitude as the size of the database - the frequency at which we access every 100 GB table is uniform; they are used in a manner that I can't imagine effective caching... half of them are mission-critical logs that we'll be accessing (<20 concurrent users, fortunately) and writing 24/7, the other half are used equally often (we sort of rotate the most frequently used 5-10 on a daily basis).

Hm but having a large L2ARC is looking more compelling now...
 
What I see is a fairly small percentage (but large amount space wise) of data that is never in either cache. The working set works fine for me, but maybe not you. Yeah, though, if you can get all but many/most of the tables in L2ARC you will be happy. Keep in mind the L2ARC is write through but still the read caching should be helpful.
 
I seem to recall that there were issues with 3 or more drives as an L2ARC. Found the post I was thinking of: http://nex7.blogspot.com/2013/10/lots-of-l2arc-not-good-idea-for-now.html

Might be something to keep in mind if you do go with a large SSD L2ARC. Otherwise, it seems best practices all point to maxing out your RAM before moving to an L2ARC solution. Might also look at getting an SLOG, but those that are recommended are also VERY expensive.
 
@raiderj: Thanks a lot! This narrows down my options... I'm thinking that a PCIe SSD may be a better L2ARC device than a regular SAS SSD. There are a few competitively priced choices of PCIe SSDs: OWC Mercury Accelsior E2, Intel 910, OCZ Z-Drive R4 C-series (no supercapacitor version), but I'm worried of using these as L2ARC now as it seems that the controllers on these present a single card as 4 (or more?) discrete drives. Any thoughts on these?

I will definitely have dedicated ZIL/SLOG devices (SLC NAND) even in my SSD pool build. My reasoning is that I imagine there will be asymmetric wear on my pool SSDs if the ZIL cache was on those instead. There's probably some configuration or hack to get around this, but I'm not a ZFS expert and I don't think it will give the best performance at the end of the day.

Already maxing out the RAM within practical reason... (256 GB registered ECC). I'll need to get other processors and >=32 GB DIMMs (expensive) to achieve more than that.
 
Just another thought: http://www.anandtech.com/show/6614/microncrucial-announces-m500-ssd-line-of-ssds I realized that the M500s eliminate the data loss issue faced by consumer SATA SSDs by having both a regular capacitor and no write cache at all. I'm very certain that I will experience fewer disk failures on 840 Pros, but making some cost-benefit analysis (each M500 yields about twice the GB/$), I decided that I will go for the M500s if I did a full-SSD pool.
 
Yeah.

I mean I didn't kill the Samsung yet. Fine. But that's a couple of hours using the "trick" that took out some others. I haven't thought about new tricks for the 840 yet :)

I maintain that consumer SATA SSDs are garbage and I only keep games on them and other recreatable mostly readonly data.

I am also thinking about writing a specific benchmark, or load program, that directly exposes the latencies that we here describe as "freezing". The way I see it the bar is just a bit higher today. No way that a cheap SSD came up with magic bullets.

Plus it is still too easy to lose the ability to TRIM, e.g. if you use any kind of raw-to-raw device layer such as encryption. I think a lot of people probably run without it and don't know it.
 
Last edited:
Use a SSD with a capacitor. Or you will LOSE data or get data corruption.
http://www.extremetech.com/computin...might-be-the-only-reliable-drive-manufacturer

Could you elaborate? It seems the M500 doesn't have a volatile write cache so even if you lose power in the middle of a write, there's nothing that needs reserve power to be flushed onto permanent storage.

The article that you linked suggested only the S3500 and S3700 have power loss protection (basically, they have capacitors on-chip for reserve power), but they didn't test the M500 which also has 20 ceramic capacitors on-chip. See: http://www.thessdreview.com/wp-content/uploads/2013/12/Crucial-M500-M.2-NGFF-SSD-Back.png

Thanks! :)
 
Yeah.

I mean I didn't kill the Samsung yet. Fine. But that's a couple of hours using the "trick" that took out some others. I haven't thought about new tricks for the 840 yet :)

I maintain that consumer SATA SSDs are garbage and I only keep games on them and other recreatable mostly readonly data.

I am also thinking about writing a specific benchmark, or load program, that directly exposes the latencies that we here describe as "freezing". The way I see it the bar is just a bit higher today. No way that a cheap SSD came up with magic bullets.

Plus it is still too easy to lose the ability to TRIM, e.g. if you use any kind of raw-to-raw device layer such as encryption. I think a lot of people probably run without it and don't know it.

Let me know what you find, I'd be interested to know!

I guess I have too much faith in commodity-off-the-shelf parts. I worked on a different team - but when I was at Google, I noticed that we had a LOT of c-o-t-s SATA disks and we didn't have a problem. The high-level strategy was to invest more in recovery than in precaution. We used to buy them in bulk from SYNNEX; I wonder where to get good volume discounts nowadays. ;)

Gave it a lot of thought and I'm guessing no one tried the M500s so I'll push a new frontier and try them! It will probably end up badly haha. :p
 
Let me know what you find, I'd be interested to know!

I guess I have too much faith in commodity-off-the-shelf parts. I worked on a different team - but when I was at Google, I noticed that we had a LOT of c-o-t-s SATA disks and we didn't have a problem. The high-level strategy was to invest more in recovery than in precaution. We used to buy them in bulk from SYNNEX; I wonder where to get good volume discounts nowadays. ;)

Gave it a lot of thought and I'm guessing no one tried the M500s so I'll push a new frontier and try them! It will probably end up badly haha. :p

I'd be curious to see how those M500s works for you too. Seems like if you had those and a UPS you could handle common power failures without needing to invest in more expensive SSDs with supercapacitors. And like you say, invest that saved cash into proper backups and recovery systems. Plus, ZFS should be able to recover on its own with minimal data loss.
 
Just another thought: http://www.anandtech.com/show/6614/microncrucial-announces-m500-ssd-line-of-ssds I realized that the M500s eliminate the data loss issue faced by consumer SATA SSDs by having both a regular capacitor and no write cache at all. I'm very certain that I will experience fewer disk failures on 840 Pros, but making some cost-benefit analysis (each M500 yields about twice the GB/$), I decided that I will go for the M500s if I did a full-SSD pool.

My home NAS has two M500s and a M4 in a raidz pool, gonna swap the M4 for another M500 in due time. Used as an Iscsi target. Freebsd/Freenas 9 has trim for zfs.
 
Back
Top