• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

RaidZ3 has much smaller available disk space than expected

MarcusXP

Gawd
Joined
Sep 23, 2007
Messages
926
Hello,

I've recently set-up a ZFS box using 23 x 4TB Hitachi drives in RaidZ3
According to this calculator:
http://www.servethehome.com/raid-calculator/
the expected size of the pool should be about 72.8TB (taking in consideration the 3 drives lost for parity and the conversion between 1000 and 1024 - TB/TiB)
However, I am only getting 65.7TB...
This is the listing showing the available space (removed some unrelated information):

root@openindiana:~# zfs list
NAME USED AVAIL REFER MOUNTPOINT
Pool-Storage1 65.7T 2.85G 460K /Pool-Storage1
Pool-Storage1/Volume-Storage1 65.7T 65.7T 230K -
rpool 16.8G 2.78G 47K /rpool


This is my pool with 23 x 4TB drives:

Pool-Storage1 5000 14237568027059623914 vdevs: 1
vdev 1: raidz3 12 92.02 TB 12843423673337650988
c10t5000CCA22EC008ADd0 11897191478853965569 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec008ad:a
id1,sd@n5000cca22ec008ad/a
PL1310LAG029NA
c10t5000CCA22EC00A7Cd0 4700561813297779659 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec00a7c:a
id1,sd@n5000cca22ec00a7c/a
PL1310LAG02TLA
c10t5000CCA22EC00F4Fd0 18078600867002041605 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec00f4f:a
id1,sd@n5000cca22ec00f4f/a
PL1310LAG042EA
c10t5000CCA22EC01A8Dd0 17529955265994712780 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec01a8d:a
id1,sd@n5000cca22ec01a8d/a
PL1310LAG0728A
c10t5000CCA22EC0210Ed0 521233862270374400 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec0210e:a
id1,sd@n5000cca22ec0210e/a
PL1310LAG08TZA
c10t5000CCA22EC021D9d0 16504663500950716529 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec021d9:a
id1,sd@n5000cca22ec021d9/a
PL1310LAG090JA
c10t5000CCA22EC02208d0 17286534736261518309 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec02208:a
id1,sd@n5000cca22ec02208/a
PL1310LAG0921A
c10t5000CCA22EC0221Bd0 11471414523034938154 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec0221b:a
id1,sd@n5000cca22ec0221b/a
PL1310LAG092NA
c10t5000CCA22EC0222Bd0 10137454426417349138 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec0222b:a
id1,sd@n5000cca22ec0222b/a
PL1310LAG0935A
c10t5000CCA22EC02234d0 9661598762332469726 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec02234:a
id1,sd@n5000cca22ec02234/a
PL1310LAG093GA
c10t5000CCA22EC02287d0 2707340434424007771 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec02287:a
id1,sd@n5000cca22ec02287/a
PL1310LAG0964A
c10t5000CCA22EC022ACd0 1409056393389450611 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec022ac:a
id1,sd@n5000cca22ec022ac/a
PL1310LAG097AA
c10t5000CCA22EC022E7d0 13341020487462083677 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec022e7:a
id1,sd@n5000cca22ec022e7/a
PL1310LAG0997A
c10t5000CCA22EC02D19d0 7239933575572325830 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec02d19:a
id1,sd@n5000cca22ec02d19/a
PL1310LAG0D0EA
c10t5000CCA22EC032E8d0 11341704217867462188 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec032e8:a
id1,sd@n5000cca22ec032e8/a
PL1310LAG0EKDA
c10t5000CCA22EC032F8d0 14003683354051041944 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec032f8:a
id1,sd@n5000cca22ec032f8/a
PL1310LAG0EKXA
c10t5000CCA22EC0330Bd0 15602634426552806117 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec0330b:a
id1,sd@n5000cca22ec0330b/a
PL1310LAG0ELJA
c10t5000CCA22EC0340Bd0 1306860508072491531 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec0340b:a
id1,sd@n5000cca22ec0340b/a
PL1310LAG0EVTA
c10t5000CCA22EC04C15d0 16571289578325581979 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec04c15:a
id1,sd@n5000cca22ec04c15/a
PL1310LAG0N89A
c10t5000CCA22EC04C26d0 14950040610680530891 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec04c26:a
id1,sd@n5000cca22ec04c26/a
PL1310LAG0N8VA
c10t5000CCA22EC04C35d0 6692930116809901728 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec04c35:a
id1,sd@n5000cca22ec04c35/a
PL1310LAG0N9AA
c10t5000CCA22EC04C48d0 10360558934820987983 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec04c48:a
id1,sd@n5000cca22ec04c48/a
PL1310LAG0N9YA
c10t5000CCA22EC04C87d0 5707996551731273821 Hitachi HDS5C404
/scsi_vhci/disk@g5000cca22ec04c87:a
id1,sd@n5000cca22ec04c87/a
PL1310LAG0NBZA

Can anyone more experienced with ZFS tell me what is going on?
Where is the ~7TB lost space? Is it reserved for snapshots, or anything else? Can I recover it?

thanks a lot,
-Marcus
 
And the answer is....after parity ZFS by default will keep 10% of available space in reserve. That is the "missing" space your seeing. You CAN change this but I wouldn't unless you really need to. This reserved free space is to allow proper function/performance of the file system should it become full. This is not unlike the way an SSD needs reserved space to maintain it's proper performance.
 
And the answer is....after parity ZFS by default will keep 10% of available space in reserve. That is the "missing" space your seeing. You CAN change this but I wouldn't unless you really need to. This reserved free space is to allow proper function/performance of the file system should it become full. This is not unlike the way an SSD needs reserved space to maintain it's proper performance.

I thought it was only 1/64th of the total space which is reserved by ZFS... which is almost negligible compared to amount of space lost that I see.
http://cuddletech.com/blog/pivot/entry.php?id=1013
Do you have a link of this statement that says 10% of disk space is reserved?
Also, how do I decrease this number, I don't need extreme performance this would be used for backups mostly and most data will rarely ever change, so high performance is not a requirement.
 
Well, you really made a very strange pool there.

with 20 data disks, a 128k block will get striped at 6.4k written per disk. That means on your 4k disks, it will use 8k, so 1.6k wasted per disk per write.

So over 20 disks, the 128k block will use 160k of diskspace, making about 20% of your diskspace wasted.

Patrick replied this on the duplicate thread I opened for same issue:
http://hardforum.com/showthread.php?p=1039865399
Could it be the case? Honestly I don't quite understand the relationship between the 128k and 160k of diskpace explained here :(
 
I'm no ZFS expert, but that seems like a lot of HDs in just 1 VDEV.

I think 23-24 drives are not too many for 1 VDEV, considering it's RaidZ3... but I am not an expert either. I've had 24 drives 3TB in Raid-6 before (hardware RAID) with no issues, and I thought ZFS would be more smart :)
 
Also, how do I decrease this number, I don't need extreme performance this would be used for backups mostly and most data will rarely ever change, so high performance is not a requirement.

Try recreating the pool with ashift = 9, see if that helps a bit with the lost space. Test the performance also, it will be lower.
 
And the answer is....after parity ZFS by default will keep 10% of available space in reserve. That is the "missing" space your seeing. You CAN change this but I wouldn't unless you really need to. This reserved free space is to allow proper function/performance of the file system should it become full. This is not unlike the way an SSD needs reserved space to maintain it's proper performance.

This is only the case when using napp-it
napp-it sets a base pool reservation of 10% per default to avoid that a filesystem
can go above 90% of poolsize, resulting in a very bad performance,

If you need the extra space, you can delete or lower this reservation.
 
I did create the pool with napp-it but I am pretty sure I un-checked the box that reserved space..
Is there a command I can run so that I see what is the reserved space on the pool?
 
I did create the pool with napp-it but I am pretty sure I un-checked the box that reserved space..
Is there a command I can run so that I see what is the reserved space on the pool?

If you open menu ZFS Filesystems in napp-it you can check/set the pool reservation under RFRES
The 10% base reservation is set on pool-level to reduce the available space for all filesystems
 
I assumed your disks are 4k sector.

If that is the case, zfs will use ashift=12 when it makes the pool, meaning the smallest write it does will be 4k.

When you stripe the 128k block of data over 20 disks, that will be 6.4k per disk, so it needs two 4k blocks to store it on, meaning a 128k will use 160k of diskspace.
 
I assumed your disks are 4k sector.

If that is the case, zfs will use ashift=12 when it makes the pool, meaning the smallest write it does will be 4k.

When you stripe the 128k block of data over 20 disks, that will be 6.4k per disk, so it needs two 4k blocks to store it on, meaning a 128k will use 160k of diskspace.

Thanks a lot for your explanation, how would you recommend to re-configure the storage to use the maximum available space?
If 23 drives are not optimal for RaidZ3, should I use another number of drives for the pool, or can I change the block size to something that fits better, or should I try using RaidZ2 instead?
 
I think 23-24 drives are not too many for 1 VDEV, considering it's RaidZ3... but I am not an expert either. I've had 24 drives 3TB in Raid-6 before (hardware RAID) with no issues, and I thought ZFS would be more smart :)

It's not a matter of it 'not working' or being 'not smart', just less than optimal in terms of performance. If you aren't doing much in the way of random IO it's okay. You could double your IOPs by splitting 24 drives in raidz3 to 2 vdevs, each raidz2, without almost the same safety. Up to you though...
 
Does ZFS allow a non-standard block size (that is, different than 16k, 32k, 64k, 128k)?
Maybe I can create the pool from command like with blocksize of 60k or 80k or 120k or anything that goes well on 20 drives?
 
It's not a matter of it 'not working' or being 'not smart', just less than optimal in terms of performance. If you aren't doing much in the way of random IO it's okay. You could double your IOPs by splitting 24 drives in raidz3 to 2 vdevs, each raidz2, without almost the same safety. Up to you though...

there is no question about the performance here, it is a question of missing/wasted space (a lot of it). That's what I meant with "not smart" :)
I never had this kind of issue with wasted space on a hardware raid controller, no matter what block size I chose.. there it was a matter of performance issue, but the space available was the same...
 
Is the pool empty? I thought it looked like it was full with only 2gigs availble in his posting.

But no, you can only select a power of two sectorsize.

This has nothing to do with raid, raid just passes the blocks given to it back to the drives.

This is more a filesystem issue, zfs uses the largest blocksize it can to store a file, upto 128k normally. Other filesystems generally default to 4k blocksizes (ntfs can go higher, but unix normally 4k is the largest).

If you where to store lots of 1k files on ntfs with 64k clustersize, you will have the same issue.

If you where to store 1k files on ext3 with 4k blocksize, again, same issue, 3k wasted per file.

Here zfs stores 128k per vdev, with 1 23disk raidz3 vdev, that 128k gets split 20 ways + partity, causes 6.4k per disk, round that up to the 4k sector size, and you used 8k to store 6.4k

So now it looks like you formatted your system with 4k blocksize, and stored LOTS of 3.2k files on it.

This is specific to zfs though, as other filesystems don't split things up like this, cause they don't join the raid + filesystem levels together.

So the goal when using raidz on zfs, if you wanted MAX space, is to use a base of 2 number of data disks, 2, 4, 8, 16, and add your parity to them.

The other option is, calculate the wasted space, and see if the extra disks you added beyond optimal gets you more than you wasted.
 
This is an EMPTY pool, that is ALL the space available on it.. I created a ZFS filesystem which almost filled it, but that volume is also empty... and when I created the ZFS filesystem it wouldn't let me assign more than ~65TB on it (or so).
So yeah, I'm a bit puzzled as to where the extra storage is... I should have around 72.8TB but I only have 65.7TB available to use.

If the pool is empty, then it has nothing to do with the block size.. ? Maybe it is the 10% reserved space.... I will check when I get back home today...
 
I think it would be better to do 2 RAIDZ2 vdev of 11 drives or 12 drives each if you could fit one more in.

Writing will be much faster with 2 vdevs. On my single vdev pool I can't saturate gigabit due to the ~5 second pause where the pool computes parity and writes the blocks.

recordsize Property

The recordsize property specifies a suggested block size for files in the file system.

This property is designed solely for use with database workloads that access files in fixed-size records. ZFS automatically adjust block sizes according to internal algorithms optimized for typical access patterns. For databases that create very large files but access the files in small random chunks, these algorithms might be suboptimal. Specifying a recordsize value greater than or equal to the record size of the database can result in significant performance gains. Use of this property for general purpose file systems is strongly discouraged and might adversely affect performance. The size specified must be a power of 2 greater than or equal to 512 bytes and less than or equal to 128 KB. Changing the file system's recordsize value only affects files created afterward. Existing files are unaffected.

The property abbreviation is recsize.

http://docs.oracle.com/cd/E18752_01/html/819-5461/gazss.html#gcfgv


https://blogs.oracle.com/roch/entry/tuning_zfs_recordsize
 
Performance of 1 vdev of 10 - 2TB drives.

aggregatedlinks_zps6e4d8602.png


More vdevs = better.
 
I think it would be better to do 2 RAIDZ2 vdev of 11 drives or 12 drives each if you could fit one more in.

Writing will be much faster with 2 vdevs. On my single vdev pool I can't saturate gigabit due to the ~5 second pause where the pool computes parity and writes the blocks.



http://docs.oracle.com/cd/E18752_01/html/819-5461/gazss.html#gcfgv


https://blogs.oracle.com/roch/entry/tuning_zfs_recordsize


I tend to agree with you, more VDEVS=faster speeds.
However, wouldn't 11 or 12 drives in a VDEV lead to the blocksize problem that patrickdk underlined previously?
 
You should avoid 23 disks in one vdev. Better to have two vdevs, each having 11 disks - that is perfect alignment (2,4,8 disks). 11 disks = 8 + 3 which is perfect.

The problem with one big vdev, is that performance will degrade heavily. On the FreeBSD forum someone tried a large server with one big vdev and got horrible performance results. For isntance, resilvering did never complete (while the server was being used at the same time). So they did 3 or 4 raidz2 vdevs and performance skyrocketed. So, resilver will take very long time, possibly weeks(?). If that is ok, continue with one huge vdev. Otherwise, split it up into at least two vdevs. and use the last 23rd disk as a hot spare.
 
You should avoid 23 disks in one vdev. Better to have two vdevs, each having 11 disks - that is perfect alignment (2,4,8 disks). 11 disks = 8 + 3 which is perfect.

The problem with one big vdev, is that performance will degrade heavily. On the FreeBSD forum someone tried a large server with one big vdev and got horrible performance results. For isntance, resilvering did never complete (while the server was being used at the same time). So they did 3 or 4 raidz2 vdevs and performance skyrocketed. So, resilver will take very long time, possibly weeks(?). If that is ok, continue with one huge vdev. Otherwise, split it up into at least two vdevs. and use the last 23rd disk as a hot spare.

VERY good suggestion!
For some reason one of the hotswap bays doesn't seem to be working so I am limited to using 23 disks (I will do more testing to see why this happens but even if I can't fix it its not a big deal).

2 x 11 disks VDEVS plus 1 hotspare would be perfect for me.
I'll do the reconfiguration today and see how it works.
But first, I will run some tests on the current VDEV and I will post the results.
The current pool is mapped to a Windows machine via 4Gb FC cards in Point-to-Point mode.
I will come back with some updates later tonight most likely (if anyone curious about performance results).
 
Last edited:
I'de be willing to be you don't have 4TB (4096GB), you have 3.90625TB (4000GB) drives which would net you ~70.9TB before filesystem losses, so ~66GB is right on the money.

All the stuff about having 2 VDEV's is correct too.
 
He has 4TB or 3,63TiB drives.

20*3,63=72,6

Remove 1/64th for ZFS => 71,4TiB for 80TB RAW + 12TB parity RAW.

About 5,7TiB are missing and that's not 20% or even 10% so I'm not convinced by the previous explanations.

I'm particularly interested since as soon as I receive a 8087-8088 cable I'm planning a 16 drives RAIDZ3, which according to previous explanations would also lead to a 20% loss in capacity, unacceptable. I might go for 19 drives if it's confirmed. I'm trying to test this with vmware but can't manage to make fake 4K drives with it.
 
I'de be willing to be you don't have 4TB (4096GB), you have 3.90625TB (4000GB) drives which would net you ~70.9TB before filesystem losses, so ~66GB is right on the money.

All the stuff about having 2 VDEV's is correct too.

I don't believe your calculations are quite correct... a 4TB drive is actually smaller:
Roughly, there is about 9% difference between the mathematical TB and computerized TB (TiB).

4TB drive is ~4,000,000,000,000 bytes which is actually ~3.63TiB (4,000,000,000,000 / 1024 / 1024 / 1024 / 1024)

20 x 4TB drives is 20 x 3.63TiB = 72.6TiB (before filesystem losses).
Filesystem losses account for 1/64th of the total space, so the remaining usable space is 72.6 * 63/64 = 71.46TiB

In any case, from 71.46 to ~66 is quite a big loss of space... I don't think filesystem losses account for ~5TB of the missing space, I would be more likely to believe the explanation provided by patrickdk in regards to the block size.
 
Last edited:
Giga is 9 zeros, we're in the 12 zeros world now (and more for us storage junkies).
 
Anyways... after I get the performance numbers with current setup, I will 'kill' the current pool and create different pool sizes with various drive numbers, and record the usable space in each case, this way we can get a better idea of what is going on.
 
Giga is 9 zeros, we're in the 12 zeros world now (and more for us storage junkies).

You are absolutely correct, I lost 3 zeroes on the way :)
I remade my calculations in the same post, and they seem to be the same as yours.
 
Could you first remake the pool and just check that you didn't check :)p) the box in napp-it ? Or make the pool manually (a bit of a bother though with 23 drives).
 
I had a similar issue with mysterious missing space on an empty pool. I just assumed it was a feature and not a bug! I think it was missing 600 GB at the time (2xvdevs, RAIDZ2, 10x1 TB + 10x2 TB hdds).

I compared ASHIFT = 9 vs 12 and the 12 definitely was missing a pile of space.
 
I've tested the current pool and I've got the following results for seqential speeds:
READ: ~360MB/sec
WRITE: ~105MB/sec

The connection is established by 4Gb FC card Point-to-Point, the client is a Windows VM, the FC card is pass-through to the Windows VM.

From what I've read so far, it seems that the ASHIFT value would cause the issue I am experiencing.
I will re-create the pool with ASHIFT=9 and see what happens. I will also re-test the performance and post the results.
 
I can't seem to find a way to format the drives with ASHIFT=9 under OpenIndiana... can anyone help please?

I've found a guide for FreeBSD, the command "gnop" is used there, but this command is not available in Openindiana....
 
I think these are 4K disks.. so I need to find a way to force the OS to format them as 512b disks..
 
I found this guide which explains how to force the 512b block size, however the guide is assuming driver "sd" is used for hard drives:
http://wiki.illumos.org/display/illumos/ZFS+and+Advanced+Format+disks

In my case, the format command shows that my disks are using scsi_vhci driver:

root@openindiana:~# format
Searching for disks...done

AVAILABLE DISK SELECTIONS:
0. c5t0d0 <ATA-MZ-5EA1000-0D3-7D3Q cyl 2607 alt 2 hd 255 sec 63>
/pci@0,0/pci1043,8362@1f,2/disk@0,0
1. c10t5000CCA22EC00A7Cd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec00a7c
2. c10t5000CCA22EC00F4Fd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec00f4f
3. c10t5000CCA22EC01A8Dd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec01a8d
4. c10t5000CCA22EC02D19d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec02d19
5. c10t5000CCA22EC04C15d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec04c15
6. c10t5000CCA22EC04C26d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec04c26
7. c10t5000CCA22EC04C35d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec04c35
8. c10t5000CCA22EC04C48d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec04c48
9. c10t5000CCA22EC04C87d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec04c87
10. c10t5000CCA22EC008ADd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec008ad
11. c10t5000CCA22EC021D9d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec021d9
12. c10t5000CCA22EC022ACd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec022ac
13. c10t5000CCA22EC022E7d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec022e7
14. c10t5000CCA22EC032E8d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec032e8
15. c10t5000CCA22EC032F8d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec032f8
16. c10t5000CCA22EC0210Ed0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec0210e
17. c10t5000CCA22EC0221Bd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec0221b
18. c10t5000CCA22EC0222Bd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec0222b
19. c10t5000CCA22EC0330Bd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec0330b
20. c10t5000CCA22EC0340Bd0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec0340b
21. c10t5000CCA22EC02208d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec02208
22. c10t5000CCA22EC02234d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec02234
23. c10t5000CCA22EC02287d0 <ATA-Hitachi HDS5C404-A250-3.64TB>
/scsi_vhci/disk@g5000cca22ec02287


so it made sense (to me, at least) to try to modify the file /kernel/drv/scsi_vhci.conf

So I put at the end of the file the following:
device-type-scsi-options-list =
"Hitachi HDS5C404", "physical-block-size:512";

then I rebooted the system, created a new pool, but it still has the same ashift=12, not ashift=9 as I wanted :(
root@openindiana:~# zdb Pool-Storage1 | grep ashift
ashift: 12
ashift: 12

So probably what I put in the scsi_vhci.conf didn't work... any suggestions?
 
Last edited:
ashift=12 seems like a known issue developers are working to fix it (or fixed already..?)
https://github.com/behlendorf/zfs/commit/b28e57cb82c5d5a992b90c56f67dd7dbf9b6f296


this option is only available on Linux not Solarish systems and introduced to force ashift=12 on older disks.
Illumos decided to include this as a disk property in sd.conf.

But I would never decide to use 4k disks with ashift=9.
You should be happy to have ashift=12 vdevs as they allow a disk - replace with actual 4 k disks.

Even with old 512B disks, I would prefer ashift=12 or you may come in to a situation that you cannot buy
old 512B disks and you cannot replace a faulted disk with a new 4k one.

For sure, 4k disks are not as space efficient like disks with a smaller physical block size.
If you decide to save a file with one character only, you have lost 4k. If you build a Raid you have also a lost
depending on number of disks or used blocksize.

My suggestion: be happy with 4k and ashift=12
It is even faster and you have no choice, even if you have lost some space for this. Any other option is worse.
 
I've forced ashift = 9 on a couple of my servers that have 4k disks. I did some tests and ashift = 12 would have cost me around 15 TB more slack space on those particular servers given the pool configuration and expected file size distribution, so I chose to sacrifice some speed instead.
 
Back
Top