• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

ZFS Performance - Is it where it should be?

Ripley

Limp Gawd
Joined
Nov 4, 2004
Messages
240
I have finished construction of my ZFS NAS. Before I start transferring large amounts of data to it I wanted to make sure that performance was appropriate for the hardware. Here's what I'm running:


  • Supermicro 846 24bay 4U Case
  • Dual 920W redundant Power suplies
  • Supermicro X9SRL-F Motherboard
  • Intel Xeon E5-2603 CPU
  • 16GBs ECC RAM
  • (3) IBM M1015 HBAs flashed to IT firmware
  • 120GB 2.5" Toshiba OS Drive
  • 12 1TB WD RE4 D1003FBYX
  • 12 2TB WD RE4 D2003FYYS


I ran a bunch of different tests using dd and bonnie++. I was hoping that you talented people with lots more ZFS experience would be able to tell me if I'm on track. I tested 3 different drive configurations a 2 drive mirror, 2 2 drive mirrors and a 12 drive RAIDZ2 (this is the planned final configuration). Someone suggested to me that setting atime=off and relatime=on could improve performance so I tested both as well. Here's what I'd like to know:

  • Did I run reasonable tests for performance?
  • If so does the performance match what I should be able to get?
  • Is there anything I should do to improve performance?

Code:
{atime=on}
[2 Drive mirror Write]
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 211.715 s, 116 MB/s

real    3m40.995s
user    0m10.357s
sys     1m53.347s

[2 Drive mirror Read]
root@muninn:/# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 291.195 s, 84.4 MB/s

real    4m51.198s
user    0m52.392s
sys     3m58.737s

[2 Drive mirror Bonnie++]
root@muninn:/# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    82  99 105071  32 59240  17   193  99 176016  19 323.3  15
Latency               107ms   14036us    2469ms   59348us     419ms     309ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024 12162  99 327812  99 12349  99 11952  99 392036  99 11692  99
Latency               176ms     661us     923us     159ms     213us   19064us
1.97,1.97,muninn,1,1421967113,32G,,82,99,105071,32,59240,17,193,99,176016,19,323.3,15,1024,,,,,12162,99,327812,99,12349,99,11952,99,392036,99,11692,99,107ms,14036us,2469ms,59348us,419ms,309ms,176ms,661us,923us,159ms,213us,19064us

[2 2 Drive mirrors Write]
root@muninn:~# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 107.992 s, 228 MB/s

real    2m0.339s
user    0m7.691s
sys     1m28.301s

[2 2 Drive mirrors Read]
root@muninn:~# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 286.534 s, 85.8 MB/s

real    4m46.537s
user    0m52.275s
sys     3m54.164s

[2 2 Drive mirrors Bonnie++]
root@muninn:/# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    62  99 205804  56 120419  40   196  99 342134  34 535.4  21
Latency               136ms   15372us    2065ms   73134us     258ms     188ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024 12365  99 335887  99 12404  99 12215  99 398375  99 12277  99
Latency               179ms     657us    4510us     183ms      23us    4280us
1.97,1.97,muninn,1,1421951305,32G,,62,99,205804,56,120419,40,196,99,342134,34,535.4,21,1024,,,,,12365,99,335887,99,12404,99,12215,99,398375,99,12277,99,136ms,15372us,2065ms,73134us,258ms,188ms,179ms,657us,4510us,183ms,23us,4280us

[12 drive RAIDZ2 Write]
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 87.3114 s, 281 MB/s

real    1m27.694s
user    0m7.208s
sys     1m20.053s

[12 drive RAIDZ2 Read]
root@muninn:/# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 294.115 s, 83.6 MB/s

real    4m54.118s
user    0m52.463s
sys     4m1.529s

[12 drive RAIDZ2 Bonnie++]
root@muninn:/# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    82  99 510005  98 226927  69   200  99 878879  99 248.3  14
Latency               105ms    8099us    1837ms   71367us   61029us     153ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024 11565  99 333760  99 12291  99 11650  99 394102  99 10196  99
Latency               233ms     833us    4132us     278ms     113us    8251us
1.97,1.97,muninn,1,1421962472,32G,,82,99,510005,98,226927,69,200,99,878879,99,248.3,14,1024,,,,,11565,99,333760,99,12291,99,11650,99,394102,99,10196,99,105ms,8099us,1837ms,71367us,61029us,153ms,233ms,833us,4132us,278ms,113us,8251us

{atime off, relatime on}
[2 Drive mirror Write]
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 215.51 s, 114 MB/s

real    3m59.471s
user    0m10.014s
sys     1m55.938s

[2 Drive mirror Read]
root@muninn:/# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 287.072 s, 85.6 MB/s

real    4m47.075s
user    0m52.511s
sys     3m54.466s

[2 Drive mirror Bonnie++]
root@muninn:/# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    81  99 104857  33 60546  18   195  99 177635  19 326.5  15
Latency               105ms   14650us    2236ms   70346us     503ms     280ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024 11838  99 327691  99 12234  99 11742  99 390507  99 11870  99
Latency               193ms     496us    1015us     180ms      30us    4225us
1.97,1.97,muninn,1,1421965202,32G,,81,99,104857,33,60546,18,195,99,177635,19,326.5,15,1024,,,,,11838,99,327691,99,12234,99,11742,99,390507,99,11870,99,105ms,14650us,2236ms,70346us,503ms,280ms,193ms,496us,1015us,180ms,30us,4225us

[2 2 Drive mirrors Write]
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 109.423 s, 225 MB/s

real    2m3.079s
user    0m7.765s
sys     1m38.359s

[2 2 Drive mirrors Read]
root@muninn:/# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 289.35 s, 84.9 MB/s

real    4m49.353s
user    0m53.412s
sys     3m55.851s

[2 2 Drive mirrors Bonnie++]
root@muninn:/# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    64  99 206422  57 119291  39   200  99 339873  33 519.5  20
Latency               133ms   13539us    2290ms   68845us     268ms     217ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024  5674  77 168728  52  5252  99 12400  99 394056  99 12481  99
Latency               204ms   38855us   25770us     171ms      25us   26036us
1.97,1.97,muninn,1,1421952667,32G,,64,99,206422,57,119291,39,200,99,339873,33,519.5,20,1024,,,,,5674,77,168728,52,5252,99,12400,99,394056,99,12481,99,133ms,13539us,2290ms,68845us,268ms,217ms,204ms,38855us,25770us,171ms,25us,26036us

[12 Drive RAIDZ2 Write]
root@muninn:/tank1# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 87.6158 s, 280 MB/s

real    1m28.047s
user    0m7.168s
sys     1m20.622s

[12 Drive RAIDZ2 Read]
root@muninn:/tank1# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 301.988 s, 81.4 MB/s

real    5m1.991s
user    0m52.308s
sys     4m2.378s

[12 Drive RAIDZ2 Bonnie++]
root@muninn:/tank1# bonnie++ -d /tank1 -r 16384 -u ripley -n 1024
Using uid:1000, gid:1000.
Writing a byte at a time...done
Writing intelligently...done
Rewriting...done
Reading a byte at a time...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn          32G    80  99 505593  98 220501  67   190  99 880084  99 238.8  13
Latency               106ms    6547us     136ms   71490us   16216us     142ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
               1024 11377  99 329742  99 12094  99 11519  99 390066  99 10420  99
Latency               219ms     471us    4182us     246ms      21us    5603us
1.97,1.97,muninn,1,1421957471,32G,,80,99,505593,98,220501,67,190,99,880084,99,238.8,13,1024,,,,,11377,99,329742,99,12094,99,11519,99,390066,99,10420,99,106ms,6547us,136ms,71490us,16216us,142ms,219ms,471us,4182us,246ms,21us,5603us
 
Here is a 14 disk consisting of two 7 disk raidz2s.

Code:
datastore4 shell-scripts # uname -a
Linux datastore4 3.17.8-gentoo-datastore4-zfs-20141220 #1 SMP Fri Jan 9 14:00:24 EST 2015 x86_64 Intel(R) Xeon(R) CPU E31230 @ 3.20GHz GenuineIntel GNU/Linux

datastore4 Images # cat /sys/module/{spl,zfs}/version
0.6.3-54_g03a7835
0.6.3-170_gd958324

datastore4 Images # free -m
             total       used       free     shared    buffers     cached
Mem:          7947       7650        296          6          0        243
-/+ buffers/cache:       7407        539
Swap:        16383         20      16363

Code:
datastore4 Images # zpool status
  pool: zfs_data_0
 state: ONLINE
status: Some supported features are not enabled on the pool. The pool can
        still be used, but some features are unavailable.
action: Enable all features using 'zpool upgrade'. Once this is done,
        the pool may no longer be accessible by software that does not support
        the features. See zpool-features(5) for details.
  scan: scrub repaired 0 in 7h53m with 0 errors on Sat Jan 17 12:13:28 2015
config:

        NAME             STATE     READ WRITE CKSUM
        zfs_data_0       ONLINE       0     0     0
          raidz2-0       ONLINE       0     0     0
            a0_d0-part3  ONLINE       0     0     0
            a0_d1-part3  ONLINE       0     0     0
            a0_d2-part3  ONLINE       0     0     0
            a0_d3-part3  ONLINE       0     0     0
            a0_d4-part3  ONLINE       0     0     0
            a0_d5-part3  ONLINE       0     0     0
            a0_d6-part3  ONLINE       0     0     0
          raidz2-1       ONLINE       0     0     0
            a1_d0-part3  ONLINE       0     0     0
            a1_d1-part3  ONLINE       0     0     0
            a1_d2-part3  ONLINE       0     0     0
            a1_d3-part3  ONLINE       0     0     0
            a1_d4-part3  ONLINE       0     0     0
            a1_d5-part3  ONLINE       0     0     0
            a1_d6-part3  ONLINE       0     0     0

errors: No known data errors

Code:
datastore4 Images # dd if=/dev/zero of=/zfs_data_0/test.dd bs=4k count=6000000 && sync
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 50.6417 s, 485 MB/s

datastore4 Images # time sh -c "dd if=/zfs_data_0/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 159.594 s, 154 MB/s

real    2m39.617s
user    0m5.410s
sys     2m16.550s

disks are a mix of 7200 RPM 2TB Seagate Constellation ES.3, Hitachi 7K200 DeskStar 2TB and Toshiba 3.5" HDD MK.002TSKB 2TB.
 
Last edited:
I reconfigured my drives to more closely match your setup:

Code:
root@muninn:/# zpool status tank1
  pool: tank1
 state: ONLINE
  scan: none requested
config:

        NAME                        STATE     READ WRITE CKSUM
        tank1                       ONLINE       0     0     0
          raidz2-0                  ONLINE       0     0     0
            wwn-0x50014ee00228aef7  ONLINE       0     0     0
            wwn-0x50014ee00228c4a7  ONLINE       0     0     0
            wwn-0x50014ee0577e156a  ONLINE       0     0     0
            wwn-0x50014ee0577e15a7  ONLINE       0     0     0
            wwn-0x50014ee0577e17f9  ONLINE       0     0     0
            wwn-0x50014ee205367b27  ONLINE       0     0     0
          raidz2-1                  ONLINE       0     0     0
            wwn-0x50014ee0acd39e8a  ONLINE       0     0     0
            wwn-0x50014ee0acd3abbe  ONLINE       0     0     0
            wwn-0x50014ee0ad432df1  ONLINE       0     0     0
            wwn-0x50014ee25a8bd248  ONLINE       0     0     0
            wwn-0x50014ee25a8c154b  ONLINE       0     0     0
            wwn-0x50014ee2afe1e347  ONLINE       0     0     0

errors: No known data errors

I reran the tests:

Code:
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=4k count=6000000 && sync"
6000000+0 records in
6000000+0 records out
24576000000 bytes (25 GB) copied, 89.3532 s, 275 MB/s

real    1m29.615s
user    0m7.226s
sys     1m22.080s

root@muninn:/# time sh -c "dd if=/tank1/test.dd of=/dev/null"
48000000+0 records in
48000000+0 records out
24576000000 bytes (25 GB) copied, 303.308 s, 81.0 MB/s

real    5m3.311s
user    0m52.532s
sys     4m10.685s


It looks like I'm getting roughly half the performance you are. All of my disks are the same but they are spread across HBAs. How are your drives connected? Any other ideas why I'm getting some much less speed?
 
I have two Intel branded 8 port LSI SAS1068E PCIe cards. I doubt that I have the drives evenly split.
 
I may have answered my own question. Looks like all my drives are 3Gb/s SATA and yours are 6Gb/s Sata.
 
The 2TB Seagate Constellation ES.3 net 198 MB/s in the outer tracks while the Hitachi and Toshiba are around 155MB/s. I believe this is largely due to 1TB platters versus 666GB.

Forgot to mention the array has 14TB+ of data on it that is not evenly distributed to the vdevs (since I created the first 7 drive vdev added data then added the 2nd 7 drive vdev)

Code:
datastore4 Images # zpool list
NAME         SIZE  ALLOC   FREE  EXPANDSZ   FRAG    CAP  DEDUP  HEALTH  ALTROOT
zfs_data_0  25.2T  14.8T  10.4T         -      -    58%  1.00x  ONLINE  -
 
I may have answered my own question. Looks like all my drives are 3Gb/s SATA and yours are 6Gb/s Sata.
1068e is a SAS1 controller (hence SATA2) so they'll be running at 3Gb/s, most spinning rust won't run into the 3Gb/s limit very easily. Either way your results do look a little on the low side.

Which OS?
 
1068e is a SAS1 controller (hence SATA2) so they'll be running at 3Gb/s, most spinning rust won't run into the 3Gb/s limit very easily. Either way your results do look a little on the low side.

Which OS?

Ubuntu 14.04 LTS with ZoL 0.6.3
 
Your tests are flawed, but as they are, they look to be expected.

You are writing 4k from dd to the os at a time.
Then zfs groups those 4k into 128k writes, and drops them on the disk, so your getting as good as expected write performance for your hardware/disks.

The reads suck, cause your reading 1byte at a time, causing 4k more calls to the os for data than your writes, this is why your reads suck. Really though, if your going bother even using dd for a test at all, atleast use like 1MB or so block sizes.

And don't test with atime on also, set atime off, this is causing an extra write for EVERY read you do, this is also killing performance.
 
Your tests are flawed, but as they are, they look to be expected.

You are writing 4k from dd to the os at a time.
Then zfs groups those 4k into 128k writes, and drops them on the disk, so your getting as good as expected write performance for your hardware/disks.

The reads suck, cause your reading 1byte at a time, causing 4k more calls to the os for data than your writes, this is why your reads suck. Really though, if your going bother even using dd for a test at all, atleast use like 1MB or so block sizes.

And don't test with atime on also, set atime off, this is causing an extra write for EVERY read you do, this is also killing performance.

If there are better tests I'm happy to run them, what would you suggest? Also I didn't see much difference in my testing with atime on or off. The numbers are very close. Lastly, do you have any idea why drescherjm is getting almost double the speed with a similar setup?

I also ran the test with 1M sized blocks with a 12 drive RAIDZ2 pool.

Code:
root@muninn:/# zfs set atime=off tank1
root@muninn:/# time sh -c "dd if=/dev/zero of=/tank1/test.dd bs=1M count=25000 && sync"
25000+0 records in
25000+0 records out
26214400000 bytes (26 GB) copied, 30.8422 s, 850 MB/s

real    0m32.315s
user    0m0.093s
sys     0m14.064s
root@muninn:/tank1# time sh -c "dd if=/tank1/test.dd of=/dev/null bs=1M"
25000+0 records in
25000+0 records out
26214400000 bytes (26 GB) copied, 23.095 s, 1.1 GB/s

real    0m23.098s
user    0m0.067s
sys     0m14.464s

You are definitely right that changing to 1M block size makes a big difference, I tried it with 512K as well and got almost identical results. Still almost no difference between atime on or off though.
 
Last edited:
Maybe it's only touching atime once per file open, not exactly sure. I know in my tests it had a huge difference, but I wasn't using dd to test.

I also have 0 need for atime, most people dont need it though.
 
Maybe it's only touching atime once per file open, not exactly sure. I know in my tests it had a huge difference, but I wasn't using dd to test.

I also have 0 need for atime, most people dont need it though.

What kind of testing do you use? I'd like to try some other tests if you have suggestions.
 
looks like you're cpu bound, see how bonnie is using like 99% cpu? bonnie isn't that efficient....

erk dunno where i got that from, maybe someone elses results :)

oh i see, you should run with -f0 for fast mode, the per char is pointless. but you were 99% cpu at raidz2 with 12 disks;

Version 1.97 ------Sequential Output------ --Sequential Input- --Random-
Concurrency 1 -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP /sec %CP
muninn 32G 82 99 510005 98 226927 69 200 99 878879 99 248.3 14
Latency 105ms 8099us 1837ms 71367us 61029us 153ms

Where it says 99 after 878879.

i hit bonnie cpu bottleneck on i7-4770 when enabling lz4 compression, but can max out disks without compression. but it looks like your cpu is more like 1.8ghz. so you may need a couple of processes or to focus more on tests like dd/cp/etc. samba is also a huge cpu hog, which you are likely to run into at 10 gigabit ethernet but not gigabit ethernet.

my guess is with bonnie giving slow speeds it's probably because it uses a block size that's too small, and that should be trivially simple to fix.

looks like no source modification necessary:
-s the size of the file(s) for IO performance measures in megabytes. If the size is greater than 1G then multiple files will be used to store the data, and each file will be up to 1G in
size. This parameter may include the chunk size seperated from the size by a colon. The chunk-size is measured in bytes and must be a power of two from 256 to 1048576, the default
is 8192. NB You can specify the size in giga-bytes or the chunk-size in kilo-bytes if you add g or k to the end of the number respectively.

so try -s 32768:262144

for comparisions sake, testing with less ram for quicker tests, and testing cpu performance of bonnie mostly:

Code:
% time ./bonnie++ -f0 -r4096M
Writing intelligently...done
Rewriting...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine        Size K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
amethyst         8G           433497  42 337157  31           7037133  99 +++++ +++
Latency                         697ms     237ms               328us     379us
Version  1.97       ------Sequential Create------ --------Random Create--------
amethyst            -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
                 16 +++++ +++ +++++ +++ +++++ +++ +++++ +++ +++++ +++ +++++ +++
Latency               309us     506us     614us     335us      41us     262us
1.97,1.97,amethyst,1,1422156650,8G,,,,433497,42,337157,31,,,7037133,99,+++++,+++,16,,,,,+++++,+++,+++++,+++,+++++,+++,+++++,+++,+++++,+++,+++++,+++,,697ms,237ms,,328us,379us,309us,506us,614us,335us,41us,262us
./bonnie++ -f0 -r4096M  0.34s user 17.52s system 36% cpu 49.101 total


 % time ./bonnie++ -f0 -r4096M -s8192:262144
Writing intelligently...done
Rewriting...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine   Size:chnk K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
amethyst    8G:256k           440625  18 559972  17           10141751  99 16046 701
Latency                         402ms     255ms                76us     706us
Version  1.97       ------Sequential Create------ --------Random Create--------
amethyst            -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
                 16 +++++ +++ +++++ +++ +++++ +++ +++++ +++ +++++ +++ +++++ +++
Latency               137us     230us     281us     185us      18us     195us
1.97,1.97,amethyst,1,1422157225,8G,256k,,,440625,18,559972,17,,,10141751,99,16046,701,16,,,,,+++++,+++,+++++,+++,+++++,+++,+++++,+++,+++++,+++,+++++,+++,,402ms,255ms,,76us,706us,137us,230us,281us,185us,18us,195us
./bonnie++ -f0 -r4096M -s8192:262144  0.05s user 8.72s system 22% cpu 39.576 total
 
Last edited:
here's some comparison tests on zfs raid array on 4 cheap ssd's with mirrored vdevs on i7-4770. it looks like 23% cpu for 64k and 20% cpu for 256k for reads. around 64 to 256k should be good. and 8k is terrible. performance looks pretty close between them.

256k block size:
Code:
# time bonnie++ -f0 -u root -s 32768:262144
Using uid:0, gid:0.
Writing intelligently...done
Rewriting...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine   Size:chnk K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
amethyst   32G:256k           441727   8 291082   8           1396770  20  4238 112
Latency                         494ms     837ms             81913us   97308us
Version  1.97       ------Sequential Create------ --------Random Create--------
amethyst            -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
                 16 +++++ +++ +++++ +++ +++++ +++ 32417  96 +++++ +++ +++++ +++
Latency             15105us     355us     283us   30482us      17us      77us
1.97,1.97,amethyst,1,1422157723,32G,256k,,,441727,8,291082,8,,,1396770,20,4238,112,16,,,,,+++++,+++,+++++,+++,+++++,+++,32417,96,+++++,+++,+++++,+++,,494ms,837ms,,81913us,97308us,15105us,355us,283us,30482us,17us,77us
bonnie++ -f0 -u root -s 32768:262144  0.19s user 22.92s system 10% cpu 3:42.17 total

64k block size
Code:
 # time bonnie++ -f0 -u root -s 32768:65536 
Using uid:0, gid:0.
Writing intelligently...done
Rewriting...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine   Size:chnk K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
amethyst    32G:64k           438411  14 290279  10           1418218  23  7379  88
Latency                         789ms     401ms             55848us   62908us
Version  1.97       ------Sequential Create------ --------Random Create--------
amethyst            -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
                 16 +++++ +++ +++++ +++ +++++ +++ 23456  96 +++++ +++  7006  99
Latency             11549us     249us     269us   37978us      12us     701us
1.97,1.97,amethyst,1,1422157299,32G,64k,,,438411,14,290279,10,,,1418218,23,7379,88,16,,,,,+++++,+++,+++++,+++,+++++,+++,23456,96,+++++,+++,7006,99,,789ms,401ms,,55848us,62908us,11549us,249us,269us,37978us,12us,701us
bonnie++ -f0 -u root -s 32768:65536  0.40s user 32.30s system 14% cpu 3:43.99 total
 
I believe I have found my final configuration.

Code:
root@muninn:/# zpool status
  pool: tank2
 state: ONLINE
  scan: scrub repaired 0 in 0h0m with 0 errors on Wed Jan 21 18:01:22 2015
config:

        NAME                        STATE     READ WRITE CKSUM
        tank2                       ONLINE       0     0     0
          raidz2-0                  ONLINE       0     0     0
            wwn-0x50014ee0024e013d  ONLINE       0     0     0
            wwn-0x50014ee057a390c4  ONLINE       0     0     0
            wwn-0x50014ee057a3fb1e  ONLINE       0     0     0
            wwn-0x50014ee057a41324  ONLINE       0     0     0
            wwn-0x50014ee057a42fcb  ONLINE       0     0     0
            wwn-0x50014ee057a5275d  ONLINE       0     0     0
            wwn-0x50014ee0ace3aaac  ONLINE       0     0     0
            wwn-0x50014ee0acf9ab09  ONLINE       0     0     0
            wwn-0x50014ee0acf9bb21  ONLINE       0     0     0
            wwn-0x50014ee0acfa2eeb  ONLINE       0     0     0
            wwn-0x50014ee65569775a  ONLINE       0     0     0
            wwn-0x50014ee6aabec49a  ONLINE       0     0     0
          raidz2-1                  ONLINE       0     0     0
            wwn-0x50014ee00228aef7  ONLINE       0     0     0
            wwn-0x50014ee00228c4a7  ONLINE       0     0     0
            wwn-0x50014ee0577e156a  ONLINE       0     0     0
            wwn-0x50014ee0577e15a7  ONLINE       0     0     0
            wwn-0x50014ee0577e17f9  ONLINE       0     0     0
            wwn-0x50014ee205367b27  ONLINE       0     0     0
            wwn-0x50014ee0acd39e8a  ONLINE       0     0     0
            wwn-0x50014ee0acd3abbe  ONLINE       0     0     0
            wwn-0x50014ee0ad432df1  ONLINE       0     0     0
            wwn-0x50014ee25a8bd248  ONLINE       0     0     0
            wwn-0x50014ee25a8c154b  ONLINE       0     0     0
            wwn-0x50014ee2afe1e347  ONLINE       0     0     0

errors: No known data errors

This configuration produced these results:

Code:
root@muninn:/tank2# bonnie++ -d /tank2/media -u ripley -s32G:512k -f0 -n 500
Using uid:1000, gid:1000.
Writing intelligently...done
Rewriting...done
Reading intelligently...done
start 'em...done...done...done...done...done...
Create files in sequential order...done.
Stat files in sequential order...done.
Delete files in sequential order...done.
Create files in random order...done.
Stat files in random order...done.
Delete files in random order...done.
Version  1.97       ------Sequential Output------ --Sequential Input- --Random-
Concurrency   1     -Per Chr- --Block-- -Rewrite- -Per Chr- --Block-- --Seeks--
Machine   Size:chnk K/sec %CP K/sec %CP K/sec %CP K/sec %CP K/sec %CP  /sec %CP
muninn     32G:512k           889964  39 600642  60           1459997  91 281.7  47
Latency                         518ms     157ms             57003us     126ms
Version  1.97       ------Sequential Create------ --------Random Create--------
muninn              -Create-- --Read--- -Delete-- -Create-- --Read--- -Delete--
              files  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP  /sec %CP
                500  7364  69 221862  91  9380  77  9687  85 312256  99 11090  98
Latency               223ms   10548us     113ms     209ms      20us   24369us
1.97,1.97,muninn,1,1422374513,32G,512k,,,889964,39,600642,60,,,1459997,91,281.7,47,500,,,,,7364,69,221862,91,9380,77,9687,85,312256,99,11090,98,,518ms,157ms,,57003us,126ms,223ms,10548us,113ms,209ms,20us,24369us

These numbers are better generally then they other setups I tried. I still though don't feel like I have a great handle on what the numbers mean or that I know much more than I did when I started. After trying multiple different benchmark programs (dd, bonnie++, LAN Speed Test, IOmeter, Phoronix Test Suite) I've generated a lot of numbers but I'm still not quite sure how to interpret them. This is one of the few things that I've ever spent this much time on and not come away feeling like I had a better handle on the process. I seem to be be getting reasonable performance. I can saturate the network connection during a file copy. But its hard to come to an absolute performance number. This has been an interesting if frustrating endeavor. At this point I just going to start using it and not spend anymore time benchmarking. I appreciate all your comments, testing and advice.
 
Back
Top