• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

ZFS Configuration Help

TechIsCool

Limp Gawd
Joined
Aug 2, 2009
Messages
167
Alright everyone I have been using zfs for my iscsi backup location. I am using Veeam to backup my virtual machines and I have had dedup enabled. Now looking back this might not have been the best idea but it is what is enabled right now. I can still write data to the system fine but I am nearing my capacity and I know for a fact that my dedupe table is not in ram/l2arc. What I am looking for is what I should do to get myself back into a place that I still have space for my backups and how I should handle year over year adding hard drives if I can't dedup my data.

I had thought using my math seen below that this was a perfect use case for dedup since almost all my data will be duplicated about 40 times before being delete.

So more information
Code:
Supermicro X9SCM-F - I think
2x M1015
12x  Hitachi HDS5C303
Intel Xeon E3-1230 Sandy Bridge 3.2GHz LGA 1155 80W Quad-Core Server Processor
2x Kingston 8GB (2 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333 (PC3 10600) ECC Unbuffered Server Memory Model KVR1333D3E9SK2/8G
SUPERMICRO SuperChassis CSE-846TQ-R900B - I think.
Code:
Name					Data	Disk Growth	Dedupe		Full Backups	Incr. Backups	Modified Data		Backup Size		Replica Size
Exchange				120 GB	0.25%		50%		41		31		10.0%			2,646 GB		306 GB
Flat File Database			12 GB	0.25%		50%		40		18		5.0%			245 GB			17 GB
User/Files				300 GB	0.25%		50%		41		31		5.0%			6,383 GB		533 GB
Sharepoint				26 GB	0.10%		50%		40		18		10.0%			543 GB			49 GB
Domain Controller			18 GB	0.25%		50%		41		11		5.0%			374 GB			23 GB
SQL Server				25 GB	0.50%		50%		40		18		10.0%			523 GB			48 GB
Website					6 GB	0.10%		50%		40		18		5.0%			123 GB			9 GB
Terminal Server				23 GB	0.25%		50%		40		18		5.0%			470 GB			33 GB
2nd Domain Controller			20 GB	0.25%		50%		40		18		5.0%			409 GB			29 GB
Total					550 GB					363		181					11,716 GB		1,047 GB

Code:
 data 	 used 	 		12.8T 	 - 
 data 	 available 	 	3.60T 	 - 
 data 	 referenced 	 	5.24T 	 - 
 data 	 compressratio 	 	1.00x 	 - 
 data 	 mounted 	 	yes 	 - 
 data 	 quota 	 		none 	 default 
 data 	 reservation 	 	276G 	 local 
 data 	 recordsize 	 	128K 	 default 
 data 	 mountpoint 	 	/data 	 default 
 data 	 sharenfs 	 	off 	 default 
 data 	 checksum 	 	on 	 default 
 data 	 compression 	 	off 	 default 
 data 	 atime 	 		on 	 default 
 data 	 devices 	 	on 	 default 
 data 	 exec 	 		on 	 default 
 data 	 setuid 	 	on 	 default 
 data 	 readonly 	 	off 	 default 
 data 	 zoned 	 		off 	 default 
 data 	 snapdir 	 	hidden 	 default 
 data 	 aclmode 	 	discard 	 default 
 data 	 aclinherit 	 	restricted 	 default 
 data 	 canmount 	 	on 	 default 
 data 	 xattr 	 		on 	 default 
 data 	 copies 	 	1 	 default 
 data 	 version 	 	5 	 - 
 data 	 utf8only 	 	off 	 - 
 data 	 normalization 	 	none 	 - 
 data 	 casesensitivity 	sensitive 	 - 
 data 	 vscan 	 		off 	 default 
 data 	 nbmand 	 	off 	 default 
 data 	 sharesmb 	 	off 	 default 
 data 	 refquota 	 	none 	 default 
 data 	 refreservation 	276G 	 local 
 data 	 primarycache 	 	all 	 default 
 data 	 secondarycache 	all 	 default 
 data 	 usedbysnapshots 	0 	 - 
 data 	 usedbydataset 	 	5.24T 	 - 
 data 	 usedbychildren 	7.58T 	 - 
 data 	 usedbyrefreservation 	0 	 - 
 data 	 logbias 	 	latency  default 
 data 	 dedup 	 		off 	 default 
 data 	 mlslabel 	 	none 	 default 
 data 	 sync 	 standard 	default 
 data 	 refcompressratio 	1.00x 	 - 
 data 	 written 	 	5.24T 	 -

Code:
  pool: data
 state: ONLINE
  scan: none requested
config:

	NAME                       STATE     READ WRITE CKSUM     CAP            Product
	data                       ONLINE       0     0     0
	  mirror-0                 ONLINE       0     0     0
	    c2t5000CCA228C2D931d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C2E76Ad0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	  mirror-1                 ONLINE       0     0     0
	    c2t5000CCA228C2EEC5d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C2FBE0d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	  mirror-2                 ONLINE       0     0     0
	    c2t5000CCA228C3036Fd0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C308BEd0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	  mirror-3                 ONLINE       0     0     0
	    c2t5000CCA228C308CEd0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C30C35d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	  mirror-4                 ONLINE       0     0     0
	    c2t5000CCA228C311CFd0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C31376d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	  mirror-5                 ONLINE       0     0     0
	    c2t5000CCA228C313B5d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303
	    c2t5000CCA228C31446d0  ONLINE       0     0     0     3 TB           Hitachi HDS5C303

errors: No known data errors

  pool: rpool
 state: ONLINE
  scan: none requested
config:

	NAME        STATE     READ WRITE CKSUM     CAP            Product
	rpool       ONLINE       0     0     0
	  c4t0d0s0  ONLINE       0     0     0     32.2 GB        Virtual disk

errors: No known data errors

Code:
DDT-sha256-zap-duplicate: 2289889 entries, size 353 on disk, 224 in core
DDT-sha256-zap-unique: 57840997 entries, size 318 on disk, 183 in core

DDT histogram (aggregated over all DDTs):

bucket              allocated                       referenced
______   ______________________________   ______________________________
refcnt   blocks   LSIZE   PSIZE   DSIZE   blocks   LSIZE   PSIZE   DSIZE
------   ------   -----   -----   -----   ------   -----   -----   -----
     1    55.2M   6.90T   6.90T   6.90T    55.2M   6.90T   6.90T   6.90T
     2    2.14M    274G    274G    274G    4.47M    572G    572G    572G
     4    39.9K   4.98G   4.98G   4.98G     175K   21.9G   21.9G   21.9G
     8    3.20K    409M    409M    409M    30.8K   3.85G   3.85G   3.85G
    16      498   62.2M   62.2M   62.2M    11.6K   1.45G   1.45G   1.45G
    32       74   9.25M   9.25M   9.25M    2.93K    376M    376M    376M
    64       36   4.50M   4.50M   4.50M    3.43K    439M    439M    439M
   128       10   1.25M   1.25M   1.25M    1.65K    211M    211M    211M
   256        5    640K    640K    640K    1.57K    201M    201M    201M
   512        1    128K    128K    128K      674   84.2M   84.2M   84.2M
    1K        1    128K    128K    128K    1.42K    181M    181M    181M
    4K        1    128K    128K    128K    6.26K    801M    801M    801M
    8K        1    128K    128K    128K    14.6K   1.83G   1.83G   1.83G
  256K        1    128K    128K    128K     420K   52.5G   52.5G   52.5G
 Total    57.3M   7.17T   7.17T   7.17T    60.3M   7.54T   7.54T   7.54T

dedup = 1.05, compress = 1.00, copies = 1.00, dedup * compress / copies = 1.05

I can provide more information if needed I have contemplated converting the mirrors to raid-z2 since it should give me atleast a little more space while still keeping the IO fine for my type of backup.

Thank you for your time.
 
Last edited:
I would stick to raid10. The random IOPS will be superior (especially for reads) to raidz*.
 
Could you explain a bit more about what you mean when you call it a "iSCSI backup location"? Does that mean it's just an offline backup dumping area, or are you expecting your iSCSI targets to fail-over to this pool if the primary location goes down?
Once we understand a bit more about what the performance requirements would be for it then we can think about whether moving to RAID-Z would be a good idea or not.

Slightly wild idea here - since this is for backups, I presume you could take the pool offline temporarily and normal service wouldn't be interrupted (except obviously you wouldn't be able to take backups)?
Why not buy a couple more drives and add another mirror? Best would be to destroy and rebuild the pool so that the data would be balanced (AFAIK ZFS doesn't spread the existing data in the pool if new vdevs are added, it'll just add new data to the new vdev since it is less full, which isn't ideal for performance).
Could the data be copied back from other backups, or?
 
Need more info on hardware. Chassis, Processor, MB, and Memory.
 
Could you explain a bit more about what you mean when you call it a "iSCSI backup location"? Does that mean it's just an offline backup dumping area, or are you expecting your iSCSI targets to fail-over to this pool if the primary location goes down?
Once we understand a bit more about what the performance requirements would be for it then we can think about whether moving to RAID-Z would be a good idea or not.

This is a Disk backup location for my virtual machines. I use Veeam which basically snapshots the VMs and then stores them on this array. There is no fail over but every once in awhile I have too boot two or three virtual machines from this array to get files or restore settings. That's why I went with a mirrored array original verse a raid-z.

Slightly wild idea here - since this is for backups, I presume you could take the pool offline temporarily and normal service wouldn't be interrupted (except obviously you wouldn't be able to take backups)?

Correct I can normally take 12 hours worth of downtime without any impact to backups. If I need more I could point all backups at a different location for a short time and then bring back the files onto this system.

Why not buy a couple more drives and add another mirror? Best would be to destroy and rebuild the pool so that the data would be balanced (AFAIK ZFS doesn't spread the existing data in the pool if new vdevs are added, it'll just add new data to the new vdev since it is less full, which isn't ideal for performance).
Could the data be copied back from other backups, or?

Right now I have space for 4 more drives without having to add new interface hardware. After that I have to get an expander.

Need more info on hardware. Chassis, Processor, MB, and Memory.

Specs.
Code:
Supermicro X9SCM-F - I think
2x M1015
12x  Hitachi HDS5C303
Intel Xeon E3-1230 Sandy Bridge 3.2GHz LGA 1155 80W Quad-Core Server Processor
2x Kingston 8GB (2 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333 (PC3 10600) ECC Unbuffered Server Memory Model KVR1333D3E9SK2/8G
SUPERMICRO SuperChassis CSE-846TQ-R900B - I think.
 
This is a Disk backup location for my virtual machines. I use Veeam which basically snapshots the VMs and then stores them on this array. There is no fail over but every once in awhile I have too boot two or three virtual machines from this array to get files or restore settings. That's why I went with a mirrored array original verse a raid-z.

If the vast majority of the load is going to be sequential backup jobs and there will only occasionally be the need to access a few files then, in my opinion, a RAID-Z configuration would be much more suitable than mirrors.

In your situation you are wasting an unnecessary amount of disk capacity by mirroring, for no noticeable benefit. Mirrors are the best choice in a high-I/O situation where capacity is not the main priority, but in your scenario it's the opposite. You need more space but performance is not a problem.

I work with a SAN that feeds iSCSI targets for over 20 VMs, as well as doing general NFS file serving, and the performance is acceptable with a single large RAID-Z2 vdev, which is not an optimal configuration.
In your case you are only accessing one or two images on odd occasions, which by comparison is a very light load. RAID-Z vdevs (or RAID-Z2 or Z3) will be plenty fast enough for what you describe.

With your current 12 drives I would change to two 6-drive RAID-Z2 vdevs. That would give you approx 21TB of usable space rather than 16, and you would arguably have better protection against multiple drive failure.
If you wanted even more capacity you could use those extra 4 connections and either give each of those vdevs 8 drives each, or maybe even do three 5-drive RAID-Z vdevs and a spare.
 
Last edited:
If the vast majority of the time you are writing backups to it and only ever need to occasionally boot one or two of the images to read a few files then, in my opinion, a RAID-Z configuration would be much more suitable than mirrors.

In your situation you are wasting an unnecessary amount of disk capacity by mirroring, for no noticeable benefit.
I work with a SAN that feeds iSCSI targets for over 20 VMs, as well as doing general NFS file serving, and the performance is acceptable with a single large RAID-Z2 vdev, which is not an optimal configuration. In your case you are only accessing one or two images on odd occasions, which is a very light load. RAID-Z will still be plenty fast enough, especially if there are several striped together in the pool.

With your current 12 drives I would change to two 6-drive RAID-Z2 vdevs. That would give you approx 21TB of usable space rather than 16, and you would arguably have better protection against multiple drive failure.
If you wanted even more capacity you could use those extra 4 connections and either give each of those vdevs 8 drives each, or maybe even do three 5-drive RAID-Z vdevs and a spare.

Alright awesome thats is what I kind of was thinking on the raid-z volume front.
 
Code:
Supermicro X9SCM-F - I think
2x M1015
12x  Hitachi HDS5C303
Intel Xeon E3-1230 Sandy Bridge 3.2GHz LGA 1155 80W Quad-Core Server Processor
2x Kingston 8GB (2 x 4GB) 240-Pin DDR3 SDRAM DDR3 1333 (PC3 10600) ECC Unbuffered Server Memory Model KVR1333D3E9SK2/8G
SUPERMICRO SuperChassis CSE-846TQ-R900B - I think.

Well I would fix this first. 16GB for 18TB + Dedup and no L2ARC, and no ZIL on SSD isn't going to perform too well. You need more memory big time and L2ARC. You've gone too cheap in your build. The maximum amount of memory you can have is 32GB, which isn't too good.

But let's try to salvage the build. Add the maximum amount of memory (32GB) and some L2ARC and ZIL on SSD. However, this is just temporary. There's only so much L2ARC you can add before it actually hampers performance. These changes will improve performance greatly but it's just not going to scale. So plan to eventually replace the board and processor.

If you need more storage go with a JBOD + external SAS. But I wouldn't do that before addressing everything else first.
 
Alright awesome thats is what I kind of was thinking on the raid-z volume front.
I made a few edits since you quoted me :D

Well I would fix this first. 16GB for 18TB + Dedup and no L2ARC, and no ZIL on SSD isn't going to perform too well. You need more memory big time and L2ARC. You've gone too cheap in your build. The maximum amount of memory you can have is 32GB, which isn't too good.

But let's try to salvage the build. Add the maximum amount of memory (32GB) and some L2ARC and ZIL on SSD. However, this is just temporary. There's only so much L2ARC you can add before it actually hampers performance. These changes will improve performance greatly but it's just not going to scale. So plan to eventually replace the board and processor.

If you need more storage go with a JBOD + external SAS. But I wouldn't do that before addressing everything else first.
In my opinion, you're focusing on the wrong area. More RAM definitely wouldn't hurt, but by increasing the capacity of the pool as I have suggested, using dedup might not even be necessary. And we've established that there is going to be very little read load, so why worry about an L2ARC? This is a backup storage dump setup. Keep it simple, IMO.
 
Well I would fix this first. 16GB for 18TB + Dedup and no L2ARC, and no ZIL on SSD isn't going to perform too well. You need more memory big time and L2ARC. You've gone too cheap in your build. The maximum amount of memory you can have is 32GB, which isn't too good.

But let's try to salvage the build. Add the maximum amount of memory (32GB) and some L2ARC and ZIL on SSD. However, this is just temporary. There's only so much L2ARC you can add before it actually hampers performance. These changes will improve performance greatly but it's just not going to scale. So plan to eventually replace the board and processor.

If you need more storage go with a JBOD + external SAS. But I wouldn't do that before addressing everything else first.

Alright lets address this. 32GB is the maximum that I can place in the system.

It would appear that there are three choices
1. http://www.superbiiz.com/detail.php?name=D38GE1333H
2.http://www.superbiiz.com/detail.php?name=D38GE1600H
3.http://www.superbiiz.com/detail.php?name=D38GE1600S

I would assume that number two is the best of the three price and quality wise. So at basically 300$ I double my ram.

The l2arc does not need to be super cap thats for the ZIL the way I understand it. More just a high speed cache that is a step between spindles and ram for caching files and the lookup table for dedup.

So maybe a M4 or Samsung 840?

I should also mention that I have 16 x 1TB Hard drives that have been used in a dell raid setup that are no longer needed and I could use them for temporary storage.


I made a few edits since you quoted me :D


In my opinion, you're focusing on the wrong area. More RAM definitely wouldn't hurt, but by increasing the capacity of the pool as I have suggested, using dedup might not even be necessary. And we've established that there is going to be very little read load, so why worry about an L2ARC? This is a backup storage dump setup. Keep it simple, IMO.
The l2arc would be used for caching the dedup table since it will not fit in ram.

I think the most interesting part of this is my deup is not working the way my math shows which means even if I am getting the worst performance right now I still am not getting deduplicated data.
 
The l2arc does not need to be super cap thats for the ZIL the way I understand it. More just a high speed cache that is a step between spindles and ram for caching files and the lookup table for dedup.

When you say the L2ARC "doesn't need to be super cap" you mean it doesn't need very much storage space?
I think you might have the two things the wrong way around - the way I understood it, the ZIL is just keeping track of write operations until they are successfully completed and doesn't get very large.
The L2ARC on the other hand will use as much space as you give it, because it's basically another tier of read cache, between the L1ARC that is kept in RAM and the disks in the pools.
In the same way that you can give the system more RAM to use as L1ARC, you can use a larger drive for the L2ARC and the system will use as much as it can.

Anyone know for sure and can correct me on that?
 
When you say the L2ARC "doesn't need to be super cap" you mean it doesn't need very much storage space?
I think you might have the two things the wrong way around - the way I understood it, the ZIL is just keeping track of write operations until they are successfully completed and doesn't get very large.
The L2ARC on the other hand will use as much space as you give it, because it's basically another tier of read cache, between the L1ARC that is kept in RAM and the disks in the pools.
In the same way that you can give the system more RAM to use as L1ARC, you can use a larger drive for the L2ARC and the system will use as much as it can.

Anyone know for sure and can correct me on that?

I mean to say that the ZIL needs a super cap that can flush the writes from the local ssd cache to the internal memory before failure but this is not a requirement for the l2arc since its a read only medium and all information is also store on the disks.
 
In my opinion, you're focusing on the wrong area. More RAM definitely wouldn't hurt, but by increasing the capacity of the pool as I have suggested, using dedup might not even be necessary. And we've established that there is going to be very little read load, so why worry about an L2ARC? This is a backup storage dump setup. Keep it simple, IMO.

Actually no. The more storage you have the larger your metadata table will be. If he didn't have dedup on then he would be OK ( still not where I would like it) . Be he doesn't. Adding even more storage increases the use of ARC and since he's limited and really he's way past the minimum ARC recommendation for enabling dedup, your suggestion will lower his performance if he leaves dedup on. His write performance is already going to suffer due to enabling raidz. Without SSD based ZIL (a really good idea for backup) or enough ARC/L2ARC (a good idea for dedup) his OS disks are going to thrash and absolutely tank performance.

Even if we are talking backups you want to decrease your backup window. The longer the backup window the larger your window of risk. If this was for home use I would let it be. I still would recommend more memory though. However, I'm seeing Exchange, and domain controllers so I'm assuming this is work related. If that's the case then we want to balance performance a bit against his storage requirements.

This is why I put in the recommendation of JBOD and external sas. I've seen the case he has and 24 slots is good for your typical file server but it's not enough for backups for an entire rack. Whatever he has in live production he needs twice that at the minimum. The OP needs more cages so that he is protected in the future. Raidz2 will only provide temporary relief. I'm thinking long term. Basically he needs your suggestion (raidz setting) + and upgrade in hardware.
 
Last edited:
Alright lets address this. 32GB is the maximum that I can place in the system.

It would appear that there are three choices
1. http://www.superbiiz.com/detail.php?name=D38GE1333H
2.http://www.superbiiz.com/detail.php?name=D38GE1600H
3.http://www.superbiiz.com/detail.php?name=D38GE1600S

I would assume that number two is the best of the three price and quality wise. So at basically 300$ I double my ram.
Yup I would go option 2.

The l2arc does not need to be super cap thats for the ZIL the way I understand it. More just a high speed cache that is a step between spindles and ram for caching files and the lookup table for dedup.

So maybe a M4 or Samsung 840?
That will work. Correct. Supercap isn't needed for L2ARC. Supercap for ZIL however would be nice if you are enabling write back.

I should also mention that I have 16 x 1TB Hard drives that have been used in a dell raid setup that are no longer needed and I could use them for temporary storage.
Keep that for your JBOD or maybe for some other use down the road.

Here's what I would do.

SHORT TERM PLAN:

Max out your RAM + add ZIL + add L2ARC and buy a chassis (JBOD) + external SAS card. All new storage added to the JBOD can be raidz2.

LONG TERM PLAN:

Go for a current board either Intel 2011 or AMD G34. Make sure the MB has 8 slots per socket. Get as much RAM as you can. You'll be able to reuse it anyway along with your current 846TQ. Keep the 8GB sticks and with 8 slots per socket your RAM will double to 64 once you outgrow the short term config. At this point you'll have taken your total storage for backup if you stayed with 3TB disks (don't know what your budget is) to 84TB using radiz2 and 45 slot JBOD (not including what's in your 846TQ). At this point you can disable dedup and be good for the foreseeable future. If you go with a 2P motherboard you'll have room to grow even if you outstrip this and it will even allow you to re-enable dedup if you want.
 
Last edited:
Yup I would go option 2.


That will work. Correct. Supercap isn't needed for L2ARC. Supercap for ZIL however would be nice if you are enabling write back.


Keep that for your JBOD or maybe for some other use down the road.

Here's what I would do.
So question the way I understand it is that zfs using iscsi mounted on a windows host there is no real use for a zil since all iscsi traffic routes around the zil.

Now I have no real reason for using iscsi vs nfs unless there is some specialty reason. Right now I have a single host connecting with no real intention of expanding that but removing the iscsi abstraction layer means I can move the data around from different volumes without having to use a windows machine.

I would also make the assumption a L2ARC of about 256GB should be good but is it better to have two 128GB in a striped fashion for better io speed?

SHORT TERM PLAN:

Max out your RAM + add ZIL + add L2ARC and buy a chassis (JBOD) + external SAS card. All new storage added to the JBOD can be raidz2.

LONG TERM PLAN:

Go for a current board either Intel 2011 or AMD G34. Make sure the MB has 8 slots per socket. Get as much RAM as you can. You'll be able to reuse it anyway along with your current 846TQ. Keep the 8GB sticks and with 8 slots per socket your RAM will double to 64 once you outgrow the short term config. At this point you'll have taken your total storage for backup if you stayed with 3TB disks (don't know what your budget is) to 84TB using radiz2 and 45 slot JBOD (not including what's in your 846TQ). At this point you can disable dedup and be good for the foreseeable future. If you go with a 2P motherboard you'll have room to grow even if you outstrip this and it will even allow you to re-enable dedup if you want.

So lets talk about short term versus long term. I have about 5k in this years budget that could be allocated without a question to this. Right now this box is onsite but it will become an offsite backup connected via a 20mbit connection or about (125GB/day). Slow I know but better than what we had before. This should factor into what type of z-raid level we use. Since in all cases except when we sneakernet a dataset back for recover do we have anything over 20mbps.

I only need a working set of 3 months onsite and then everything else would be farmed off to this box. At my current rate my 14 day working set uses about 2.2TB I would assume that I could get away with about 5-10TB for a 3 month set with no deduplication for the onsite location.

Also I have two extra cases that are lying dormant they both could be converted to JBOD

1x Norco RPC-4020
1x Norco RPC-4224
Single 500-600w Power supply


So maybe what I should be looking at replacing this host with a new host stealing its hard drives for the new host and placing all the older 1TB drives into it.
 
He's using ZFS for iSCSI - a SLOG for your ZIL would be a waste of money.

Could the reason that you're not seeing the dedupe ratios you were hoping for, be the result of dedupe and compression used in Veeam? Usually incremental backup data (which is a feature of Veeam) is not an optimal candidate for backups. Full backups, yes, but I doubt that is the case here.


If your current performance is acceptable, why not just disable dedupe, add more RAM and then wait and see what happens to your dataset? It might not grow much, and over time the amount of deduped data that ZFS needs to keep track of would be less and less.
 
He's using ZFS for iSCSI - a SLOG for your ZIL would be a waste of money.

Could the reason that you're not seeing the dedupe ratios you were hoping for, be the result of dedupe and compression used in Veeam? Usually incremental backup data (which is a feature of Veeam) is not an optimal candidate for backups. Full backups, yes, but I doubt that is the case here.


If your current performance is acceptable, why not just disable dedupe, add more RAM and then wait and see what happens to your dataset? It might not grow much, and over time the amount of deduped data that ZFS needs to keep track of would be less and less.

Veeam has compression in dedup friendly mode. I actually am doing more Full datasets then incrementals the only actual incrementals is my 6 weeks worth but everything else will be full backups.
Code:
Age of Data	Granularity	Amount
14 Days		12 Hours	3x Full then 28x Incremental
6 Weeks		1 Week		2x Full then 5x Incremental
12 Months	1 Month		12x Full
6 Years		3 Months	24x Full
 
Why not use incrementals?

I could use reverse incrementals I have been bitten too many times by backup exec breaking incremental to trust it I know veeam is a different product But at what point do you change from incremental to full. Once a year.
 
We do daily incremental forever of thousands of servers (using TSM) and have yet to encounter an issue due to this. I wouldn't be afraid to run Veeams reverse incremental.

As with most things, and probably more important regarding backup: Test it.
 
So question the way I understand it is that zfs using iscsi mounted on a windows host there is no real use for a zil since all iscsi traffic routes around the zil.
Whoops sorry I missed the part where Veem is performing the backup. You are correct you can forgo the ZIL as long as your basic setup remains the same.

I would also make the assumption a L2ARC of about 256GB should be good but is it better to have two 128GB in a striped fashion for better io speed?
Not really needed.

So lets talk about short term versus long term. I have about 5k in this years budget that could be allocated without a question to this. Right now this box is onsite but it will become an offsite backup connected via a 20mbit connection or about (125GB/day). Slow I know but better than what we had before. This should factor into what type of z-raid level we use. Since in all cases except when we sneakernet a dataset back for recover do we have anything over 20mbps.

I only need a working set of 3 months onsite and then everything else would be farmed off to this box. At my current rate my 14 day working set uses about 2.2TB I would assume that I could get away with about 5-10TB for a 3 month set with no deduplication for the onsite location.

Also I have two extra cases that are lying dormant they both could be converted to JBOD

1x Norco RPC-4020
1x Norco RPC-4224
Single 500-600w Power supply


So maybe what I should be looking at replacing this host with a new host stealing its hard drives for the new host and placing all the older 1TB drives into it.

That connection is so slow then I just wouldn't even worry about it. Max out the memory, change to a higher raidz level, turn off dedup and call it day. $5,000 will buy you a considerable amount of storage. If it was remaining onsite and local attached there is a case to be made. If all of that is going away then well there's isn't too much to worry about on the performance side of it.
 
We do daily incremental forever of thousands of servers (using TSM) and have yet to encounter an issue due to this. I wouldn't be afraid to run Veeams reverse incremental.

As with most things, and probably more important regarding backup: Test it.

Alright I will setup a job that does this just for testing. Are you running reverse or forward incrementals? since it seems either way you have to go through almost all of them depending on forward/reverse incremental backup method.

Whoops sorry I missed the part where Veem is performing the backup. You are correct you can forgo the ZIL as long as your basic setup remains the same.

Not really needed.

Do you think I will get better performance via NFS Vs. iSCSI for dedup since the iSCSI file system is muddling it up.

That connection is so slow then I just wouldn't even worry about it. Max out the memory, change to a higher raidz level, turn off dedup and call it day. $5,000 will buy you a considerable amount of storage. If it was remaining onsite and local attached there is a case to be made. If all of that is going away then well there's isn't too much to worry about on the performance side of it.


So basically I need two pieces of hardware. This device was originally intended for an offsite backup location. What happens is that I had a device fail and before I could place this at the off site location I started using it for on site backups. Now what I do need to do is keep 3 months worth of my backups on site with the rest residing off site. I don't mind replacing the hardware either way but it would seem that I need more ram if I want to do dedup on the off site location. I guess I could just use L2ARC and make the assumption that I will never need to write to it fast but I think that is a bad decision.
 
TSM only has support for forward incremental. I think (as in, not sure at all) that reverse incremental is a proprietary Veeam tech.
 
Do you think I will get better performance via NFS Vs. iSCSI for dedup since the iSCSI file system is muddling it up.
Not likely. Performance will be similar on reads but iSCSI will be better for writes. I would max out the memory first before trying to assess any additional performance issues though.


So basically I need two pieces of hardware. This device was originally intended for an offsite backup location. What happens is that I had a device fail and before I could place this at the off site location I started using it for on site backups. Now what I do need to do is keep 3 months worth of my backups on site with the rest residing off site. I don't mind replacing the hardware either way but it would seem that I need more ram if I want to do dedup on the off site location. I guess I could just use L2ARC and make the assumption that I will never need to write to it fast but I think that is a bad decision.

Hmm if this is already your back up box then you kind of have some things to address beyond just this storage unit. Restore your alternate location first, then address the additional hardware concerns for this current box.,
 
Back
Top