Increase iSCSI Performance over 1 GB Ethernet?

KapsZ28

2[H]4U
Joined
May 29, 2009
Messages
2,114
I've asked about stuff like this before, but my boss and I can't seem to agree. We have a client running Windows VM's and needs the disk speed to exceed 100 MB/s. His VMs reside on an OpenFiler SAN with six Intel 320 Series SSDs in RAID 5. Software RAID unfortunately. Below is a speed test ran from the Windows VM.



Currently the datastore is mounted as NFS with two 1 GB NICs. We will most likely be switching to iSCSI as all our iSCSI datastores seem to perform better.

Here is a speed test of a NetApp SAN setup with iSCSI over 1 GB ethernet.



Theoretically we should be able to get up to 125 MB/s on 1 GB ethernet with zero network congestion. My boss thinks that if we setup LAG and MultiPath that we should be able to exceed 200 MB/s through two 1 GB connections. From what I've been told, it is not possible to push a single stream of data over two separate NICs to double the speed. So when running a speed test like above, all the traffic will still go over only 1 NIC at a time. Is that correct?
 
keep in mind TCP overhead, which can be 10% of your total 125MB/s.

From my experience LACP doesn't scale like you think it will.
 
Multi-path is for fail over or multiple connections. You don't add up the speeds together for one bigger pipe.

I run iSCSI disk off a 6x2TB RAID 5 array and saturate 1Gb links at about 115MB/s rw

What do the benchmarks look like if you run them locally on the target server?
 
Excuse me, are you saying single GbE and you want more than 100 MB/sec?

Comparing NFS and iSCSI performance is very tricky.

What kind of software RAID?

Using a compressed filesystem on top of iSCSI, if you have compressable data, is one way.
 
MPIO will achieve what you want and is very easy to set up.

I have an openindiana SAN box at work with MPIO across 2 GbE to a hyper-v host and get up to 230MBps in-vm on a single file transfer/benchmark.

All you need is 2+ NICs in each end, configure each with its own IP, then set up each path in ms iscsi initiator.


Forget LACP
 
The other question looking at current write speeds would be, is there any write caching going on at the software raid or iscsi target level?

Looks like there may not be looking at the throughput you have.
 
He doesn't list what is running the vm's.

But esxi runs nfs in sync, and iscsi in async. So slow write speeds with nfs would be expected over iscsi.

But yes, nfs won't scale over one nic, you have to use multipath, and nfs doesn't support that.
 
MPIO will achieve what you want and is very easy to set up.

I have an openindiana SAN box at work with MPIO across 2 GbE to a hyper-v host and get up to 230MBps in-vm on a single file transfer/benchmark.

All you need is 2+ NICs in each end, configure each with its own IP, then set up each path in ms iscsi initiator.


Forget LACP

This is interesting. I have been fussin around with LACP and find it to pretty much suck for what I need, even if 4 hosts are accessing the same server under a 4 NIC lag. I'm going to test this approach instead as it's simpler and I find single copper performance to be very good in my env with iSCSI, over 100MB no problem.
 
How long of a run are you trying to accommodate?

Can you just use 10/40Gb?

I wish. Not really in the budget. We did acquire a Dell blade chassis and PowerVault SAN in another datacenter that is all 10 GB. Not in the right location for this project. Below is the performance.



Not bad. A little disappointed on write speed.
 
Multi-path is for fail over or multiple connections. You don't add up the speeds together for one bigger pipe.

I run iSCSI disk off a 6x2TB RAID 5 array and saturate 1Gb links at about 115MB/s rw

What do the benchmarks look like if you run them locally on the target server?

Are you asking what are the benchmarks if I was running the VM on local storage to the ESXi host? We don't run any VMs that way.
 
Excuse me, are you saying single GbE and you want more than 100 MB/sec?

Comparing NFS and iSCSI performance is very tricky.

What kind of software RAID?

Using a compressed filesystem on top of iSCSI, if you have compressable data, is one way.

No, saying with at least two 1 GB NICs. Right now it is NFS, but I know that doesn't load balance and it would be changed to iSCSI.
 
MPIO will achieve what you want and is very easy to set up.

I have an openindiana SAN box at work with MPIO across 2 GbE to a hyper-v host and get up to 230MBps in-vm on a single file transfer/benchmark.

All you need is 2+ NICs in each end, configure each with its own IP, then set up each path in ms iscsi initiator.


Forget LACP

It sounds like you are talking about using MPIO at the OS level which we don't use. In other locations we have the SAN setup with multiple NICs, different IP address. Then connected to the ESXi host as separate iSCSI Storage Adapters. This doesn't give us any disk improvement.
 
LACP does not increase individual transfer speeds, and depending on setup may not even increase per-host transfer speeds. That is not what it is for. It is for link redundancy, period. If you're making lots of connections and are using LACP with certain settings it might improve your aggregate throughput across the connections/streams, but it isn't going to do that for a single file copy or benchmark test. No single stream is going to outdo the speed of the individual links, because they're only passing over one of the links.

For this reason, NFS is not going to ever go faster than one link's worth - you can stick 8 x 1 Gbit NIC's in there and LACP the whole thing, and your NFS file transfers are still going to go at 100 MB/s. Depending on LACP setup, you could have say 8 clients or maybe even 8 separate NFS mounts pushing 100 MB/s each, but no individual file copy is going over 100 MB/s.

iSCSI has something called MPIO - if you properly enable it on both the client and the server, it can utilize multiple independent NIC's simultaneously, thus a single file copy could go faster than a single link. It isn't a perfect return, you'll lose some off the top, but it'll pretty easily push the majority of the speed of all the links, assuming proper setup. From a network bottleneck perspective, iSCSI thus will generally win over NFS.

As some people have already touched on, you should also be careful with these benchmarks. The default settings typically for NFS are going to be synchronous, which is safer in the event of a hardware failure somewhere in the stack. iSCSI is often by default asynchronous, which is less safe and more prone to data loss/filesystem corruption. Before comparing the two, you should decide if you're OK with async or if you desire sync, and modify the settings appropriately. I think you'll find that before you bring MPIO into the picture, NFS and iSCSI are similar in performance when operating similarly.

If your budget precludes 10Gbit networking, and you really need > 100 MB/s, iSCSI really is your only option (there's no MPIO NFS solution that I'm aware of). If you can go 10Gbit, however, it's rare to find storage solutions that can even really drive more than 1 GB/s or clients that need more than that on a single transfer, so the 'win' of MPIO on iSCSI is lost, and then honestly NFS has the edge, because if performance isn't your sole and only goal, NFS is just a better way to go (it gives the storage more semantics, it is far more forgiving of network quality issues and network loss events, you've got more introspection into the data on the storage side, etc, etc).
 
He doesn't list what is running the vm's.

But esxi runs nfs in sync, and iscsi in async. So slow write speeds with nfs would be expected over iscsi.

But yes, nfs won't scale over one nic, you have to use multipath, and nfs doesn't support that.

So how do you multipath at the hardware level or is that not possible?
 
iSCSI has something called MPIO - if you properly enable it on both the client and the server, it can utilize multiple independent NIC's simultaneously, thus a single file copy could go faster than a single link. It isn't a perfect return, you'll lose some off the top, but it'll pretty easily push the majority of the speed of all the links, assuming proper setup. From a network bottleneck perspective, iSCSI thus will generally win over NFS.

When you say to enable it on the client, are you referring to adding to NICs to a Windows VM and enabling MPIO within Windows? Or is this something that can be completely done at the hardware level?
 
Kaps, you need single-stream throughput from an OpenFiler Box to ESXi hosts of over 100MB/s right? How many NICs are in the OpenFiler box and how many NICs do you have available on the Hosts?
 
Kaps, you need single-stream throughput from an OpenFiler Box to ESXi hosts of over 100MB/s right? How many NICs are in the OpenFiler box and how many NICs do you have available on the Hosts?

Yes, that is correct. More than 100MB/s. Right now we have three ESXi hosts with six 1GB NICs. The OpenFiler SAN only has two 1GB NICs, but that can be upgraded. We plan on rebuilding it anyway and using a real RAID card rather than software RAID.
 
What is this hardware level your talking about?

Openfiler is hardly hardware level. nfs can never be done hardware level. Iscsi can be done hardware level, including mpio even, but normally isn't these days, expecially on esxi.

Hardware level anything restricts you to the features of that hardware, and that is normally very expensive and limiting compared to software level.

Like, esxi, hardware level iscsi works, is fast, but no mpio support (if remember right), and no jumbo frames. Using software level iscsi can use normal nics, supports mpio, and you can even use jumboframes.

I would not do anything your asking *hardware level*.

The only *hardware level* things I would recommend is a battery backed raid card, if you need one. Or a real hardware iscsi/fc card if you needed iscsi/fc diskless boot.

Otherwise using fancy hardware, gets to be a very very specific usecase, and normally isn't good for most other uses.
 
What is this hardware level your talking about?

Openfiler is hardly hardware level. nfs can never be done hardware level. Iscsi can be done hardware level, including mpio even, but normally isn't these days, expecially on esxi.

Hardware level anything restricts you to the features of that hardware, and that is normally very expensive and limiting compared to software level.

Like, esxi, hardware level iscsi works, is fast, but no mpio support (if remember right), and no jumbo frames. Using software level iscsi can use normal nics, supports mpio, and you can even use jumboframes.

I would not do anything your asking *hardware level*.

The only *hardware level* things I would recommend is a battery backed raid card, if you need one. Or a real hardware iscsi/fc card if you needed iscsi/fc diskless boot.

Otherwise using fancy hardware, gets to be a very very specific usecase, and normally isn't good for most other uses.

Sorry, I probably shouldn't have said "hardware". I just mean, not at the OS level such as doing it in Windows. Doing it on the ESXi host is what I mean. We have a NetApp in another site using iSCSI and they are connected to each host with 2 NICs and 2 IP addresses. The data store path is setup for round robin and shows both NICs active. However it still doesn't exceed 100 MB/s.

So my question is how would configure MPIO on an ESXi host so a single data-stream utilizes both NICs?
 
You won't be able to exceed 100ish in a single stream with ESXi. The native Round Robin MPT in ESXi won't aggregate multiple paths, and it set to bounce between connections every 100 IOs, if I remember right.

With 3 hosts, your options are limited. Are you stuck on Openfiler? Assuming that OF is used because of the price, I'll assume that cheap is the order of the project. Might look as some used/Refurb 4GB fiber cards. A 4 port for the NAS box and a card for each host. Single higher bandwidth paths. Should be able to do it on the cheap also.
 
You won't be able to exceed 100ish in a single stream with ESXi. The native Round Robin MPT in ESXi won't aggregate multiple paths, and it set to bounce between connections every 100 IOs, if I remember right.

With 3 hosts, your options are limited. Are you stuck on Openfiler? Assuming that OF is used because of the price, I'll assume that cheap is the order of the project. Might look as some used/Refurb 4GB fiber cards. A 4 port for the NAS box and a card for each host. Single higher bandwidth paths. Should be able to do it on the cheap also.

So, is there any scenario with MPIO where you could double the speed of a single data stream? One of the other members mentioned 230 MB/s using MPIO in Windows.
 
I dont know about OF for sure, but ESXi just won't aggregate or bond iSCSI connections together natively. So you are stuck with 10GbE or Fibre for added bandwidth. There are a few prorietary plugins for MPIOthat are more robust, but they are paired woth specific hardware. i.e. Equallogic HIT kit MPIO plugin.
 
Last edited:
I dont know about OF for sure, but ESXi just won't aggregate or bond iSCSI connections together natively. So you are stuck with 10GbE or Fibre for added bandwidth. There are a few prorietary plugins for MPIOthat are more robust, but they are paired woth specific hardware. i.e. Equallogic HIT kit MPIO plugin.

What about with NetApp?
 
Does NetApp have an ESXi MPIO plugin? I'm not a NetApp dealer, so I don't know. Both the Host (ESXi) and the SAN/NAS have to support whatever method of MPIO/bonding in order to exceed the single link limit. The basic installs of ESXi are not capable, there's your limitation.
 
Nate7311, shut your nonsense.

If you don't know how to configure esxi for mpio, don't blame it on esxi.

esxi will do whatever you tell it to do. Yes, default setting is 100iops per path.

It's easy enough, and recommended by netapp and several others to lower it to 3 iops or even 1iop per path, in the roundrobin scheduler.

You can adjust it per iop and per block size. Since I use jumboframes, I normally set it to 8k and 1iop and set every lun to roundrobin.
 
Patrick, absolutely that's possible, but that's a manual setting and that still doesn't get past the single stream barrier.
 
Yes, esxi won't aggergate or bond iscsi together, cause the whole point to MPIO is to NOT DO THAT.

over 4 1gbit connections, I normally get 380mb/sec on esxi.

I have no idea what this said equilogic module does, but mpio is fully supported with alua. If you can't take full advantage of all your paths, either you configured it that way, or you lacked to configure it.
 
Do you have any idea what MPIO is? there is no way you can compare it to *single stream* or aggrates, or bonding.

Enough of this, let the blind lead the blind for the rest of this thread.
 
Well then, I'll let you refine the OPs goal and explain the configuration then...
 
Greetings

What about using two cans and a piece of string? what I have in mind are two Intel QLE7340 Infiniband 40Gbs adapters connected with a Mellanox Active Fiber Cable, IB QDR/FDR10, 40Gb/s, QSFP cable, all up cost just under a grand. Would this be within the allowable budget and get the job done? I presume since its just point to point you don't need a switch as all you have to do is run a fabric manager at one end.

If latency is not much of a problem then perhaps two cheaper 10Gbe ethernet adapters might work also.

Cheers
 
To clarify, our setup is set to use mpio in 'lowest queue' mode. This is configured on a hyper-v host.

I have read in the past about esxi mpio and changing the queue per path to achieve better low qd throughput....
 
Well, I changed on one of our NetApp LUNs that is using iSCSI and I am not seeing any improvement.

~ # esxcli storage nmp psp roundrobin deviceconfig get -d naa.60a980003753496c422442494f394267
Byte Limit: 10485760
Device: naa.60a980003753496c422442494f394267
IOOperation Limit: 1
Limit Type: Iops
Use Active Unoptimized Paths: false

Although, it may be because of the way the network is setup. (FYI, I didn't configure this.)

So, in this first screen shot is the standard switch that was setup for our NetApp storage. Both IP addresses are on the same VLAN and subnet. Shouldn't they be on different subnets? Also, should it be 1 standard switch for both NICs, or a standard switch for each NIC?

The second screen shot shows the datastore path and both NICs on the ESXi host are going to the same IP address on the NetApp rather than two different IPs on different subnets. Again, that appears incorrect to me.




 
Given the networking screenshot above, I see 2 vKernels and 2 NICs. Look at each vKernel. Do they each only have a single NIC bound? i.e. the opposite one set in the "Unused Adapters" section for each vKernel.
 
Given the networking screenshot above, I see 2 vKernels and 2 NICs. Look at each vKernel. Do they each only have a single NIC bound? i.e. the opposite one set in the "Unused Adapters" section for each vKernel.

Yes, they both only have a single NIC bound. The other is set to Unused. Is that adequate? I do need to verify the NetApp configuration though.
 
Looks like the tests you where doing used 1MB blocks.

Even though you set the iops=1 on esxi, you still left the iop size at 10MB (Byte Limit: 10485760)

You probably want to turn that down to like 4k or 8k, so that those writes get split into multible iops over your multible paths. Or that test likely isn't going show much improvement.

It does sound like only 1 port on the netapp is configured, so still a bottlenech there. You will have to refresh on esxi when it's corrected to show up atleast 4 paths.
 
Back
Top