• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

ESXi Host and SAN connectivity

Adam12176

Limp Gawd
Joined
Dec 27, 2007
Messages
205
I'm seeing a lot of really strange behavior out of the two Dell R810 hosts I'm using.

Both hosts are using the integrated Broadcom NICs for vMotion/Mgmt/VM traffic, and an Intel ET2 Quad NIC for iSCSI.

Installed the Dell OMSA Agent and patched two hosts with vUM. Host A came back up fine, after a management network issue. Waited a few days, and now patching host B - same exact set of patches, same server, bios level, everything.

Host B suddenly can't connect to the EqualLogic group IP on reboot. It had the same issues as the other host when it was rebooted as well, can't ping the management address until you log into the DCUI and restart management network. Pings immediately after that.

SSH'ed into the Array and two hosts, results below:
Host A: Can ping everything. 4 VMKernel ports per host, and the group IP
Host B: Can ping all VMKernel ports, cannot ping the group IP.
Array: Can ping all VMKernel ports on Host A. Error message is actually "No route to host" when pinging Host B, thought that was weird.

Rescanning the HBA just comes back "iSCSI Initiator couldn't establish a network connection to the discovery address", which is what I would expect if it won't even ping. The EQL log shows a timeout yesterday when patching, but after that there are no logged connection attempts from Host B at all.

Dell just wants to update the switch firmware. Which is fine, but this was working pre-patch, and I have one host working post patch, I'm guessing there is something I'm missing in the config or a vital step. Any ideas?
 
Show me your networking config.

esxcfg-vswitch -l
esxcfg-vmknic -l
esxcfg-nics -l
(output from the command line - paste here).
Gratzi.
 
Not sure if this might also help in determining the issue:

To use ping and/or traceroute from the PS Aarray,

telnet/ssh to one of the eth interfaces (do not use the group ip, if you have more than one array member in the group, connect to each member to perform the test)

login as groupadmin, and at the groupname prompt:

Ping:
to ping out of each of the specific ETH port interfaces add the switch to choose the interface IP as shown below:
ping "-I <source_ETH_IP> <dest IP>"
(that is a &#8211;I as in Capital letter &#8220;eye&#8221;, and don't forget the quotes after ping and at the end of the argument string).

Test all IP combinations (eth0, eth1, eth2, etc.)

Traceroute:
To traceroute out each of the specific ETH port interface add a switch to choose the interface IP as shown below (note the cli command is "support traceroute"):

GrpName>support traceroute "-s [ETH port source IP] [destinationIP]"

(don't forget the quotes after support traceroute and at the end of the argument string).

Ensure to test each ETH interface combination from all members of the group

Test all IP combinations (eth0, eth1, eth2, etc.)

-joe
 
Note: This is the problem host. Two vmk nics per machine are in use, vmnic 6/7 down is expected for now.

~ # esxcfg-nics -l
Name PCI Driver Link Speed Duplex MAC Address MTU Description
vmnic0 0000:01:00.00 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e6:1a 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic1 0000:01:00.01 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e6:1c 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic2 0000:02:00.00 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e6:1e 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic3 0000:02:00.01 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e6:20 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic4 0000:08:00.00 igb Up 1000Mbps Full 00:1b:21:ba:f5:9c 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic5 0000:08:00.01 igb Up 1000Mbps Full 00:1b:21:ba:f5:9d 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic6 0000:09:00.00 igb Down 0Mbps Half 00:1b:21:ba:f5:9e 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic7 0000:09:00.01 igb Down 0Mbps Half 00:1b:21:ba:f5:9f 1500 Intel Corporation 82576 Gigabit Network Connection


~ # esxcfg-vmknic -l
Interface Port Group/DVPort IP Family IP Address Netmask Broadcast MAC Address MTU TSO MSS Enabled Type
vmk1 vMotion IPv4 10.1.2.10 255.255.255.0 10.1.2.255 00:50:56:7d:c0:8c 1500 65535 true STATIC
vmk2 iSCSI Port 1 IPv4 10.1.1.10 255.255.255.0 10.1.1.255 00:50:56:70:47:52 1500 65535 true STATIC
vmk3 iSCSI Port 2 IPv4 10.1.1.11 255.255.255.0 10.1.1.255 00:50:56:7f:b5:54 1500 65535 true STATIC
vmk4 iSCSI Port 3 IPv4 10.1.1.12 255.255.255.0 10.1.1.255 00:50:56:71:b8:7f 1500 65535 true STATIC
vmk5 iSCSI Port 4 IPv4 10.1.1.13 255.255.255.0 10.1.1.255 00:50:56:70:ed:27 1500 65535 true STATIC
vmk0 Management Network IPv4 192.149.115.222 255.255.255.0 192.149.115.255 14:fe:b5:c9:e6:1a 1500 65535 true STATIC

~ # esxcfg-vswitch -l
Switch Name Num Ports Used Ports Configured Ports MTU Uplinks
vSwitch0 128 7 128 1500 vmnic0,vmnic1,vmnic2,vmnic3

PortGroup Name VLAN ID Used Ports Uplinks
ProdVM 0 0 vmnic2,vmnic3
Management Network 0 1 vmnic0
vMotion 0 1 vmnic1,vmnic0

Switch Name Num Ports Used Ports Configured Ports MTU Uplinks
vSwitch1 128 9 128 1500 vmnic4,vmnic5,vmnic6,vmnic7

PortGroup Name VLAN ID Used Ports Uplinks
iSCSI Port 4 0 1 vmnic7
iSCSI Port 3 0 1 vmnic6
iSCSI Port 2 0 1 vmnic5
iSCSI Port 1 0 1 vmnic4
 
Good host, still 6/7 down as expected.

~ # esxcfg-nics -l
Name PCI Driver Link Speed Duplex MAC Address MTU Description
vmnic0 0000:01:00.00 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e1:01 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic1 0000:01:00.01 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e1:03 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic2 0000:02:00.00 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e1:05 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic3 0000:02:00.01 bnx2 Up 1000Mbps Full 14:fe:b5:c9:e1:07 1500 Broadcom Corporation Broadcom NetXtreme II BCM5709 1000Base-T
vmnic4 0000:08:00.00 igb Up 1000Mbps Full 00:1b:21:ba:f6:4c 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic5 0000:08:00.01 igb Up 1000Mbps Full 00:1b:21:ba:f6:4d 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic6 0000:09:00.00 igb Down 0Mbps Half 00:1b:21:ba:f6:4e 1500 Intel Corporation 82576 Gigabit Network Connection
vmnic7 0000:09:00.01 igb Down 0Mbps Half 00:1b:21:ba:f6:4f 1500 Intel Corporation 82576 Gigabit Network Connection


~ # esxcfg-vmknic -l
Interface Port Group/DVPort IP Family IP Address Netmask Broadcast MAC Address MTU TSO MSS Enabled Type
vmk1 vMotion IPv4 10.1.2.11 255.255.255.0 10.1.2.255 00:50:56:79:8e:dc 1500 65535 true STATIC
vmk2 iSCSI Port 1 IPv4 10.1.1.14 255.255.255.0 10.1.1.255 00:50:56:77:13:4d 1500 65535 true STATIC
vmk3 iSCSI Port 2 IPv4 10.1.1.15 255.255.255.0 10.1.1.255 00:50:56:75:6b:16 1500 65535 true STATIC
vmk4 iSCSI Port 3 IPv4 10.1.1.16 255.255.255.0 10.1.1.255 00:50:56:76:1d:62 1500 65535 true STATIC
vmk5 iSCSI Port 4 IPv4 10.1.1.17 255.255.255.0 10.1.1.255 00:50:56:7c:4c:25 1500 65535 true STATIC
vmk0 Management Network IPv4 192.149.115.231 255.255.255.0 192.149.115.255 14:fe:b5:c9:e1:01 1500 65535 true STATIC


~ # esxcfg-vswitch -l
Switch Name Num Ports Used Ports Configured Ports MTU Uplinks
vSwitch0 128 16 128 1500 vmnic0,vmnic2,vmnic3

PortGroup Name VLAN ID Used Ports Uplinks
ProdVM 0 10 vmnic2,vmnic3
Management Network 0 1 vmnic0
vMotion 0 1 vmnic0

Switch Name Num Ports Used Ports Configured Ports MTU Uplinks
vSwitch1 128 9 128 1500 vmnic4,vmnic5,vmnic6,vmnic7

PortGroup Name VLAN ID Used Ports Uplinks
iSCSI Port 4 0 1 vmnic7
iSCSI Port 3 0 1 vmnic6
iSCSI Port 2 0 1 vmnic5
iSCSI Port 1 0 1 vmnic4
 
Why are all your iSCSI vmkernels on the same subnet? Only the first one will get used when you connect to your iSCSI SAN since it's the first in the routing table.

Log into a box or run the following via the VMA:

esxcfg-route -l

Notice that the 10.1.1.0/24 subnet is set to only use vmk2?

Each iSCSI interface out of your host should be on a separate subnet so you can multipath.

Can you also verify via the vSphere client that vmk0 on each host has "Management Traffic" checked?
 
That config is per Dell's setup guide. I know what you mean, but I followed their guide. Shrug.

More detail: The EQL arrays use that group IP for all connections. The 4 interfaces on the actual array don't get referenced. I do remember in even the MD3000 setup guides indicating each interface would need to be in its own address space - my SAN experience is limited but I do think this setup differs from other vendors in that respect. Some back end trickery.
 
Last edited:
That config is per Dell's setup guide. I know what you mean, but I followed their guide. Shrug.

More detail: The EQL arrays use that group IP for all connections. The 4 interfaces on the actual array don't get referenced. I do remember in even the MD3000 setup guides indicating each interface would need to be in its own address space - my SAN experience is limited but I do think this setup differs from other vendors in that respect. Some back end trickery.

Weird. I just read over Dell's Equalogic and VMware guide and it says just that. Must be something the EQL Load Balancer does to make all three paths work.

EDIT:

Putting some thought into this I suppose ESXi is smart enough to know that any bound vmk's should be used to attempt to contact the iSCSI target IP. If it fails to connect, don't display that vmk as a legitimate path. It would also explain why you cannot use any routing once a vmk is bound to the iSCSI adapter.

I'd still prefer to use multiple subnets for iSCSI MPIO, but since your EQL only has the 1 clustered IP this works, too.
 
Last edited:
Weird. I just read over Dell's Equalogic and VMware guide and it says just that. Must be something the EQL Load Balancer does to make all three paths work.

EDIT:

Putting some thought into this I suppose ESXi is smart enough to know that any bound vmk's should be used to attempt to contact the iSCSI target IP. If it fails to connect, don't display that vmk as a legitimate path. It would also explain why you cannot use any routing once a vmk is bound to the iSCSI adapter.

I'd still prefer to use multiple subnets for iSCSI MPIO, but since your EQL only has the 1 clustered IP this works, too.

Bound nics should be used to connect for storage traffic, unbound nics can do logins. EQL has to be set up this way, but you're missing a vmkernel port

Create a vmkernel port on the same vswitch, ~unbound~ to any individual nics (call it iSCSI Heartbeat), and let it float. It'll need to be vmk0, so you may also have to do some fiddling with the other vmkernel ports (remove 0, create the HB one, recreate whatever was 0 before), due to a networking code bug and an oddity with how the EQL arrays work. That port is for responding to ICMP traffic, as odd as it sounds >_<

give it an IP in the same range, of course.
 
Hey lopoetve,

EQL did note that vmk0 issue. However when they checked my config they did mention that having vmk0 be the management port was acceptable, because it kept it from being a valid kernel interface (which I believe is what you are referencing, that vmk0 should NOT be used for iSCSI on EQL for whatever reason) - however they didn't mention that the heartbeat HAD to be on vmk0.

Any idea if this would create the erratic behavior I've seen so far? This has been pretty crazy, honestly. I'm deathly afraid to reboot the second esx host and lose all visibility to all luns at this point.
 
Hard to say, to be totally honest. Wouldn't be hard to create to test and find out though.

Generally, that just causes problems with failovers, so... dunno :(

You could also nuke the iSCSI config and rebuild it, but I don't know if you feel comfortable doing that.
 
So I made a few changes, and got a little more drastic than intended. First, a powerconnect tech said in no uncertain terms that the 55xx series is absolute garbage under a certain firmware version. At their request I stopped all VMs on the LUNs and upgraded the stack firmware.

I also changed vmk0 over to vswitch1 on both hosts, using all adapters as active.

When the switches came back up after the firmware upgrade I had connectivity on both hosts again. Hard to say where the fault lies, but my gut says it had something to do with that switch firmware. That's based solely on their unwillingness to troubleshoot at all, and just wanting to upgrade it.
 
Why are all your iSCSI vmkernels on the same subnet? Only the first one will get used when you connect to your iSCSI SAN since it's the first in the routing table.

Log into a box or run the following via the VMA:

esxcfg-route -l

Notice that the 10.1.1.0/24 subnet is set to only use vmk2?

Each iSCSI interface out of your host should be on a separate subnet so you can multipath.

Can you also verify via the vSphere client that vmk0 on each host has "Management Traffic" checked?
worst advice, ever.
 
worst advice, ever.

No, that's accurate for many arrays actually - MD series, any other LSI Engenia based array, Clariion, NX series, etc.

For those, port binding causes a login problem based on how we do SIDs.
 
I was horribly sick yesterday - meant to respond to this.

Switch firmware was 4.0.0.3 IIRC, upgraded to 4.0.1.12. Like I said the powerconnect guy was pretty adamant that it was a firmware issue. The behavior I was seeing trying to ping from various points wasn't instilling a lot of confidence either.

Are there actually other arrays that are setup like the EQL series? I understood Child's advice, as I'm sure it applies to the majority of other arrays.
 
Not meaning to bash anyone in particular, but it's not particularly useful to make one line replies like 'worst advice ever'. If you know what you're talking about, presumably you can give a cliff notes explanation as to *why* you believe this.
 
I was horribly sick yesterday - meant to respond to this.

Switch firmware was 4.0.0.3 IIRC, upgraded to 4.0.1.12. Like I said the powerconnect guy was pretty adamant that it was a firmware issue. The behavior I was seeing trying to ping from various points wasn't instilling a lot of confidence either.

Are there actually other arrays that are setup like the EQL series? I understood Child's advice, as I'm sure it applies to the majority of other arrays.

I believe the HP Lefthand series uses a clustered IP for iSCSI. Been a while since I've worked on one.

http://h20195.www2.hp.com/v2/GetPDF.aspx/4AA3-0261ENW.pdf

But what do I know?
 
No, that's accurate for many arrays actually - MD series, any other LSI Engenia based array, Clariion, NX series, etc.

For those, port binding causes a login problem based on how we do SIDs.

specifically for the clariionss they fixed the bug a long time ago that required the multi subnet work around. everyone supports MPIO via vmware NMP now and have for a long time.
 
specifically for the clariionss they fixed the bug a long time ago that required the multi subnet work around. everyone supports MPIO via vmware NMP now and have for a long time.

Flare 31 only, last I checked. Might be Flare 30. But for anything older, still have to do it :)
 
Back
Top