• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

My networking woes

haxzorian

n00b
Joined
May 18, 2011
Messages
19
Let me hit the ground at a slow crawl. I'm not a network administrator. I stepped up into the role after the previous admin left and I realized that without someone who at least knew the vision and where to begin; my colleges would have busted out the witch doctor chants and animal carcasses, only to eventually burn the data center to a fine ashy substance. Now on to the point. Intermittently, however frequently at 4:00 AM on Tuesdays my critical enterprise services (Exchange 2010 and Veritas Cluster (AKA file server)) poops the bed. Normally, Exchange recovers without too much heartache and woe, but does present an inconvenience for my users the next morning... when they have to type their passwords to reconnect. Oh, the horror. Needless to say I take the heat, be it wrongfully placed in ignorance, that “my” network is refuse. So a little about the network, it’s comprised of Juniper switches at the core and top of rack, we have 8 Dell fabric switches. Our edge firewall, PCI firewall, and router are Cisco devices. There is a HA pair of F5 load balancers and various other Dell and IBM servers. The critical enterprise services are all virtual machines, being housed by two Dell blade chassis, using the Dell fabric switches. I’ve uplinked a server running Wireshark to the active 1 Gb Dell fabric with a constant capture running. I’ve so far been able to discern that when we experience an outage its due to packet loss on the heartbeat network. I was able to capture my Exchange systems receiving duplicate ACK packets and then subsequently firing off retransmits. From what I’ve read, that normally indicates some sort of network congestion. In conjunction with that, I have a NetApp connected to two of the core Junipers and on those interfaces I show MAC pause frames incrementing rather frequently. Which could lead me into a tangent about how that shouldn’t be happening given the amount of traffic pushing to and from the NetApp, but I won’t go there. I’ve actually got a case open with damn near all my vendors, who’ve had me try a ton of different things… which usually ends up to them pointing the finger at the other vendors.

In summary, I have a really weird network related issue occurring frequently. No real direction to point to what is causing it. I could really use some fresh ideas, clear eyes, troubleshooting advice, or even a ‘sucks to be you snarky comment’ for a good laugh.
 
Is the Exchange server restarting, or just losing network connectivity? Have you checked the events log on the Exchange server? Packet loss, when it is some small percentage, is usually network congestion or faulty wiring. 100% packet loss is usually a network segment down.
 
Its just losing network connectivity between the MS Cluster heartbeat as well as the iSCSI channel connected back to the NetApp. Which show up in the logs as a disconnect. But from the NetApp side, there is nothing that says we lost connectivity. I can only assume the MAC pause frames we see on the NetApp are to blame. But that just circles around to what is hammering the NetApp to cause it to say slow down.
 
Where are the user mailboxes, and what size are they? Is there some Exchange operation in the logs happening before it goes pear-shaped? An Exchange Server performing an operation on large user mailboxes can definitely hammer hardware.

When I get in the multi-vendor finger pointing game, I like to get them to give some concrete numbers/conditions for their product to be working correctly. When your product is designed to interoperate or act as fabric, you need to support everyone else's product(s).

I don't suppose you bought all this from a single vendor/contractor?
 
I don't know about you, but I performance metric everything I deploy and monitor everything like a hawk. I'd get some SMNP traps set and some monitoring software like ptrg, Nagios or at least cacti running. You should be able to narrow down the issue fairly quickly if you have everything online and being monitored.

Without something like a visio map its hard to understand where everything is and how it is connected up.
 
Sorry I kind of dropped the ball on this thread. Just to update anyone who still cares. Turned out to be an over provisioned Net App. Once we turned flow controls off the problem ceased to occur. After which we got on the phone with Net App and are now battling the imbalanced filers and looking at upgrading. Thanks to all who posted ideas and got me out of the burned up funk i was in to re analyze with a clear set of specs. :D
 
Back
Top