• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

VDI Datastore issues

Vader

Supreme [H]ardness
2FA
Joined
Dec 22, 2002
Messages
5,137
I'm having some issues with a particular VDI datastore. The current data store is:

1. 5 x 300GB FC Drives RAID5
2. Hosts are 2Gb FC multipath RR to my EMC Clariion CX4-240
3. View Desktops are XP SP3 running basic office and terminal emulation software
4. Average 60 Desktops on this datastore all users have between 2GB-3GB of memory.

We are having sporadic issues where in the middle of the day, the datastore Latency spikes up to around 50ms, hit's a very high queue, etc. This does not happen every day, and based on load/usage, it's sporadic. Monday is our heaviest day, with Friday being our lightest, but it could happen on Friday very easily or any day during the week.

It's happened twice over the last two weeks. Nothing has changed, the amount of users has been the same..etc VDI desktop IOPs haven't changed from the average daily I was seeing over the last couple of months or so.

The alarms show External i/o workload. We are not pusing DAT's or WSUS updates during this window etc. This is also beyond what I see during our small bootstorms, like first thing in the morning where I see it get higher amount of latency then normalize. This actually freezes up the VM's where they all have to manually reset and then it goes away and normalizes again. This is typically what I see in the morning:




When the issue occurrs, it stays at upwards of 70ms of latency..etc and just sits there until it freezes the VM's.

I have a ticket in with EMC at the moment to see if we have disk issues but they haven't found anything. VMware is also not seeing anything, of course, the problem is getting historical data so i'm going to setup a syslog server so that they can see if first hand.

Any input or thoughts on my next course of action?
 
Last edited:
Have you pulled NAR data off the array while this happens? I'd like to see that.
 
I didn't get any .NAR files when I ran the analyzer for a 24hr period, did get the .naz file and uploaded them to EMC. Will that help?

Unfortunately we aren't licensed for analyzer so couldn't get the nar files. When it occurrs again, i'll setup a perf analyzer to get a .naz file during the issue time period.

Update:

EMC got back to me after i provided them a 24hr analyzer run and they stated that we were hitting peak drive utilization on the RAID group associated with this datastore from 6PM to 10PM my time. I can understand that becuase that's when we run our DAT file updates and WSUS updates. They are recommeding I redistribute VM's to another datastore. I'm not buying this as of yet...again, i've already verified DAT and WSUS are not running during both of the times of issue occurrence.
 
Last edited:
Sounds right. Randomize those. Incredibly common problem. You can't run that many updates at the same time to disks and not expect them to cry.

Looks like the queue algorithm is throttling things to try and make it better, which will only make it really really slow.
 
We are randomizing WSUS i know for sure, the EPO DAT updates..i'll have to check, however, neither are scheduled during the time this happens...so I not betting that's the issue. There is some other underlying issue here.

FAST is coming when the VNX arrives..but that's not for another few weeks..unfortunately...;) It was either new VNX or upgrade the Clariion....the latter wasn't even considered..lol...6gb SAS, FAST VP...newer UI....no brainer..lol...

Ya know..I was just thinking to myself why the hell are we doing windows updates on Vm's in the first place. We should be doing it on our master image, then recompose..doh!!..lol

We are going to turn off Windows Updates on the VM's themselves and add the Master to our Update Weekend Schedule. I'll let you know if this has any impact.

EPO DAT updates are scheduled for 10:00PM.

Thanks fellas.
 
Last edited:
You're going to hate me for saying this, but screw raid 5 and go to raid 10. I've watched an 8 disk raid 10 group kick the crap out of a 12 disk raid 5 group on my old CX-500. I witnessed the same on a NetApp 2050a. Parity calculations add latency. Latency = pissed off users.

Hate to be the bearer of bad news, but being someone who runs a 250 desktop View deployment, just sharing what I have seen.
 
BTW- I actually ended up abandoning both the CX and the 2050a for VDI and went Nexenta running off of SSD's. I currently have 1ms average disk latency spiking as high as 2ms as reported by VMware. Full HA clustered solution for under 20 grand. Even then, the SSD's are RAID10. Note, I only needed about 2.5tb of useable space for my VDI pools.
 
Hopefully I won't have to worry about that with the new VNX w/FAST. I realize that you get a pretty significant write penalty with RAID 5, but my VDI deployment is very read heavy..around 70/30 read/write ratio.
 
Hopefully I won't have to worry about that with the new VNX w/FAST. I realize that you get a pretty significant write penalty with RAID 5, but my VDI deployment is very read heavy..around 70/30 read/write ratio.

You're forgetting that all writes are subpage writes with VDI. They suck ass. Even more so on RAID5. :)

Kills the CPUs on the arrays too.
 
Hopefully I won't have to worry about that with the new VNX w/FAST. I realize that you get a pretty significant write penalty with RAID 5, but my VDI deployment is very read heavy..around 70/30 read/write ratio.

FAST Cache will help a lot with VDI. But be stingy about what you assign FAST Cache to. Not all work loads benefit from it and you present more workload to the SPs with every Pool or LUN you give FAST Cache to.
 
Hi there - my first post, but I am a longtime lurker - love this forum. We have VNXs with FAST enabled, purchased solely for VDI. I am quite familiar with the challenges of VDI, having dealt with it now for nearly three years. We first ran it on our DMX3 and worked through FA overloading/sharing issues, then finally were able to isolate it on its own FAs, but could never dedicate separate spindles to it.

VDI by itself is a challenge, but VDI on VNX is even more involved. We have learned in the past week that EMC advises using thick LUNs on the VNX for anything that is performance-sensitve, i.e., VDI. Thin LUNs are far less efficient on the VNX because it allocates space in chunks of 1GB. FAST then moves the 1GB chunks around. As the array fills up, FAST is less and less able to move the 1GB chunks around. EMC is advising us to maintain 30% free space in the array for FAST to work effectively. Refer to kb article emc279316 on Powerlink for the recommendations on thin vs. thick LUNs. We've used thin LUNs with a lot of success on DMX4s and 3PARs so this was a bit of a surprise. Our VNX started with 20 SSD drives, five 15k drives, and five SATA drives. The SATA drives became so saturated in the past couple weeks since the array went over 90% consumed that Analyzer showed them to be 100% utilized, 24 hours/day, 7 days a week. I thought this was surely an anomaly, but it was not until we got the entire workload moved off to another array that I saw utilization go down below 100% on these drives; even with less than 10 VMs left on the VNX utilization was at 100% and avg. seek distance peaked at 1TB multiple times per day. Our SSDs were no more utilized than our 15k drives. No wonder users were feeling such pain. FAST cache was doing its thing, but I have not found a way in Analyzer to view FAST cache IOPs, etc., just the number of dirty pages in Unisphere.

We are planning to lay the pool out again with thick LUNs, then will move some of the workload back, but will make sure to leave more free space on it.

So... word of caution, make absolutely sure that the VNX Is sized correctly for your expected performance and space requirements and takes that 30% capacity buffer into account. Do not fill the VNX past that point. At 90% FAST is nearly crippled.
 
Thanks for the input! Will definately factor this in!
 
Back
Top