• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

VM SMP problems

APOLLO

[H]ard|DCer of the Month - March 2009
Joined
Sep 17, 2000
Messages
9,089
I'm experiencing unusual slowdown issues with my Linux SMP clients in VMs the past few days. The WUs are all A2 core types. Some VMs are processing at ordinary speeds but other ones on the same systems are twice as slow or far slower. One or two VMs were taking over an hour per step! :eek:

Has Stanford released new A2 WUs I'm not aware of that are much larger than previously, or is there something wrong with my systems? I deleted the WUs and restarted the VMs but it seems that these VMs will continue to process at unusually slow speeds even with a newly downloaded WU. :confused:

Any help would be appreciated.
 
I have not seen anything like that. Which units are you seeing this with? P2669?
 
I think the last one was a P2669 but I believe I saw the problem manifest with other WUs as well. Starting to wonder what the heck could be causing this....?

At the slowest rate of processing I witnessed so far, there's no way these WUs can make the deadline.
 
Is this happening on several PCs or just one? Have you tried getting rid of one of the VMs and making a new one to see if the issue goes away?

How about temperatures? Are any of your CPUs overheating? And you can also try disabling EIST and C1E just in case some weird issue is causing your CPUs to clock down.
 
Apollo - I tapped into my three today as well and noticed that they were runnning much longer frame times (i.e. more than 15 minutes). I didn't have time to check the WU numbers though, so I'm not sure what it was crunching on.
 
Is this happening on several PCs or just one? Have you tried getting rid of one of the VMs and making a new one to see if the issue goes away? How about temperatures? Are any of your CPUs overheating?
Yes, actually I shut one VM down and a console client to see if there was heat related issues, and no effect, unfortunately.

And you can also try disabling EIST and C1E just in case some weird issue is causing your CPUs to clock down.
Both are disabled as well as any other throttling settings in the BIOS. Since these are Supermicro boards, there's quite a lot of features like that and made sure I checked through all of them when I received the boards over a year ago. That's very good advice though.

Apollo - I tapped into my three today as well and noticed that they were runnning much longer frame times (i.e. more than 15 minutes). I didn't have time to check the WU numbers though, so I'm not sure what it was crunching on.
OK, something must definitely be up at Stanford. It seems that a lot of the WUs downloaded today are running much slower frame times, even though the project numbers haven't changed at all. There are some WUs running in the roughly ~30 min range, and then there's the obscenely long WUs running over one hour frames! :eek:

All these WUs still display the same designations, and makes one wonder if we are truly processing the same units and what the deadlines are. My previous times with the older units was also 15 minutes or less, so this is a big jump.

The fact that only one or two VMs are processing slowly per system with different motherboards in each system, and the fact that the frame times are the same across affected VMs, points towards new WUs. I will have to browse over at the official forum for additional information if there is any.
 
That is really weird. I've gone through I think about 6 2669 workunits today and I haven't seen anything like that. I hope you can find something on the F@H forum.
 
That is really weird. I've gone through I think about 6 2669 workunits today and I haven't seen anything like that. I hope you can find something on the F@H forum.
I'm going to search for info now and post anything I see. I'm concerned for meeting deadlines, I don't really care about points. What Axdrenalin posted has me thinking new WUs are in the wild. BTW, how many Linux SMP clients do you have running?

As a side note, the slowdown has now 'affected' a VM in a third system, again different config. This is definitely something different.
 
i 2 have the slow down. im running 2 VMa on my phenom and 1 on my laptop. the laptop went from 1000ppd to 200 :eek: and one of the VMs on my phenom whent from 1700ppd to 600 :eek:

i restart the laptop but im 45 minutes away from it and my phenom atm.

 
i 2 have the slow down. im running 2 VMa on my phenom and 1 on my laptop. the laptop went from 1000ppd to 200 :eek: and one of the VMs on my phenom whent from 1700ppd to 600 :eek:
OK, well I'm sort of relieved it's not only me and Ax, not that I wanted this to happen, LOL, but it does speak of something afoot. I'm looking through the FF ATM, so far nothing except a thread about the new A2 core v2.8. I wonder if this is somehow connected to our issues?

i restart the laptop but im 45 minutes away from it and my phenom atm.
I tried just about everything today, restarted client, restarted VM, restarted system, downloaded new WU, deleted and reinstalled client - nothing. :(
 
My three Linux SMP's are all set for -smp 8 to utilize all 8 cores, but as I mentioned these are at work on my servers and things were so busy that I didn't really have time to check them out. I just remember that I was seeing at least one 15 minutes checkpoint on each of them.

Find anything out over at the official forums yet?
 
I had a weird one the other day running at like 300ppd. I just thought it was a local issue with my machine or software. I ended up deleting the entire vm and building a new one. It downloaded a new WU and was back to normal. Maybe just a bad WU or corrupt VM or who knows... Sorry I don't have any more detail on the project number or anything.
 
My three Linux SMP's are all set for -smp 8 to utilize all 8 cores, but as I mentioned these are at work on my servers and things were so busy that I didn't really have time to check them out. I just remember that I was seeing at least one 15 minutes checkpoint on each of them.
Are you really getting better PPD than paired off cores per VM? I can't do that with my systems because they are all OC through windows using software, and if I install Linux as the only OS I will lose the OC, effectively dropping the clock speed to a level that might not make the preferred deadline.

Find anything out over at the official forums yet?
Not yet, but I only had time for a brief look. I'll check again later this evening.

I had a weird one the other day running at like 300ppd. I just thought it was a local issue with my machine or software. I ended up deleting the entire vm and building a new one. It downloaded a new WU and was back to normal. Maybe just a bad WU or corrupt VM or who knows... Sorry I don't have any more detail on the project number or anything.
I might have to reinstall some of my VMs the way it looks now. It's something that I rarely had to do before. I just find it weird that this occurred across several systems and all at roughly the same time.

I've been away from the forum much of this year, and it appears there's some catching up I need to do. I only discovered yesterday there was a new A2 core. I'm just wondering if there is some connection, and it's only affecting a few systems with specific WUs...??
 
Are you really getting better PPD than paired off cores per VM? I can't do that with my systems because they are all OC through windows using software, and if I install Linux as the only OS I will lose the OC, effectively dropping the clock speed to a level that might not make the preferred deadline.
I can't speak for him, but I had worse PPD running a single 4-core VM compared to running a pair of 2-core VMs. It was only a difference of a few hundred though.
 
Are you really getting better PPD than paired off cores per VM?

I can't speak for him, but I had worse PPD running a single 4-core VM compared to running a pair of 2-core VMs. It was only a difference of a few hundred though.

Sorry for the confusion on this. I'm not actually running under VM's on these three machines. I'm running native CentOS 5.3 installs that my Linux SMP console clients run under. I was contributing based on the issue with Linux SMP's slowing down on frame times without thinking about the VM side of things. ;) The servers are dual Xeon 5320 CPU's with a total of 12 Gigs of RAM per machine, so I wasn't too concerned with not meeting deadlines.
 
Sorry for the confusion on this. I'm not actually running under VM's on these three machines. I'm running native CentOS 5.3 installs that my Linux SMP console clients run under. I was contributing based on the issue with Linux SMP's slowing down on frame times without thinking about the VM side of things. ;) The servers are dual Xeon 5320 CPU's with a total of 12 Gigs of RAM per machine, so I wasn't too concerned with not meeting deadlines.
From what I read, the WUs obtained by using the -8 switch have a preferred deadline that grants the full value, but exceeding the initial deadline results in less points. Correct me if I have my information wrong regarding the points allotments, but I remember reading about these special WUs earlier in the summer.
 
I am seeing much the same on 2 of my VMs. I deleted the VMs and copied from two other working VMs and STILL have the same problem. I'm not really sure what the problem is. These particular machines have been running without issue for MONTHS!!!

 
I am seeing much the same on 2 of my VMs. I deleted the VMs and copied from two other working VMs and STILL have the same problem. I'm not really sure what the problem is. These particular machines have been running without issue for MONTHS!!!
Yeah, I know what you mean. It's troubling to read that reinstalling VMs didn't fix the problem.

Well, I reinstalled all the affected clients again last night and they are back to normal. I was going to delete the VMs but thought of wiping the clients with issues first, to see if the problem would go away. So far, they are working OK and it has been close to 12 hours.

Try wiping the clients off your VMs. Delete every F@H related file and folder in the affected VMs and reinstall the clients from scratch. It's possible that corrupted WUs were dispatched from Stanford and a few people received them. I really can't think of any other cause. I've never seen this happen since I started using VMs over a year ago.
 
check my SMP2 VM below

fprob.jpg


 
ALL4AMD,

Today, I just had a VM affected on yet another computer. I think this makes all the systems I run with VMs. The good news is the clients I wiped clean and reinstalled over the weekend are still running OK, thankfully. There must be some corrupt WUs circulating. It's the only explanation I can think of. For the time being, we just have to be vigilant and hope Stanford spots the problem and fixes it. :(
 
Well I did try to "clean install" on one of them and after two days, it was back to the same ole shit. I left this morning for MSP and I just shut it down. I'll deal with on Friday morning when I am back in my home office.

I'm leaning your way in thinking it is a bad WU, but can't confirm as I haven't been tracking what WU's I've been getting.
 
I shut it down. i have to go to work in a few hours. when i get home ill try a reinstal.

 
check my SMP2 VM below

fprob.jpg

When I first read that, coupled together with the SMP2 that Evil has been talking about some, I thought you had actually received an SMP2 WU for a minute. Then I went back and looked just to see that was what you *named* it on FAHmon. :D

Had me going there for a minute ALL4AMD...two hour frame times is crazy though. Are you running that on your sig machine or some other boxen?
 
When I first read that, coupled together with the SMP2 that Evil has been talking about some, I thought you had actually received an SMP2 WU for a minute. Then I went back and looked just to see that was what you *named* it on FAHmon. :D

Had me going there for a minute ALL4AMD...two hour frame times is crazy though. Are you running that on your sig machine or some other boxen?

sig

 
Hopefully sooner rather than later.

I don't have the time to babysit the SMP clients. My GPU2 Clients run hard and strong for weeks/months at a time. It is only of late that the SMP clients have had issue and only a couple of them, but it is those couple that require the most babysitting.
 
Hopefully sooner rather than later.

I don't have the time to babysit the SMP clients. My GPU2 Clients run hard and strong for weeks/months at a time. It is only of late that the SMP clients have had issue and only a couple of them, but it is those couple that require the most babysitting.

Agreed...the virtual machines put out much better ppd...but after averaging out all the times I have to delete them and create new ones due to various issues it's becoming a major PITA.
 
I think I caught one of these errant WUs from the start. Here's what happened. I noticed that a client in one VM had an EUE. Looking at the log.txt file for the client revealed that this WU was just received and crashed right away before completing the first step. After a few restarts that ended in further EUE's, I decided to delete the A2 core since it was the older v2.07. The client DL the new v2.08 A2 core and restarted the WU. There were no errors this time but after almost one hour, the client was stuck at 0%. :(

I can't say for certain that the new core is somehow connected but find it weird the older core just causes errors with some recent WUs, while the new core is super slow. I still think the best option is to wipe and reinstall the client. Reinstalling the VM takes longer, is more involved and simply not necessary. VMs are not the problem. I think it's almost certain now the culprit is WUs that either crash or process very slowly. Another thing I find unusual is only a few people seem to be reporting about this problem. :confused:

Agreed...the virtual machines put out much better ppd...but after averaging out all the times I have to delete them and create new ones due to various issues it's becoming a major PITA.
I realize I cannot speak for everyone because there aren't two machines built exactly the same, but I almost never had to delete and recreate a VM since I started running them over a year ago. The Linux clients installed in VMs are far more stable than anything I ever had to work with from Stanford in a Windows environment. If they can ever get the Windows SMP client to be as stable and productive, I'd consider dropping the VMs. Until then, I will continue using them knowing that every MHz of CPU power is being utilized to the maximum efficiency possible. :cool:
 
Agreed...the virtual machines put out much better ppd...but after averaging out all the times I have to delete them and create new ones due to various issues it's becoming a major PITA.
I realize I cannot speak for everyone because there aren't two machines built exactly the same, but I almost never had to delete and recreate a VM since I started running them over a year ago. The Linux clients installed in VMs are far more stable than anything I ever had to work with from Stanford in a Windows environment. If they can ever get the Windows SMP client to be as stable and productive, I'd consider dropping the VMs. Until then, I will continue using them knowing that every MHz of CPU power is being utilized to the maximum efficiency possible. :cool:
I have personally never had to recreate a single VM since I started using them, which was around the end of last year IIRC.
 
I have not had any problems. I get mostly 2677s which do about 9.5 minute frame times on my phenom II. That's the only SMP client I have running.
 
I have not had any problems. I get mostly 2677s which do about 9.5 minute frame times on my phenom II. That's the only SMP client I have running.
So far, I've had about a quarter of my clients affected by this problem, but when one runs many clients that's quite a few that needed to be deleted and reinstalled. Fortunately, it only takes about a minute to reinstall and restart the client. I haven't yet seen this problem resurface with the clients I reinstalled.
 
Ooh, I just got a 5101 that is running about 21:40 per frame. Looks like it is worth 2165 points so it's slower ppd than my usual units.
 
Ooh, I just got a 5101 that is running about 21:40 per frame. Looks like it is worth 2165 points so it's slower ppd than my usual units.
P5101 is an A1 unit so it will naturally provide lower PPD. What you're seeing isn't anything out of the ordinary.
 
I received a P5101 today on one of my clients as well. The A1 units are downloaded a couple of times a month on my systems. In the beginning of the year, if I remember correctly, Stanford was sending only A1 WUs and then they reverted to a 99% A2 scheme sometime by the end of winter. It all depends on what they need processing when.
 
Well I shut down the last remaining VM that was having problems last Monday. I restarted the VM this morning and it is chugging along fine.

At least for me, this proves it is a WU issue as the VM is just fine on a better WU. The problem is a lot of time is wasted on the "bad" WU and then on top of that, I don't meet deadlines so it is wasted time with NO points!!!
 
At least for me, this proves it is a WU issue as the VM is just fine on a better WU. The problem is a lot of time is wasted on the "bad" WU and then on top of that, I don't meet deadlines so it is wasted time with NO points!!!
I had yet another one early this morning. The same situation that I outlined in post #31. A client was running v2.07 A2 core with a new WU that would just crash and stop the client at 0%. I delete and download the new core. I restart the same WU with the new v2.08 A2 core. I check back about an hour and a half later, the WU has only completed one step in that time. I shut down, wipe and reinstall the client from scratch and everything works fine. I have a strong feeling I will need to go through this with every one of my remaining clients still running the v2.07 core. :(
 
Just had this happen to me. I nuked the VM and remade it; hopefully it'll be back to normal.
 
Back
Top