• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

monitoring my machines?

FLECOM

Modder(ator) & [H]ardest Folder Evar
Staff member
2FA
Joined
Jun 27, 2001
Messages
15,828
is there any application that could handle the number of machines i need to monitor (about 500 computers)

from what i have gathered EM wont be able to do this...

all the machines are on the same lan and i have administrative privilages on them so installing something isnt an issue, although it would be nice if it could happen remotely....

any suggestions?
 
Would FAHMon do that? It monitors the directories you put into it at the interval you set.
 
Not really.

I've given up long ago monitoring a high number of boxes. For me > 20 gets to be a PITA. I've always imaging and such (RD environment) so it's double PITA for me.

EM3 has serious issues for large numbers of boxes, > ~20 or so. EM3 has a limit depending on the OS, something for XP and more for 2K I think-- 50 max either way on the highest. What some folks do is run a few copies of EM3, breaks the instances up into 40 or so. What really gets annoying is all the false positives, basically ends up being a network error or something-- you need to click on it to shut it up. I think the latter versions stabilized the quantity issues I was having.

FAHMON is more stable I think, not as pretty but gets the job done pretty well from what I recall. It's been sometime since I've played with this one, frankly I'd start looking here.

For me, constant changing computer names makes hitting the shares a mess. I always ended up having to redo the monitoring app I just gave up. Now every few weeks I just RDP into the boxes and check. This is only ~ 30-50+ at times so it's not too bad.

YURT I don't think actually installs as a service (?), and the reporting side is not really conclusive to "is this box running" deal. I might be wrong, never really did much on YURT end.

Take 20-40 boxes and run 1x instance of FAHMON and 1x EM3 and try. You need shares on all boxes (admin share I assume) and you will use up some LAN side bandwidth.
 
Yurt does install as a service. And the report out will show you the last time it reported information to the server.

I think though adding 500 boxes to the data set might put Mike over his DB size requirements.:eek:
 
The biggest bitch I see is actually knowing/finding all 500 boxes. Lets say you find all the names (on domain pretty easy), set things up to hit the admins shares and all that. Your good to go...now someone unplugs a machine/images a machine what have you--- your stuck digging around to find that one box pissing off the monitoring program.

Across that many, I fear the monitoring programs will always be bitching about something driving a guys nuts.

Let us know how things go.
 
If Mike could build something into YURT that would alert you to a box that is down (hasn't reported results in a predetermined time frame) that would be awesome.
 
I cant use Yurt becuase it publishes the machine names...

I tried F@H LogStats and it just keeps crashing and using a whole lot of ram

im gonna try FAHMON
 
not really, issue is i got permission to run FAH on the machines on the condition that i didnt share the fact that i put FAH on all the machines here with anyone outside here... unfortunately the computer names give away identifiable information as to their location

and while i honestly dont think they would ever find it, or probably even care, im not willing to risk 99% of my production... would rather play it safe...

anyhow, FAHMon seems to be working pretty well, it updated pretty quickly

i was able to export a list of machines from Look@LAN so i just made a little app to put it in the proper "Name" "path" format that FAHMon wants so it took me 20 seconds to add 700 machines :)
 
Does this create a security issue for you?

yes, F boy brought that up in the development threads. I'm pretty sure Mike added a deal in there to hide machine names.

I always wanted a sinple thing, let me know if a given box has not returned a result in ~ 4-5 days or what not. Just really, really basic.
 
yes, F boy brought that up in the development threads. I'm pretty sure Mike added a deal in there to hide machine names.

I always wanted a sinple thing, let me know if a given box has not returned a result in ~ 4-5 days or what not. Just really, really basic.
I think that would be a nice feature.
 
Here's an idea that's kinda off-the wall:

1) Add a scheduled task to each image that sends a copy of fahlog.txt to whatever location you specify (FTP/LAN share/whatever) every 15/30/60 minutes. Alternatively, write a short app that does all the parsing and sends a much smaller amount of data to some location (protein, time started, progress %, for example)
2) Write a short webapp (jsp, php, asp, whatever) to parse all the data files and display the results.

Yeah, it sounds a lot like YURT, but you determine the destination point...
 
but those FAHlog.txt files can get fairly large... 150-300k a piece.... multiply that by the 700-some machines, and you've got a heck of a network load every 15/30/45/60 minutes....



Keep on Folding!! For the [H]orde!!

 
EM3 does do the HTTP publish deal, outputs it's display to a web site or something. Kidna got it working once.

Computer name changes screws up schedule tasks.
 
Here's an idea that's kinda off-the wall:

1) Add a scheduled task to each image that sends a copy of fahlog.txt to whatever location you specify (FTP/LAN share/whatever) every 15/30/60 minutes. Alternatively, write a short app that does all the parsing and sends a much smaller amount of data to some location (protein, time started, progress %, for example)
2) Write a short webapp (jsp, php, asp, whatever) to parse all the data files and display the results.

Yeah, it sounds a lot like YURT, but you determine the destination point...

All you would need is to add your box name to the line "CoreStatus = XX" and upload that to a central site.
That wont tell you if the boxen's not folding but will tell you of finished WU vs.early ends.

Maybe two other lines to look for could be .........
"Extra SSE boost OK. " <- for start of CPU client
"Starting GUI Server" <- for start of GPU client.

That way you would only record two lines with box name & times for each work unit.

Luck .......... :D
 
Just curious, why the interest in monitoring the boxes now after so long not doing so?
 
Yurt has the option of sending the Machine ID assigned by Stanford vs the one you have setup. It would be difficult from your end to decipher which one it was, but you would have one way of doing it.
 
Yurt has the option of sending the Machine ID assigned by Stanford vs the one you have setup. It would be difficult from your end to decipher which one it was, but you would have one way of doing it.

Its true, YURT can submit each box folding ID instead of the box name or anything like that. I'm sure you could write something small and execute it on each machine to install it. I don't think it even requires a reboot. I really think YURT is your best option, you'll never have issues with shared folders or names changing to worry about. Mike's site also lets you sort most of your stuff pretty well. You might do a trial run.
 
Just curious, why the interest in monitoring the boxes now after so long not doing so?

cause my stats are in the trash and my number of machines shouldent have changed

FAHMon seems to be working pretty well, i might look into YURT again now its been pointed out that you can use something besides the machine name
 
cause my stats are in the trash and my number of machines shouldent have changed

FAHMon seems to be working pretty well, i might look into YURT again now its been pointed out that you can use something besides the machine name

I didn't want to say it, but yea you've been sucking ass of late! :D

We all are down, my PPD is the poo. 212x series does affect PPD, not that they are undervalued it's just that the others are over values. I think your also doing -big packets (nut job) and they got reduced a touch (~ 50 points less).

My best advice is to block off an hour once or twice a week and look/act upon the monitoring stuff. If you (I) were to check often and try to fix all the little things one by one I'd go nuts. Let a few gather up, fix'em and repeat.

The only real problem I've seen with some regularity is the client goes into a crap state where is says it can't reach the AS. The box has no network issues or anything, the client just gets stuck. Once in that state only fix is service restart/boxen reboot.

Nightly reboots are great for unattended things. XP patches should catch them every few weeks anyways.
 
Yurt will also put it all under your user name. So you can easily find all your boxen in one place once they update. To my knowledge though, it dosen't do the SMP and GPU clients.
 
I didn't want to say it, but yea you've been sucking ass of late! :D

We all are down, my PPD is the poo. 212x series does affect PPD, not that they are undervalued it's just that the others are over values. I think your also doing -big packets (nut job) and they got reduced a touch (~ 50 points less).

My best advice is to block off an hour once or twice a week and look/act upon the monitoring stuff. If you (I) were to check often and try to fix all the little things one by one I'd go nuts. Let a few gather up, fix'em and repeat.

The only real problem I've seen with some regularity is the client goes into a crap state where is says it can't reach the AS. The box has no network issues or anything, the client just gets stuck. Once in that state only fix is service restart/boxen reboot.

Nightly reboots are great for unattended things. XP patches should catch them every few weeks anyways.

all my updates are done either via ghost (AI Packages usually) or our SUS/WUS/WSUS (whatever its called this week) server

i think i will try and check them like twice a week, mondays and fridays seem good, some dont survive the weekened, and want to make sure as many as possible are good to go for the weekend (when the machines are idle the longest)

YURT dosent seem like it will be easy to deploy since all the options are windows and such... unless im missing something?
 
The only real problem I've seen with some regularity is the client goes into a crap state where is says it can't reach the AS. The box has no network issues or anything, the client just gets stuck. Once in that state only fix is service restart/boxen reboot.

Nightly reboots are great for unattended things. XP patches should catch them every few weeks anyways.
QFT. A simple reboot may dramatically and quickly push your ppd back closer to where it should be.
 
QFT. A simple reboot may dramatically and quickly push your ppd back closer to where it should be.

well, i will push a reboot task this afternoon and see if it makes a difference
 
Taking a dump is right! Here's your PPM for the last 3 months. :eek:
....................NOV..............DEC.............. JAN
FLECOM...3,276,576......2,566,256.......1,643,930
 
Taking a dump is right! Here's your PPM for the last 3 months. :eek:
....................NOV..............DEC.............. JAN
FLECOM...3,276,576......2,566,256.......1,643,930

that's not exactly realitive since the boxes were off for almost a month in DEC/JAN

Maybe you should push a FOLD [H]ARDER task?
 
ya they were off for two weeks, last week of dec and first week of jan
 
There is still a drop there...not trying to be a pain, just going to your point that you are missing production somewhere.

yep, i have found a bunch of locked machines already from that infinate loop problem from a few months back
 
yep, i have found a bunch of locked machines already from that infinate loop problem from a few months back

no shit, specifically which problem? The core_a0 deal or a different one. I figured that all fixed themselves...
 
no shit, specifically which problem? The core_a0 deal or a different one. I figured that all fixed themselves...

yep the core_a0 thing, apparantly some didnt... oh well... the other problem i am having is some are getting stuck...

like FAH is still using 100% but its stuck at

Code:
[16:53:52] Folding@Home Gromacs Core
[16:53:52] Version 1.90 (March 8, 2006)
[16:53:52] 
[16:53:52] Preparing to commence simulation
[16:53:52] - Ensuring status. Please wait.
[16:54:09] - Looking at optimizations...
[16:54:09] - Working with standard loops on this execution.
[16:54:09] - Previous termination of core was improper.
[16:54:09] - Files status OK
[16:54:09] - Expanded 291025 -> 1461493 (decompressed 502.1 percent)
[16:54:10] 
[16:54:10] Project: 3039 (Run 5, Clone 754, Gen 3)
[16:54:10] 
[16:54:10] Entering M.D.
[16:54:30] (Starting from checkpoint)
[16:54:30] Protein: p3039_supervillin-03
[16:54:30] 
[16:54:30] Writing local files
[16:54:30] Completed 4000000 out of 5000000 steps  (80)

for like a day or so until i push a restart to that machine...

the other problem i am having is this one...

Code:
Launch directory: PATH
Service: PATH\FAH502-Console.exe
Arguments: -svcstart 

Launched as a service.
Entered PATH to do work.

[12:07:32] - Ask before connecting: No
[12:07:32] - User name: FLECOM (Team 33)
[12:07:32] - User ID: 5FE6E4183DEC4D2D
[12:07:32] - Machine ID: 1
[12:07:32] 

A potential conflict was detected:

Process 1656 is currently running and may also be a client with Mach. ID 1.
Program will now exit. Upon restart, this check will not be done -- 
you may wish to check that no client is currently running in
PATH before restarting.

but if i check the other instance of folding on that machine

Code:
[12:19:56] - Ask before connecting: No
[12:19:56] - User name: FLECOM (Team 33)
[12:19:56] - User ID: 5FE6E4183DEC4D2D
[12:19:56] - Machine ID: 2
[12:19:56] 
[12:19:56] Loaded queue successfully.
[12:19:56] + Benchmarking ...
[12:20:00] 
[12:20:00] + Processing work unit
[12:20:00] Core required: FahCore_78.exe
[12:20:00] Core found.
[12:20:00] Working on Unit 01 [January 19 12:20:00]
[12:20:00] + Working ...

so i dont know why the first one is complaining about another machine id 1...
 
I've had a couple of those process ID 1 already running. Its stupid, cause I can restart the process from the services menu and then it fires up just fine. It might be some kind of erratic startup issue where maybe both start too close together and get messed up somehow, who knows. It only happens on restart though from what I can tell. Doubt that helps any, other than letting you know that you aren't the first to run into that.
 
Welcome to the world of paying attention! Take it easy of yourself, these will drive you nuts. Frankly this is why I don't pay attention anymore.

a0= these should have fixed themselves...huh.

Stuck= can you post the whole log? Close the client then grab the log. Need to see a start point, and a "when you look at it" point. Generally my "stucks" are in getting work, not actually doing work.

Conflicts= yep, does that every now and again. See that on all sorts of borgs, stupid. Maybe Windows tried to light up 2x instances or something I dunno. Either way, the box does nothing.

I'm also seen corrupted WU's or something. Teh process used ~ 5% CPU and ALL available RAM. It's not actually doing anything, just brings the whole box down to a crawl. Posted 2x at official forums no one will touch that issue.
 
i got rid of all the a0's by pushing a restart command to all my machines... (some havent been rebooted in a long time)

the stuck ones, that is the whole log... if i get to one of those today i will cycle the client and see what the log sais...

i havent had any issues getting work thankfully... yay for gigantic dedicated interweb connection :D

well, all these machines have 2 instances, since they are dual core and i borged them before the SMP client came out...

but they are in seperate folders and one is setup as machine id=1 and one as =2... i know they are all like this becuase folding is in my ghost image and if it was broken every machine would be broken, not just 2 or 3...
 
I forgot to add: if possible with your network just do a nightly or weekly reboot via GPO or scheduled task or something. Me not all that domainish savy.

Get the bitches to reboot every now and again and odds are most issues will sort them selves out. Fixes a0/no getting work/conflicts

Don't forget, no SMP for windows.
 
I've seen the "A potential conflict was detected:" error in a number of cases.
1. Between two CPU clients.
2. Between a CPU & GPU client.
3. Between two SMP clients.
Also it can be either client shutting down.

So now after every reboot, I go through all the logs of boxen running two clients and check they've both restarted ok.
Its a PITA but ..........

Luck .......... :D
 
Do you use MOM2005 at all? I think you could use it to monitor the service, there would be alot of custom configuration but then it would email you whenever a client had an issue. Does F@H throw any event IDs when it goes south?
 
Back
Top