Official YURT Design Thread (Development Discussion ONLY)

Wow, what luck. About 20min after I posted a tree fell across the lines and I lost power for the rest of the night. Found out Verizon DSL is down to $15/month so I'm gonna try to get that this week so I can answer questions quicker.

Alright, down to bussiness.....

A little word of advice about the prot.dat file from EMIII, its an archive of ALL work units since EMIII has been around. Stanford has (had) a bad habbit of reusing old work unit names which causes a problem. If I submit unit P1234 worth 20 points 1 year ago, say Stanford changes unit P1234 to be worth 365 points. Now there are 2 units named P1234 worth 2 different point values in the database. If the program happens to grab the unit worth 20 points then the PPD and PPD/Ghz will be wrong. What I need to do in order to solve this is to get a PHP script that will parse the Psummary page and only add new units and if the names are reused then delete the old and use the new one. Since I havn't been around I havn't kept my database updated with the new units so this is why some units wernt found on it.

PPD is Points Per Day which is calculated server-side using the average time per frame the user/program reports

/Ghz is Points Per Day Per Ghz which is also calculated server-side using the PPD/Ghz

The easiest thing I can see to do to upload the data to the server would be to use a form. I can make a simple, basic PHP form that will be hidden from public view and will only be used by the program. It would be like a "back door" i guess. I'm not sure if this would be easier or not, but I should be able to setup an FTP account for the program to upload a .txt containing the data and have a PHP script read the file, add to the database and delete the file.

With the program, the idea to save the data in a text file before uploading to the database might be a good idea. What if the user looses internet? or if they only connectevery few days? FAH has trhe option when running timeless tinkers to download 10 work units and run them all before connecting again. If your on a broadband connection then the text file wouldn't be needed as much but i think it would be harder to decide if you needed the text file or not based on the connection than it would be just to do it. Maybe even only do it if no connection is detected, I think that would be best. Then check every XX time to see if there is a connection, if so send the data and delete the file.

"What if I add a folding client after the initial setup? do I just run the setup again? or so there some way to add it in while it's running?" I would think that there should be a way to config the program after its been setup. Even some thing like the -config flag used for FAH would be fine, just a simple command line, text only setup menu or something.

Data to be collected....
These are all the variables that are currently in my database and a description of each

ID - auto increment built into the database server-side
manuf - CPU manufacturer, ex: AMD Intel, collected by program
cpuid - CPU model, ex: Athlon 2100+, collected by program
core - CPU core, ex: Prescott, collected by program
skt - CPU socket, ex: 478, collected by program
speed - CPU speed in Mhz, ex: 1995, collected by program
nm - CPU manufacturing process, ex: 90, collected by program
multi - CPU multiplier, ex: 16, collected by program
L1 - CPU L1 cache, ex: 8, collected by program
L2 - CPU L2 cache, ex: 16, collected by program
L3 - CPU L3 cache, ex: 1024, colleted by program
fsb - system bus, ex: 200mhz, collected by program
ram - amount of system RAM in mb, 1024MB, collected by program
ramspeed - speed of RAM in mhz, ex: 133mhz, collected by program
cas - CAS or RAM, ex: 2-2-2-5, collected by program
dual - is the RAM in dual channal mode? (yes/no), collected by program
wu - current work unit from system, ex: 1234, read from logfile
cpuusage - CPU Usage, ex: 100%, read from config file
hrs - average number of hours the client runs, ex: 24, not sure how to get this, maybe have the user enter it on setup??
os - system operating system, ex: Win XP, collected by program
client - what FAH client is being used, ex: text-only, I think each client is named different so maybe we could use this name to tell what one it is?
user - username, ex: enhanced08dotcom, read from config file
frtime - average time per frame, ex: 00:15:34, not sure how to get this, may need to ask TheWeatherMan about that
date - date work submitted, ex: 03/12/06, written in to the PHP script, server-side
PPD - points per day of current unit, ex: 234, calculated using 'updated' work unit table and frtime, server-side
PPDgig - points per day per ghz, ex: 110, calculated using PPD and speed, server-side

I think I had mentioned it before but if not, the program CPUz (link) can be used and bundled with this program. CPUz can be run in "ghost" mode to create a text file of system info. All system info that is in my database can be found via CPUz. This could be an option if you want to go that way.

As far as EUE's (Early_Unit_End) go, I would not even bother with them. If one is detected then dont upload that data, it will just throw off the database. Maybe have some kind of popup or something that tells the user that there is an EUE. A lot of users dont watch the logfiles closly and as far as I know there is no program that will tell you if there has been an EUE, this could be helpful.

Beta units may also cause a problem. Maybe have the program check if its working on a beta unit or not. If it is then dont upload the data.

Thats about all I can remember of the questions that didn't get answered. I printed out this thread and will be going over it again. If there are anymore unanswered questions that I missed, ask away.

Thanks for keeping this thread alive while I have been away and thanks for the help! :)

Happy Folding!
 
I think we should track BETA proteins just the same..... they'll eventually be regular WUs, so why not track them? if a batch has alot of EUEs, then they'll be thrown out anyhow...

e08, take a look here.... we were trying to come up with a name for this project....


Keep on Folding!!

 
I dont know what Stanford would say about tracking them tho. I'll have to ask them about it in the Beta forum. The problem I see is that the points change for beta units quite often and most of the time they are re-benchmarked just before being released to the public which changes the points yet again. It may also cause people to try and join the beta team just for higher points. If a beta unit is shown on fah-database as being a high point unit then some people may try to join just for the higher points.
 
I guess that's true...

I didn't think about the re-benching....


Keep on Folding!!

 
enhanced08 said:
I think I had mentioned it before but if not, the program CPUz (link) can be used and bundled with this program.
I'm not too interested in taking a dependency on another program. We'll have to track it as it changes and update it, and so on.
 
I'll be ready for a couple of alpha testers this weekend, I think. (I don't even have all the functionality yet; I just want to have some people try the GUI "setup" program.)

Who can help?
 
definately.... just let me know how I can help


Keep on Folding!! For the [H]orde!!

 
OK, the installer program is ready for a few people to look at.

Here's the status:

1) It shouldn't crash. Error messages, sure; but it shouldn't fall over.

2) If you don't have a FAH UserID, it should complain about it. IF you do have one (eg, if you have the client installed) it should find it and let you run the program without complaint.

3) You should be able to create a list of FAH directories. Only directories that have FAHLog.txt should be allowed in the list.

4) You should be able to remove files from that list.

5) You should be able to send your machine's description and see it show up on http://beta.blaszczak.com/ShowListOfMachines.asp.

6) All buttons work -- you can copy the text, you should be able to navigate around the app.

The installer is done, in other words, unless I hear some bug reports or good suggestions. (Or, discover I need more features when turning my attention to finishing the progress reporter.
 
Sounds good... I'll give it a shot on all of my PCs if you like.... What do I need to do?


Keep on Folding!!

 
Check your mail, then see if you show up at the link above.
 
Checked... and done....... works well

This only runs on the local machine? or can I map a network drive and add that as a path?


Keep on Folding!!

 
Also curious as to how the network drive thing would work. Send me a copy at hardfolding@users.sf.net, please; I actually have a windows box, I promise!

Edit: Two comments: Microsoft VBScript runtime error '800a000d'

Type mismatch

/ShowListOfMachines.asp, line 117

and the MachineID would be better displayed as a hex value, so people can match it to their machine and make sure it's the right one.

 
The ASP error was fixed a little while ago. (Sorry; things are flakey here because they're widening the highway, and my power has gone out for a couple minutes every 10 minutes
over the last hour.)

The MachineID is displayed as a hex value, as well -- same update. The problem is that VBscript doesn't have an 8-byte integer type, so I have to play games with data conversion in T-SQL.

I'll send you an email as soon as I can.

It won't work over the network. The file change monitoring APIs are very flakey over shares, and I need to dig through the regsitry on the remote machine -- which would require credentials, and so on.

Please let me know if you can identify your machines in the list, and that the data looks reasonable/correct.
 
I have:

A
B
Foldingbox1
Foldingbox2

and everything seems okay on mine....


Keep on Folding!!

 
Foldingbox1 is an AMD rig. It has An Athlon with (K7 with the TH core) and is Stepping 1. Does it have a 256 KB L2 cache? a 64 KB L1 cache? Is that right?
 
unhappy_mage said:
Also curious as to how the network drive thing would work. Send me a copy at hardfolding@users.sf.net, please; I actually have a windows box, I promise!
Mail to that address bounces with some crazy "Postmaster verification failed while checking" error; and a little lecture about my not having a postmaster address at blaszczak.com. Do you have an account somewhere less draconian?
 
Hmm, I've never heard from anyone having problems with that address. Maybe their complaints just bounced? :p Guess it's time to start giving out a different one. Try willm1@umbc.edu. They're both forwarding to a gmail account; have you emailed anyone with gmail before and had a similar error?

 
mikeblas said:
Foldingbox1 is an AMD rig. It has An Athlon with (K7 with the TH core) and is Stepping 1. Does it have a 256 KB L2 cache? a 64 KB L1 cache? Is that right?

256K L2
64K L1
It's a Sempron 2200+ OCed a bit....

Everything looked right...


Keep on Folding!!

 
unhappy_mage said:
Hmm, I've never heard from anyone having problems with that address. Maybe their complaints just bounced? :p Guess it's time to start giving out a different one. Try willm1@umbc.edu. They're both forwarding to a gmail account; have you emailed anyone with gmail before and had a similar error?
Yep; gmail is fine. users.sf.net is what's sending the bounce. I'll forward it to you at the edu address.
 
OK, so here's what I want to work on. Not really in any order.

1) Fixing the filter, so we don't list drives on remote machines.

2) Add a ValidServers.TXT file hosted someplace, which will let you download a list of valid servers. That way, we can change hosts without changing binaries.

3) Finish up the parsing code in the monitoring program.

4) Get the monitoring program to post progress to the server.

5) Get the monitoring program to log.

6) Convert the monitoring program from a log to a service.

7) Get the monitoring program tested.

8) Write some queries.

9) Enhance the installer to detect physical and logical processors seperately. That is, cores and hyperthreaded cores.
 
mikeblas said:
1) Fixing the filter, so we don't list drives on remote machines.

Is the monitoring program going to track networked PCs? or do we have to have run the mon. program on each machine? If the mon. program tracks networked PCs, how do you gather the info needed without having odd privaledges?

What about Linux boxen with SAMBA shares? (like mage's foldserver and it's clients)?

mikeblas said:
2) Add a ValidServers.TXT file hosted someplace, which will let you download a list of valid servers. That way, we can change hosts without changing binaries.

Not a bad idea... didn't think about the possible host-change....

mikeblas said:
3) Finish up the parsing code in the monitoring program.

4) Get the monitoring program to post progress to the server.

5) Get the monitoring program to log.

6) Convert the monitoring program from a log to a service.

7) Get the monitoring program tested.

8) Write some queries.

9) Enhance the installer to detect physical and logical processors seperately. That is, cores and hyperthreaded cores.

Looks like a lot.... sorry that I'm not much help on that list other than #7......

...............................
We've got the name on the way..... and volunteers to make a few logos after we've decided the name.... Is there anything we're overlooking?




Keep on Folding!! For the [H]orde!!

 
OSUguy98 said:
Is the monitoring program going to track networked PCs? or do we have to have run the mon. program on each machine? If the mon. program tracks networked PCs, how do you gather the info needed without having odd privaledges?

What about Linux boxen with SAMBA shares? (like mage's foldserver and it's clients)?
Maybe it's better in Win XP and newer, but my experience with the file change notification APIs that they're unreliable over the network. See my previous response.

OSUguy98 said:
Not a bad idea... didn't think about the possible host-change....
Obviously, I can get SQL Server really cheap; but I can't commit the hardware or the bandwidth (to my personal connection at home!) for it. My hosting plan for blaszczak.com seems really itchy about database size and loading, so I'm not sure how long I can host it there.
 
Well, I'm getting nervous. Without web space, this is kind of a non-starter. So why am I still working on it?

I found where to get a new Emprotz.ZIP file. What are the fields in this guy? There are four; the first appears to be the protein number, the third is the number of points. The fourth is always 100. What is it? What's the second field?
 
example from emprotz.dat:

"2080"--------- the protein
2937600------- number of atoms? maybe?
226.0--------- point value
100---------- number of frames (usually 100.... sometimes 400 for tinkers)

I can't seem to find a match in atoms from stanford's psummary page for the second number...... none of the ones I checked work out........ so I dunno?

As for webspace, I guess we'll have to wait until this weekend for e08 to pop back in and answer that question....



Keep on Folding!!

 
mikeblas said:
Well, I'm getting nervous. Without web space, this is kind of a non-starter. So why am I still working on it?

I found where to get a new Emprotz.ZIP file. What are the fields in this guy? There are four; the first appears to be the protein number, the third is the number of points. The fourth is always 100. What is it? What's the second field?
If you look closely, not all of them have 100 in the fourth field. It's the number of frames, and you ought to be able to find some that have 400--those are the Tinkers.

 
mikeblas said:
Well, I'm getting nervous. Without web space, this is kind of a non-starter. So why am I still working on it?
I have some webspace available, but I'm not sure what databases are available for it. I'll check on it and get back to you.

 
OK; I've got some code that will "manually" insert frame completion data from the log into the database on the website, and I modified the ShowListOfMachines script to link to it. This is pretty fragile, but what's there works.

"foldingbox1" is the only one that has data. (And I think I transposed the time in seconds with the protein number somewhere along the line, but I'm too tired to fix it now.)

So the next step will be moving the "manual" log parsing code to the file update notification watcher project, and getting the update to happen in response to a file change.

There's no spec for what the website should look like; after no feedback from ya'll, I've gone ahead with my own site since I otherwise had no back end to inserting the data into and wouldn't have made any progress. I'll need to hear about what queries ya'll want to run soon, though I guess it doesn't make sense to do this until we've secured hosting.
 
Mike, that's crazy awesome! Sorry we haven't given you a spec on what to make things look like; the current style is cool as far as I'm concerned.

With regard to hosting, are you planning on using enhanced08's layout? It looks pretty comprehensive in terms of being able to search/sort results, and though we'll have to change a few fields around here and there I think having a full interface to the data like that makes it worthwhile to keep that code. What does the code that you're running on blaszczak.com need to run? Asp.net? MSSQL, I assume? Are you using any DB-specific features? I think most people have MySQL databases on their webhosting, so as far from a fan of it as I am, I think that's the easiest target. Just looking at a few at random (1&1, host298 (where my hosting is), prohosting.com) they all offer MySQL only. Would that suffice for the needs of the DB, and if not can you recommend a host that does have an appropriate DB system? I'd chip in for hosting, and I think a few other people would as well. If MySQL is good enough, host298 has a $20/year plan with fairly good bandwidth, unlimited usage, and good reliability.

Are you planning to release source when you finish with this project? I'm not begging you to, you're certainly not required to, but it'd make future maintainence possible if you aren't able to keep working on it in the future for whatever reason.

Early morning edit: Well, it doesn't appear to work on network drives, as promised - unless I'm running it in completely the wrong way. I ran the executable, selected the two folders to monitor, and hit send. Then I hit okay, and it disappeared. So I opened it again and let it sit while open. No frames have been reported, although a few have finished on both clients while it's been running. My fault, or expected behavior?

 
unhappy_mage said:
Mike, that's crazy awesome!
Thank you for your kind words.

unhappy_mage said:
Early morning edit: Well, it doesn't appear to work on network drives, as promised - unless I'm running it in completely the wrong way. I ran the executable, selected the two folders to monitor, and hit send. Then I hit okay, and it disappeared. So I opened it again and let it sit while open. No frames have been reported, although a few have finished on both clients while it's been running. My fault, or expected behavior?

The program you have configures the service. It does nothing more; you have the installer and you don't even have the service executable yet. I'm certain of that because I haven't even written it. In my previous note, I mention that I've got the pasring code working manually; I can run something and have it suck down one of the sample TXT files that OSUGuy gave me and shove it to the website.

The idea is that you'd run FAHInstaller.EXE once per machine to setup the service. Then, you'd run FAHUpdater.EXE as a service (though probably as a little console app initially -- easier to debug). FAHUpdater watches for directory changes in the dirs you've configured with FAHInstaller. It sees the changes, parses the file, and sends new records to the server.

So you've just written to the registry some information that the serivce, and when you press OK, you've left nothing running -- you're done running FAHInstaller and ready to run FAHUpdater. Or, more accurately, you're ready to start waiting for me to finish writing FAHUpdater.

And, of course, those programs need to get renamed pending the results of the renaming thread.

unhappy_mage said:
What does the code that you're running on blaszczak.com need to run? Asp.net? MSSQL, I assume? Are you using any DB-specific features?

It's plain ASP (and that's very bad code, because I'm not much for ASP development). If it was up to me myself, I would've written an ISAPI extension in C++.

I'm not yet using any SQL-specific features, but since my knowledge of MySQL isn't strong, and given MySQL's notoriously weak featureset, I probably am. For obvious reasons, I'd rather not use MySQL; it's very easy for me to get a SQL Server license, and I was thinking of running the server here at home. Then, it dawned on me that would mean my own home connection would get all this traffic, and I'm not sure that's acceptable. Plus, I might want to be using the machines for something else, and so on.

unhappy_mage said:
Just looking at a few at random (1&1, host298 (where my hosting is), prohosting.com) they all offer MySQL only.

blaszczak.com is at 1and1. I'm using their "Microsoft Developer Hosting Package", which includes a SQL Server database, thought it's limited to 200 MB and they have some dodgy wording sometimes. From their FAQ, for instance:

Please note that under
no circumstances may the database be used for log evaluation processes (e.g. ivw),
add-clicks, chat systems, banner rotations or any other applications that put high
loads on the database.

which is a bunch of weasel words; what does it really mean? A high I/O load, or a high CPU load? And so on. Will they close my account just because I did a GROUP BY? Wouldn't anything sitting on 200 megs of data start putting a dent in their server?

Meanwhile, their Unix Developer package offers MySQL, but it's got a lower file size limitation -- 100 megs, half what I've got for SQL Server.

The next step up from them is to get a dedicated server, which is prohibitively expensive. The cheapest package that supports SQL Server is $190 per month.

For two month's dedicated server fee from 1and1, I could get a couple of nice drives, grab a 2.4 GHz machine that's lying around here, and be ready to host. Getting it off my connection would mean colocation, then. I asked about Seattle colocation, as I'm not finding anything that's affordable -- it's all about twice the cost of 1and1.

unhappy_mage said:
Would that suffice for the needs of the DB, and if not can you recommend a host that does have an appropriate DB system? I'd chip in for hosting, and I think a few other people would as well. If MySQL is good enough, host298 has a $20/year plan with fairly good bandwidth, unlimited usage, and good reliability.
Thanks for offering. The thing is the database size, not bandwidth.

Sending one work unit or frame as being completed takes about 33 bytes and then an 8 byte response. (Not including headers.) I sent a batch of 71 items using 2322 bytes, 2322 / 71 == 32.7. The response is just "SUCCESS" (or "ERROR MACHINE 1234 NOT REGISTERED", and so on).

According to this thread, Team #33 is 4.5 years old with 2.8 million work units completed. Assuming everyone participates with their full load, 632,280 work units per year and 1730 per day. It's only 5 megs per month, then, of bandwidth for the work units. Unfortunately, it's 100 times (er, or 400 times, for tinkers?) that amount for frames, so maybe we should drop counting frames or come up with a different strategy.

Perhaps, at 5 megs/month, I can just host it at home and quit crying. The problematic part of commercial hosting is the database server; perhaps I should have the notifications sent to a database server here at home, then have it cough up static pages and send them to blaszczak.com a couple of times each day. Doing it at home makes bakup and maintenance very simple.

Another alternate strategy would be to collect all the updates at the public site, have my site here at home download them all and do the necsesary aggregations and computations, then upload either database data or static pages back to the public site.

The simplest might be to modify the client program to aggregate the frame work reports itself. Instead of reporting each frame completed and its time, it can collect a count and a total time: 35 frames took 3500 seconds, so we must be spending about 100 seconds per frame. That way, only one record eneds to be sent -- though it's not sent as often and isn't as exciting or "live", we still get partial work units into the database.

Doing some more envelope math, the 2.8 million work units isn't so bad. Each record in the completed work table has:

machine ID (8 bytes)
protein ID (4 bytes)
completed datetime (8 bytes)
work time seconds (4 bytes)

I don't think I need anything more than a clustered index, though that will depend on what queries you guys decide you want.

So I've got 24 bytes per record, plus some overhead; let's go crazy and call it 64 bytes per record. 2.8 million rows is 179 megs. The other table grows 100 times faster (again, assuming we're storing the ).

Another alternative there, then, is to have the database do the aggregation. Client #123 says it did one frame in 35 seconds for protein 11, so we create a row: {123, 1, 35, 11}. The same client later says it did another frame from protein 11 in 39 seconds, so we modify that row from {123, 1, 35, 11} to {123, 2, 74, 11}. This helps storage, but it don't help bandwidth.

unhappy_mage said:
Are you planning to release source when you finish with this project? I'm not begging you to, you're certainly not required to, but it'd make future maintainence possible if you aren't able to keep working on it in the future for whatever reason.
I don't know. My C++ code is pearly white, but my ASP code feels like the suxxor. Plus, I have to ask around about it at work. And if anyone was going to do "maintenance", they'd have to know all my account passwords at the hosting place, and I'm not down with that.
 
mikeblas said:
The program you have configures the service. It does nothing more; you have the installer and you don't even have the service executable yet.
Ah, okay. That makes sense.
mikeblas said:
it's very easy for me to get a SQL Server license, and I was thinking of running the server here at home. Then, it dawned on me that would mean my own home connection would get all this traffic, and I'm not sure that's acceptable. Plus, I might want to be using the machines for something else, and so on.
True, no need to DDOS yourself. I think the load problem would turn out to be the generated html pages - they could easily be a thousand times bigger than a frame report, so it doesn't take many to push a lot of data over an uplink that's maybe 1 mbit (?).
mikeblas said:
Thanks for offering. The thing is the database size, not bandwidth.
Database/disk usage is theoretically unlimited, too. And as far as I've been able to push the limits, they mean unlimited when they say it - we pushed ~30 gigabits for each of the last 3 months, no complaints or anything; we've pushed 5GB of data around on the site (hosting a bunch of pics for a thread in the Anime subforum), and like I said it's $20 a year. Tech support even responded within a few hours when we asked if it was okay to host said bunch of pics.
mikeblas said:
Another alternate strategy would be to collect all the updates at the public site, have my site here at home download them all and do the necsesary aggregations and computations, then upload either database data or static pages back to the public site.
That would indeed make things simpler. I think the problem would come with finding a way to automate the pushing back and forth of data in a reliable manner over an unreliable network.
mikeblas said:
The simplest might be to modify the client program to aggregate the frame work reports itself. Instead of reporting each frame completed and its time, it can collect a count and a total time: 35 frames took 3500 seconds, so we must be spending about 100 seconds per frame. That way, only one record eneds to be sent -- though it's not sent as often and isn't as exciting or "live", we still get partial work units into the database.
This would be good, but you'd want to do some nasty statistical analysis to discard outliers and such - what if someone turns their machine off overnight, so one frame appears to take 8 hours? That'd skew the results pretty badly unless it were discarded.
mikeblas said:
what queries you guys decide you want.
A topic for discussion. I think the one goal here is being able to pick several WU numbers and see a distribution of how long it took machines of a certain type to complete it. Here's my MSPaint version:
untitled1.PNG

The idea (in my head) is to keep track of how many machines report a certain time / mHz ratio. So if machine 1 is a venice at 2 gHz and reports a frame time of 300, and machine 2 is a venice at 3 gHz and reports a time of 200, those would count toward the same bar in the chart.
mikeblas said:
I don't know. My C++ code is pearly white, but my ASP code feels like the suxxor. Plus, I have to ask around about it at work. And if anyone was going to do "maintenance", they'd have to know all my account passwords at the hosting place, and I'm not down with that.
With regard to releasing source, I can understand the need to check with work people. Don't go getting into trouble over this ;)

If it's on DC-sponsored machines, I think several people would probably be entrusted with the password. Also, I know on host298 you can create multiple ftp user logins with different permissions; this means that you could create a group allowed to modify certain scripts, and give each user a seperate login name and password. Then you'd at least know who modified the script last based on the ftp logs or the ownership of the file.

 
If you looked at a log file of say, a tinker, the frame times would vary at minimum a few seconds each way... but most likely up to around 20-30 seconds especially if the user was browsing/etc/etc/etc....... If we set some "outlier" criteria, we could kick out most of the wild frames...... even if it was based on something like a std. deviation.... if you have 100 frames, and 97 of them range from 5min to 7min and the other 3 frames are in the 10s of minutes, then we should kick those out..... This would require parsing the entire log file, rather than just seaching for a start and finish time...... The other option, I guess, would be to discard the entire WU...... or a combo of both (if it's REALLY bad, just toss the WU, but if it's within reason, kick out the outlier frames)

I agree with mage, there are a number of people on the team who we could trust to maintain scripts/etc.... But in the end, those types of decisions are above my head, I know very little about the server-side of life......



Keep on Folding!! For the [H]orde!!

 
unhappy_mage said:
This would be good, but you'd want to do some nasty statistical analysis to discard outliers and such - what if someone turns their machine off overnight, so one frame appears to take 8 hours? That'd skew the results pretty badly unless it were discarded.
Actually, the FAHUpdater program notices that the machine was shut down. The client writes a startup/shutdown msesage into the log, and when it appears between the work item (or frame) starting and it ending, it isn't used for timing information.

If a frame or work unit is slow because of other activity on the system, there's nothing we can do about that and it just gets reported.

BTW, I guess some people run the folding client on super slow machines, and I don't think I can support that because there's no date information in the log. I don't see a way to tell one work unit takes more than a day.

unhappy_mage said:
With regard to releasing source, I can understand the need to check with work people. Don't go getting into trouble over this ;)
There's a lot of weird rules around here. Host298 sounds neat, by MySQL is a showstopper for me ... I don't even think I could spin it as "competitive research".
 
Even with the database size limits, I really doubt that everyone will participate with the info. I think its worth a trial run on something like 1and1 or godaddy hosting for it and revist the issues at the end of Q3 to see if their is an immediate danger of it getting out of hand. at 5 MB a month that gives at least 20 months est. before hitting the limit.
 
mikeblas said:
Actually, the FAHUpdater program notices that the machine was shut down. The client writes a startup/shutdown msesage into the log, and when it appears between the work item (or frame) starting and it ending, it isn't used for timing information.
Cool! That removes that area of uncertainty. But someone could still start playing a game and make completion of the frame drag out.
mikeblas said:
BTW, I guess some people run the folding client on super slow machines, and I don't think I can support that because there's no date information in the log. I don't see a way to tell one work unit takes more than a day.
If you're calculating WU time as the sum of the frame times, I think we'll be okay. One frame taking longer than 24 hours would be a *really* slow machine.

Just to make sure, you do account for frames that happen across midnight, right?
mikeblas said:
There's a lot of weird rules around here. Host298 sounds neat, by MySQL is a showstopper for me ... I don't even think I could spin it as "competitive research".
Yeah, I'd heard about the rules - my neighbor's son works for MS in some respect. host298 is unix only hosting - so I guess that means it doesn't work for you, even if they had say PgSQL? Perhaps the guy who wrote e08's php scripts could collaborate with you to do the web interface.

 
unhappy_mage said:
If you're calculating WU time as the sum of the frame times, I think we'll be okay. One frame taking longer than 24 hours would be a *really* slow machine.
Probably; I don't have a feel for it. My Opteron (dexter on the machines list) just got handed a work unit where each frame takes about 5 hours.

The parsing code reads lines. If It notes a work unit start message, it makes note of the time. If it notes a restart, it discards the start message time. If it finds a work unit finish message, and it still has a work unit start time, it calls it good and reports it to the server.

Almost the same thing for frames: If it notes a frame progress message, and that message indicates a full frame is starting, it will make a note of the time. It will discard the time if it sees a restart. If it finds another frame progress mesage and has a valid frame start time from before, it reports the work to the server and resets.

It'll also discard any of the times pending if it notices an EARLY_EXIT message (or something like that; it was in the files that OSUGuy sent).

unhappy_mage said:
Just to make sure, you do account for frames that happen across midnight, right?
Yes, because I figured we didn't want negative completion times. Hmm; maybe I can use the same event to adjust the current oustanding work unit start time.
 
Does anybody have a dual core (not hyperthreaded) machine?
 
I read most of that, skimmed the last part so I could post real quick here.... I have the server setup. The server has 9GB of space and unlimited bandwidth. I dont see a problem with bandwidth even if it was limited. The way I had thought about this was to have the client program comeup with an average time per frame like the program FahMon does. At the end of the WU the average time per frame would be sent to the server along with the system data. From what I understand, you want to send data after each frame? Or did I read this wrong? The way I see this, each time the client program uploads data it will only be sending 5kb or less. Even with 2000 work units finished a day, this is 10mb and 300mb a month. A single Linux ISO would be equal to this.

The server uses MySQL. I should be here till sunday, I can setup some kind of account on the server for you to upload data or whatever you need to do, let me know.

I did all the PHP (99% anyway) myself. I talked to a guy who said he would help me with it since I'm new to it, but I havnt heard from him in awhile.


I'm really sorry for not being around. I didn't plan on just handing off the idea onto you guys like this and I thank you for sticking with it..... I'm here for now, so lets talk!

EDIT: btw, I have a dually PII 500mhz system
 
enhanced08 said:
The server uses MySQL.
I work on the SQL Server team at Microsoft, so it's inappropriate for me to work on a project involving MySQL.

Pulling through some more of the numbers, I think ICE_9 is right; we can probably stay at beta.blaszczak.com and not have much of a problem—or, at least, not have a problem any time soon. My only conern is their weasel-wording about not using the server for anything busy.

enhanced08 said:
I'm really sorry for not being around.
In your absence, in the interest of making progress, I've made most of the decisions myself -- almost always after posting here. As you catch up in the thread, let me know if there's anything you feel strongly about and well see if it can be changed; but I've got most of the code written. It's now a matter of doing putting the monitoring client together and getting the website done.

enhanced08 said:
From what I understand, you want to send data after each frame? Or did I read this wrong?
Yes. I'll probably batch it up a little, though -- 20 frames or so per connection, maybe.
 
mikeblas said:
I work on the SQL Server team at Microsoft, so it's inappropriate for me to work on a project involving MySQL.

Hmmm.... I understand you side of this, my uncle works for Coca-Cola and they really dont like it when their employees buy Pepsi. What I see is that my host is paid for for the next year, I paid it in advance, If I do all of the database work would you still continue to help? I'm not even sure what you would need to do with the database anyway. How do you plan to send the data? It was mentioned to enter the data into a web form, is this still an option?

mikeblas said:
In your absence, in the interest of making progress, I've made most of the decisions myself -- almost always after posting here. As you catch up in the thread, let me know if there's anything you feel strongly about and well see if it can be changed; but I've got most of the code written. It's now a matter of doing putting the monitoring client together and getting the website done.

Thats fine, as far as I can tell, most everything is how I wanted it anyway.

mikeblas said:
Yes. I'll probably batch it up a little, though -- 20 frames or so per connection, maybe.

Why? Whats the idea behind this? I think it would be much easier to upload once at the end of the unit, also the fact that some are not connected 24/7.
 
enhanced08 said:
If I do all of the database work would you still continue to help?
The database work is the most interesting aspect of the project to me. (That, and studying the data once it's collected.)

enhanced08 said:
It was mentioned to enter the data into a web form, is this still an option?
I'd rather not rewrite the code.

enhanced08 said:
Why? Whats the idea behind this? I think it would be much easier to upload once at the end of the unit, also the fact that some are not connected 24/7.
This feature is actually what takes care of people who aren't connected all the time. If you're not connected, the completed work item will be added to the batch, to be sent when you're not connected. If you are connected all the time, it prevents connecting to send just one unit.

The cost of sending one unit or frame's data is the size of the HTTP headers plus about 33 bytes of payload data; then, the size of the HTTP response headers plus about 10 bytes of status and error info.

So sending one item six times is six times the total of all those sizes. Sending six items one time factors out the size of the headers round trip. It's not much difference, but increase the numbers to 1000 users and 20 items, and it starts to add up. Plus, it's necessary to help the users, anyway.
 
Back
Top