• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

YURT 1.0 Released

Status
Not open for further replies.
mikeblas said:
Fixed and added.

You're the man! Thanks.

How about adding a PPD for the protein Report?

http://yurt.blaszczak.com/ShowMachinesReportingProtein.asp?ProteinID=962

I really only care about it for the WU section, not frames. Allows me to compare my machine to other machines for a specific protein. I personally don't need the Version and CSD Version on this report either. Removing those might tighten it up a little.
 
mikeki said:
How about adding a PPD for the protein Report?
I want to normalize everything to PPD (that is, get rid of the places where "time to complete" is used -- it's the reciprocal of PPD), but the Protein Report page has a subtle problem or two that'll make it a bit harder to fix.
 
mikeblas said:
I want to normalize everything to PPD (that is, get rid of the places where "time to complete" is used -- it's the reciprocal of PPD), but the Protein Report page has a subtle problem or two that'll make it a bit harder to fix.

No problem. I understand these things get complicated when you consider accuracy is a goal.

Another question regarding the "Details for Machine" report. I'm trying to understand the Detailed Work Unit History section. After looking at it for a few of my machines it appears to reporting more WU's than a machine can possibly be doing.

For example on mikeblk it reports 31 WU's submitted since 6/27/2006. These WU's appear to require in excess of 20 days of compute time (I didn't add up the exact numbers from the report).

This machine is XP 1800+ and there is no way it crank these out that fast. I'm not at my office so I can't check the logs, but I doubt they go back that far anyway. It is also a dedicated folder so it never gets rebooted or used for anything else.

Under the Submitted Work Units section of this same report, it lists a total of 245 days worth of work. As this machine has only been on YURT for far less than that, this is troubling. It's been a dedicated folder for about 3 months, if YURT is somehow digging up stats on old units.

Looking at it even more, look at the bottom of the detailed work unit history section. Below is a list of all the 2505 proteins submitted by the machine above. Notice the number WU's that took exactly 1 d 18:25:06 to complete. There are also a bunch that took 1 d 07:31:53 and 1 d 08:03:49. I'm thinking these are all duplicates.

This same issue appears to exist on the Protein Report also. mikeblk is about 75% down the list and shows 32 2505's reported on that list.

This isn't specific to 2505's as it appears there are duplicate 1706 and 1704 proteins for this same machine. The same issue is appearing on other machines also. Alexspc started folding on 6/17 yet somehow as managed to submit 179 days worth of work.

My goal is to supply some detail to make it easier to explain and/or find an issue if there is one.

- Mike

Code:
Last Submitted Date	Protein ID	Points	Work Time	Points Per Day
			Days HH:MM:SS	
2006-06-28-22:09:19	2505	200	1 d 08:03:49	149.7
2006-06-28-22:09:18	2505	200	0 d 18:48:24	255.23
2006-06-27-15:18:17	2505	200	1 d 08:03:49	149.7
2006-06-27-15:18:17	2505	200	1 d 11:50:58	133.89
2006-06-23-20:01:03	2505	200	1 d 08:03:49	149.7
2006-06-23-20:01:03	2505	200	2 d 04:14:09	91.89
2006-06-20-10:42:21	2505	200	1 d 08:03:49	149.7
2006-06-20-10:42:21	2505	200	1 d 16:29:26	118.55
2006-06-19-19:42:46	2505	200	1 d 08:03:49	149.7
2006-06-19-19:42:46	2505	200	2 d 00:30:26	98.95
2006-06-18-11:42:42	2505	200	1 d 08:03:49	149.7
2006-06-18-11:42:42	2505	200	1 d 15:53:52	120.31
2006-06-17-23:15:10	2505	200	1 d 08:03:49	149.7
2006-06-14-20:37:10	2505	200	1 d 07:31:53	152.23
2006-06-14-20:37:09	2505	200	1 d 18:25:06	113.16
2006-06-13-11:44:08	2505	200	1 d 07:31:53	152.23
2006-06-13-11:44:07	2505	200	1 d 18:25:06	113.16
2006-06-10-03:08:50	2505	200	1 d 07:31:53	152.23
2006-06-10-03:08:49	2505	200	1 d 18:25:06	113.16
2006-06-08-11:39:11	2505	200	1 d 07:31:53	152.23
2006-06-08-11:39:10	2505	200	1 d 18:25:06	113.16
2006-06-05-13:56:23	2505	200	1 d 07:31:53	152.23
2006-06-05-13:56:22	2505	200	1 d 18:25:06	113.16
2006-06-04-04:51:32	2505	200	1 d 07:31:53	152.23
2006-06-04-04:51:31	2505	200	1 d 18:25:06	113.16
2006-06-03-03:59:09	2505	200	1 d 07:31:53	152.23
2006-05-24-20:02:54	2505	200	1 d 18:25:06	113.16
2006-05-21-21:52:02	2505	200	1 d 18:25:06	113.16
2006-05-20-04:33:15	2505	200	1 d 18:25:06	113.16
2006-05-20-04:33:03	2505	200	1 d 18:25:06	113.16
2006-05-19-15:54:15	2505	200	1 d 18:25:06	113.16
2006-05-19-12:54:20	2505	200	1 d 18:25:06	113.16
 
mikeki said:
Another question regarding the "Details for Machine" report. I'm trying to understand the Detailed Work Unit History section.
It's very simple: it is a list of every completed work unit the machine has reported over time. That's it.

mikeki said:
This machine is XP 1800+ and there is no way it crank these out that fast. I'm not at my office so I can't check the logs, but I doubt they go back that far anyway. It is also a dedicated folder so it never gets rebooted or used for anything else.
It's not possible for me to do anything without the logs from both YURT and FAH. Either the machine is doing the work faster than you thought, or the YURT client is getting it wrong when trying to observe what the FAH client is doing. I can't guess what's happening -- I need to see the logs to know for sure.

Obviously, I already have all the data that's at the website, so quoting it doesn't help. If the client is reporting duplicates, I need to see a debugging log and the FAH log to figure out why.

I'll see if I can check my own machines for evidence of the problem in the meantime.

mikeki said:
My goal is to supply some detail to make it easier to explain and/or find an issue if there is one.
No problem. If there's something wrong, I want to get it fixed. I just wish you had been involved in the beta program!
 
mikeblas said:
It's not possible for me to do anything without the logs from both YURT and FAH. Either the machine is doing the work faster than you thought, or the YURT client is getting it wrong when trying to observe what the FAH client is doing. I can't guess what's happening -- I need to see the logs to know for sure.

Hmmm, sounds like I need to run to my office for a bit. :)

I just wish you had been involved in the beta program!

Yeah, I wanted to be, but had two huge projects I was working at the time. Both are now finished.
 
I just e-mailed your yurt account both the yurt and fah logs from one of dual core machines that posted a duplicate very recently. One item of note is that I run my clients using verbosity level 9, so maybe that is generating extra information in the fah log file that causes the duplicates.
 
mikeblas,

I'm a bit confused by the Detailed Work History.
It shows a gob of protein 1163, which was the protein
folding when I started YURT for the first time. I know
that I haven't done 25 1163's since starting YURT,
and the Detailed Work History has a number of entries
at the same time.

Fold on!

 
Celerator,

I guess you're asking
me to debug what
you're seeing on the
website for you.

But I can't do that
unless you also include
the logs for both YURT
and the FAH client you're
running. If you have the
"debug level" YURT
logs, that helps the most.

As you can see from the
very recent exchange
between mikemi and
myself in this thread,
it sounds like your
question is very much
the same.

Please let me know if
it isn't the same, and if
you have a more specific
question. Otherwise,
I'll look forward to receving
your log files in email.

Thank
you!
 
mikeki said:
I just e-mailed your yurt account both the yurt and fah logs from one of dual core machines that posted a duplicate very recently. One item of note is that I run my clients using verbosity level 9, so maybe that is generating extra information in the fah log file that causes the duplicates.
Thanks, I'll have a look. Undoubtedly, using "verbosity level 9" will be related to the problem because non-default verbosity levels weren't tested. In fact, they weren't even known to me until after the beta.

Do you know of a document (a specification?) tha explains what the FAH will (or won't) write to the log file at the different versobisty settings?
 
mikeblas said:
Do you know of a document (a specification?) tha explains what the FAH will (or won't) write to the log file at the different versobisty settings?

Unfortuneately I don't know where one is. I just spent about an hour digging around for one and couldn't find anything. Only descriptions like verbosity 3 is the default and verbosity 9 is the maximum.

Since I'm naturally lazy, I use a simple one-click installation procedure I found here and it uses verbosity 9.

The official FAH pages just says verbosity 9 gives the most amount of data when submitting issues to Stanford.

However, I think the verbosity issue may be a red herring.

Looking at one of your machines it appears that Protein ID 772 has a few duplicates. All are exactly the same work time, which even for the same WU seems highly unlikely.

Code:
[COLOR=DarkOrange]2006-06-10-03:57:39	772	149	5 d 20:07:01	25.52[/COLOR]
2006-06-10-03:57:22	773	44	1 d 05:52:19	35.35
[COLOR=DarkOrange]2006-06-10-03:57:21	772	149	5 d 20:07:01	25.52[/COLOR]
2006-06-07-03:09:24	2304	48	3 d 17:14:16	12.91
2006-06-05-22:00:52	773	44	1 d 05:52:19	35.35
2006-06-04-11:29:21	2110	376	26 d 04:21:24	14.36
2006-06-04-11:29:20	1163	241	5 d 18:07:41	41.87
2006-06-04-11:29:20	2107	404	9 d 01:23:33	44.6
[COLOR=DarkOrange]2006-06-04-05:56:07	772	149	5 d 20:07:01	25.52[/COLOR]
2006-06-04-01:27:38	2107	404	16 d 02:46:27	25.07
[COLOR=DarkOrange]2006-06-04-01:27:38	772	149	5 d 20:07:01	25.52[/COLOR]

It also appears that this list contains 71 days worth of a work over approximately 6 or 7 days. Part of that may be an anomoly in the 2110 protein, but the duplicate issue appears to be popping up here.

My goal is to try an help narrow down the problem and it seems it would be far easier to debug if you had easy access to the rigs. That is the point in using the above box as an example.

I hope this helps.
 
mikeki said:
However, I think the verbosity issue may be a red herring.
Yeah, if the problem repros on one of my mahcines, then it probably isn't related to verbosity.

mikeki said:
I hope this helps.
Yeah, it does. Thanks for digging around! With the issu isolated, I should be able to figure out what the client is doing wrong.

I'm not sure how much time I can put into it; I have gotten sick (which always seems to happen to me on holiday weekends).

Looking at the logs on MIKEBLAS9, it seems like the problem is related to the high water mark. When YURT reads through the FAH log, it remembers how far it read. It needs to see if the log has changed (eg, been reset by FAH to a new file) or if the file is really just adding new data at the end.

In stuations where the extra work is identified, we appear to be detecting the file as reset when it really might not be. We end up parsing the whole file again, as if it was new, and that is causing the duplicate rows.

Unfortunately, the FAH work isn't easily uniquely identifiable. If it were, or if there were some documented way to observe the client from another process, we wouln't have all thse challenges.
 
If you are comfortable e-mailing me the client code, I'll take a look at it. It's been a while since I've done any real coding, but I can generally figure things out (assuming it's written in C, C++, VB, C# or some something similar).

Also, I'm out of town for two weeks starting Friday, I'll go dark for a bit starting then.

It sucks being sick on the holidays. I hope you feel better soon.

Edit: Also, I won't change any of the code. I will only suggest possible solutions to you via e-mail.
 
mikeki said:
For example on mikeblk it reports 31 WU's submitted since 6/27/2006. These WU's appear to require in excess of 20 days of compute time (I didn't add up the exact numbers from the report).

On MIKEBLAS9, I have a couple of good leads and I can pursue them.

Unfortunately, the logs you sent don't include activity for 6/27. They only cover a couple of days, and I don't get any clues to the problem you're reporting.

For now, the best way for you to help would be to get complete log information to me, if you still have the files lying around.
 
mikeki said:
It sucks being sick on the holidays. I hope you feel better soon.
No fucking kidding. I work my ass off, looking forward to a few days off the whole time. Then, the minute I relax, I blow a 101^F fever. Freaking great. On top of it, I pulled a muscle and can't even lie flat comfortably. I'm feeling better ... just in time to go back to work.

Fucking piece of fuck!

Anyway: if YURT is shut down, it writes its state to disk. When it starts up, it tries to find its state and reads it. If it reads it successfully, it uses that information to start where it left off in the FAHLog.TXT file.

If it can't read the state, there appears to be a bug. First, if the file ends up being not there, then it's not a bug -- this is really the first time that you've started YURT and we have to suck down the FAHLog.TXT file from the beginning.

If the file was there, then we try to read it and see if its content matches what we expect. If it doesn't match, then we call it "corrupt" and quit reading. Problem is, we don't discard what we read so far. The code then starts cold, meaning it reads the FAHLog.TXT file from the beginning. Since the FAHLog.TXT file actually was previously read and loaded into the cache (at least partially) from the state file, it ends up with two copies of the data. The queue fills, and we send two finishing records to the server in a row, bang bang.

Only a couple of users have reported the "Corrupt" message happening. It happened twice in the beta, and I thought I had the problem licked; after I made a related change, the issue went away, and there were no further reports, so I thought I was fine. One person has found it since, and I also saw it on MIKEBLAS9.

I think it's causing the duplicate reporting that you've seen. Problem is, I can't be 100% sure that's the problem because I've never figured out what causes the file to fail to read in the first place. I can't repro the problem with the state file that was lying around on MIKEBLAS9 (and I'm lucky to have even that, since MIKEBLAS9 was about to be decomissioned!)

The only clue that makes me not believe this explanation is that there is sometimes lots of time between the duplicate work units being reported. Two are reported together in MikeKi's example, and that fits the pattern. But then those two are reported again about seven days later. Maybe it's bad luck, maybe it's a different problem.

So, it would be great if everyone could have a look at their YURT log files and check to see if they find the "file is corrupt" message. If we can find a correlation between the machines with the error message in the log and the duplicated records, then I think we've found the cause of the problem.

Unfortunately, MikeKi says he's going out of town for two weeks, so I have to start from scratch and rely on everyone else to check their logs.
 
I'm pretty certain that I've got it figured out. If anyone wants to send along logs which have reported "corrupt" and a *.DAT file, I'll still entertain them, but after thinking about it (and after my fever broke) I'm positive that I know the main problem was.

Now I just have to figure out how to do the upgrade. I can post YURT 1.1, but since this issue invovles data corruption, I'll need to reset the statistics for everyone and not allow previous versions of the client to send their data to the database.
 
I just e-mailed you a complete set of logs for this machine.

Protein 2409 appears to be a duplicate.

Code:
2006-07-04-02:18:53	2409	600	2 d 20:16:25	210.92
2006-07-03-20:11:52	2409	600	2 d 20:16:25	210.92

There are numerous other duplicates in the logs also. To identify them there is an excel spreadsheet with the Detailed Work Unit History list sorted by PPD. This view clusters the duplicate WU's together and may help you choose which WU's you want to debug. This sheet is in e-mail to you also.

If you're going to release a version 1.1, I have one request that may or may not be possible. Can you make the dates and times in YURT match those in the FAH logs? This would make tracking WU's much easier. If this breaks your architecture, then NP. If it's relatively easy, then that may make things less confusing.

Also, you may be able to clean up the database with some form of de-dupping process. Basically remove all records that have the identical Protein ID and Work Time (leave the earliest one in there). I know this would be custom one off code that wouldn't be used again. I think it would be a select unique in SQL Server 4.2 (that's how old my experience is). :)

I hope this information helps and thanks for all your work on this.

I leave Friday afternoon for Two weeks, so I'm around a little longer.

- Mike
 
mikeki said:
I just e-mailed you a complete set of logs for
Thanks. I think these logs show the same problem. Each time that you restarted YURT, it read its state. For your machine, it is sucessfully reading the whole file -- but the state data is corrupt and that makes YURT think the file is new. So it re-scans everything. All the proteins in the re-scan get duplicated, since they were previously sent (at one time or another).

mikeki said:
If you're going to release a version 1.1, I have one request that may or may not be possible. Can you make the dates and times in YURT match those in the FAH logs? This would make tracking WU's much easier. If this breaks your architecture, then NP. If it's relatively easy, then that may make things less confusing.

FAH inexplicably uses UTC in the local log file. The YURT website does, too, so I don't have to fool around with tracking and adjusting user time zones. Since the log file is only viewed by the local human being at the local machine, I left it as local time. If you think ti's better to make it UTC to match FAH, it's trivial to do so. Nothing else reads it besides the user.

mikeki said:
Also, you may be able to clean up the database with some form of de-dupping process. Basically remove all records that have the identical Protein ID and Work Time (leave the earliest one in there). I know this would be custom one off code that wouldn't be used again. I think it would be a select unique in SQL Server 4.2 (that's how old my experience is). :)

You're thinking of SELECT DISTINCT. Unfortunately, that wouldn't get the earliest; it'll just make sure the tuple you get is unique.

Problem is, there's no reason that duplicates aren't allowed. It's possible (and even likely) that a work unit takes the same amount of time to execute in two different passes. It's safest to delete everything. I might delete all the duplicates (leaving none), or find a way to space out the duplicates measuring the reported work time to make sure the work units don't overlap. The problem with that approach is that we don't know when the work was done; it could've sat in queue with it's friends for weeks.
 
mikeblas said:
FAH inexplicably uses UTC in the local log file. The YURT website does, too, so I don't have to fool around with tracking and adjusting user time zones. Since the log file is only viewed by the local human being at the local machine, I left it as local time. If you think ti's better to make it UTC to match FAH, it's trivial to do so. Nothing else reads it besides the user.

The only time I've used it is to compare results with the FAH log. But that's just me. Your call. You've used it more than me.
mikeblas said:
You're thinking of SELECT DISTINCT. Unfortunately, that wouldn't get the earliest; it'll just make sure the tuple you get is unique.

That's what happens when I use 15 year old knowledge on a 3 day old problem. :)

mikeblas said:
Problem is, there's no reason that duplicates aren't allowed. It's possible (and even likely) that a work unit takes the same amount of time to execute in two different passes. It's safest to delete everything. I might delete all the duplicates (leaving none), or find a way to space out the duplicates measuring the reported work time to make sure the work units don't overlap. The problem with that approach is that we don't know when the work was done; it could've sat in queue with it's friends for weeks.

It's probably best just to nuke it. We'll get the data back soon enough. This is a long term project.
 
Over the weekend (probably Saturday morning), I'll be posting YURT 1.1. This should fix the problems with repetaed updates (with a few other little bugs, like the -1 protein ID problem).

The website now (even though you can't see it) records the client version whenever a work unit is submitted. This, coupled with the fact that we store work-unit historized data, could enable me to clean up the site for work units.

I'll delete all frames. Period; certainly nothing I can do about it, since individual frame progress isn't recorded. There's no way to tell which are duplicates and which aren't.

I'll also take a snapshot of the database and play with it locally. I might be able to weed-out duplicated values, but as of right now I'm convinced that it's not a good idea to try. If I delete too little, things are just as bad as they were. If I delete too much, the work rate of a given machine (and its processor, and whatever other stats we build) is too low -- the long-term average is affected.

It's kind of a bummer to tear down all the data, but it seems to be the right thing to do. Unless I'm struck blind by a huge insight, probably sometime Sunday, I'll delete the data submitted by 1.00 clients leaving only data from 1.10 clients.
 
When you download the 1.1 client, you can stop the 1.0 client, copy the 1.1 executable over it, and restart. The client will ignore any old (possibly corrupt) data file you've got lying about.

That's all there is to it. You don't need to reregister, you don't need to recreate your machine accounts, nothing.

Your 1.0 client's submissions will no longer be accepted.
 
Mike,

Sounds good to me!

Question: I've raised the overclock on my machine. When will the new values be indicated in YURT?

Also, you might want to repost the website URL when you get 1.1 ready for use.

Thanks for all your work,

 
Celerator said:
Question: I've raised the overclock on my machine. When will the new values be indicated in YURT?
Isn't this described in the manual?


Celerator said:
Also, you might want to repost the website URL when you get 1.1 ready for use.
Where?
 
I'm guessing he means in a new thread. Link to the main download and stats page. So people know that they have to install a new program. And it would be good to put in your previous post about how easy it is to just put the program over the old one.

 
Oh. Yeah, I was going to make a "YURT 1.1 Released" thread. The website will make it quite obvious that your machines aren't updating. (Even for people who don't read the manual, in fact.)
 
Read the manual? :eek: OK, I'll give it a look. :)

As for the website URL, I'm suggesting that periodically you include a link in your messages so that potential new users can easily visit, download, and start adding to the stats.

edit: OK, I read the manual and it has a line about running the YURT Installer to refresh one's registration information. I'll give it a try.

 
Worked! RTFM... :)

Mike, I believe that I'm the only one submitting stats for a Semperon 3300+. Take a look at the stats for that CPU (Work Rate by Processor Name) and I think you will agree that something is amiss. I've copied the line below:

AMD Sempron(tm) Processor 3300+ 1 13 1969.00 13 d 14:49:55 144.59 12900.00 54 d 19:52:27 235.28

Note the PPD of 235.28! I wish. I wonder if this has to do with the multiple entries issue? Will YURT 1.1 fix this? I want my stats to be helpful.

Thanks for your continued work on this project.

 
There's a couple more updates for me to take care of at the website, but it's safe to get back into the water. YURT 1.10 has been released. Please see the YURT 1.10 Released! thread for details and instructions on upgrading.
 
Status
Not open for further replies.
Back
Top