• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

SMP with Dual GPUs problem

Alpha736

Limp Gawd
2FA
Joined
Aug 17, 2004
Messages
224
I have a Q6600 a HD 4870 and an HD 4850 which I added a few weeks ago, in one machine. Each GPU client uses one core of my Q6600, but lately they have been spreading themselves across all cores... I have each GPU client set to "Slightly Higher" priority and "Do NOT lock cores to specific CPU" is turned off, although I have tried using the it with these options reversed.

My problem is that lately it seems that my SMP client has sorta gone wacko... It doesn't fully utilize even one core, and I noticed it is no longer running 4 instances of FahCore_a1.exe it is now running ONE instance of FahCore_82.exe. What I'm thinking is that my SMP client wasn't getting enough cpu time since I added the HD 4850, and has switched to a single uniprocessor core...

What do I need to get the most PPD out of this setup? Do I need to reconfigure my SMP client? I really don't want to resort to running anything in a VM, since this is a machine that I use as my workstation.

Thanks guys... :)
 
Have you changed the method in which you start your "smp" client? The same client will load as a single core client if it isn't run with the -smp deino flag or the other one (for 64 bit os) in the icon or command line.
 
Have you changed the method in which you start your "smp" client? The same client will load as a single core client if it isn't run with the -smp deino flag or the other one (for 64 bit os) in the icon or command line.

I haven't changed anything, I checked the service and it still is set to run as -smp. I'm on 64bit vista so I'm running the MPICH client.
 
I just reran -configonly and as I was going through the command line it didn't have any of my original variables remembered, but I set them as they were originally.

Now this is what I get in my logfile when I try to run the service...

--- Opening Log file [April 23 05:55:25 UTC]


# Windows SMP Console Edition #################################################
###############################################################################

Folding@Home Client Version 6.23 Beta R1

http://folding.stanford.edu

###############################################################################
###############################################################################

Launch directory: C:\FAH\SMP
Service: C:\FAH\SMP\FAH-SMP-MPICH-6.23
Arguments: -svcstart -d C:\FAH\SMP -smp -verbosity 9 -d "C:\SMP\FAH"

Launched as a service.
Could not enter C:\FAH\SMP! Working in launch directory.

[05:55:25] - Ask before connecting: No
[05:55:25] - User name: Alpha736 (Team 33)
[05:55:25] - User ID: 412E095C6439F10D
[05:55:25] - Machine ID: 1
[05:55:25]
[05:55:25] Loaded queue successfully.
[05:55:25]
[05:55:25] + Processing work unit
[05:55:25] Work type 82 not eligible for variable processors
[05:55:25] Core required: FahCore_82.exe
[05:55:25] Core found.
[05:55:25] Using generic mpiexec calls
[05:55:25] - Autosending finished units... [April 23 05:55:25 UTC]
[05:55:25] Trying to send all finished work units
[05:55:25] + No unsent completed units remaining.
[05:55:25] - Autosend completed
[05:55:25] Working on queue slot 04 [April 23 05:55:25 UTC]
[05:55:25] + Working ...
[05:55:25] - Calling 'mpiexec -np 4 -channel auto -host 127.0.0.1 FahCore_82.exe -dir work/ -suffix 04 -checkpoint 15 -service -verbose -lifeline 1500 -version 623'

[05:55:26]
[05:55:26] *------------------------------*
[05:55:26] Folding@Home PMD Core
[05:55:26] Version 1.03 (September 7, 2005)
[05:55:26]
[05:55:26] Preparing to commence simulation
[05:55:26] - Looking at optimizations...
[05:55:26] - Files status OK
[05:55:26] - Expanded 20658 -> 130596 (decompressed 632.1 percent)
[05:55:26]
[05:55:26] Project: 4615 (Run 35, Clone 38, Gen 42)
[05:55:26]
[05:55:26] Assembly optimizations on if available.
[05:55:26] Entering M.D.
[05:55:44] OK
[05:55:44] - Expanded 20658 -> 130596 (decompressed 632.1 percent)
[05:55:44]
[05:55:44] Project: 4615 (Run 35, Clone 38, Gen 42)
[05:55:44]
[05:55:44] Error: Could not write local file. Exiting.
[05:55:49] - Shutting down core
[05:55:49]
[05:55:49] Folding@home Core Shutdown: FILE_IO_ERROR
 
Arguments: -svcstart -d C:\FAH\SMP -smp -verbosity 9 -d "C:\SMP\FAH"

Needs to be -svcstart -d C:\FAH\SMP -smp mpich -verbosity 9
 
It was working before without the -smp MPICH flag... I'll try it out though...
 
When I add the -smp mpich flag it says "invalid parameters," I'm pretty sure it's just -smp that I need to type. I did just notice that I targeted an invalid directory though, but that still doesn't solve the original problem.
 
I've got the core running again, but it's still running FahCore_82.exe. Do I need to delete my work, queue, or cores?

This is what my log looks like now on a run...

--- Opening Log file [April 23 07:46:10 UTC]


# Windows SMP Console Edition #################################################
###############################################################################

Folding@Home Client Version 6.23 Beta R1

http://folding.stanford.edu

###############################################################################
###############################################################################

Launch directory: C:\FAH\SMP
Service: C:\FAH\SMP\FAH-SMP-MPICH-6.23
Arguments: -svcstart -d C:\FAH\SMP -smp -verbosity 9 -d "C:\FAH\SMP"

Launched as a service.
Could not enter C:\FAH\SMP! Working in launch directory.

[07:46:10] - Ask before connecting: No
[07:46:10] - User name: Alpha736 (Team 33)
[07:46:10] - User ID: 412E095C6439F10D
[07:46:10] - Machine ID: 1
[07:46:10]
[07:46:10] Loaded queue successfully.
[07:46:10]
[07:46:10] + Processing work unit
[07:46:10] Work type 82 not eligible for variable processors
[07:46:10] Core required: FahCore_82.exe
[07:46:10] - Autosending finished units... [April 23 07:46:10 UTC]
[07:46:10] Trying to send all finished work units
[07:46:10] + No unsent completed units remaining.
[07:46:10] - Autosend completed
[07:46:10] Core found.
[07:46:10] Using generic mpiexec calls
[07:46:10] Working on queue slot 04 [April 23 07:46:10 UTC]
[07:46:10] + Working ...
[07:46:10] - Calling 'mpiexec -np 4 -channel auto -host 127.0.0.1 FahCore_82.exe -dir work/ -suffix 04 -checkpoint 15 -service -verbose -lifeline 2952 -version 623'

[07:46:10]
[07:46:10] *------------------------------*
[07:46:10] Folding@Home PMD Core
[07:46:10] Version 1.03 (September 7, 2005)
[07:46:10]
[07:46:10] Preparing to commence simulation
[07:46:10] - Ensuring status. Please wait.
[07:46:10] - Looking at optimizations...
[07:46:10] - Working with standard loops on this execution.
[07:46:10] - Previous termination of core was improper.
[07:46:10] - Going to use standard loops.
[07:46:10] - Files statusProtein: p4615_T0_BBL-16_minout
[07:46:10]
[07:46:10] Completed 250000 out of 2500000 steps (10%)
[07:46:10] Writing local files
[07:46:10] Completed 275000 out of 2500000 steps (11%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 300000 out of 2500000 steps (12%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 325000 out of 2500000 steps (13%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 350000 out of 2500000 steps (14%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 375000 out of 2500000 steps (15%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 400000 out of 2500000 steps (16%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 425000 out of 2500000 steps (17%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 450000 out of 2500000 steps (18%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 475000 out of 2500000 steps (19%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 500000 out of 2500000 steps (20%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 525000 out of 2500000 steps (21%)
[07:46:10] Writing checkpoint files
[07:46:10] Writing local files
[07:46:10] Completed 550000 out of 2500000 steps (22%)
[07:46:10] Writing checkpoint files
 
I restarted the service and it started running 4 instances of FahCore_a1.exe but it's still running 1 instance of FahCore_82.exe. This is kinda weird, I wonder if it'll stop running that core once it finishes it's work unit?
 
Okay, I restarted one more time and now it's not running FahCore_82 anymore, I have no idea what was up with that...

So now that I have all that out of the way (For now at least), can someone tell me how to configure these clients to give the best PPD out of this machine?

Currently I have both GPU clients configured to be slightly higher priority than the SMP client. I also changed it so that the GPU clients aren't locked to a single core, It seems like each GPU client is using about 25% of each CPU core. So that's about half of the CPU time outta my Q6600.

I have the SMP client running 4 instances of FahCore_a1, 2 of them are using about 25% each of my Q6600 and the other 2 are sitting at 0%.

Since 2 of my SMP cores aren't getting any cpu time I'm wondering if I should just scrap the SMP client all together and just run 2 uniprocessor clients.

Any ideas?
 
ur shortcut target should look like "C:\fah\Folding@home-Win32-x86.exe -smp"

may be u can help me and let me know how to limit the smp to 2 core so that my gpu clients can have the other 2 cores.

 
ur shortcut target should look like "C:\fah\Folding@home-Win32-x86.exe -smp"

may be u can help me and let me know how to limit the smp to 2 core so that my gpu clients can have the other 2 cores.


It's running as a service, it doesn't use a shortcut, but the service has the SMP parameter set. It is definitely running in SMP mode, if I shutdown my GPU clients all 4 FahCore_a1 will run at 25%(of the whole CPU), meaning each FahCore_a1 will run at 100% on it's own CPU core.
 
Since it is running a Uni WU, it won't switch over to SMP until it finishes it. You can delete the WU, queue.dat, unitinfo.txt, and your "Work" folder. When you start (with the -smp flag), it will pick up a new SMP WU.

 
Since it is running a Uni WU, it won't switch over to SMP until it finishes it. You can delete the WU, queue.dat, unitinfo.txt, and your "Work" folder. When you start (with the -smp flag), it will pick up a new SMP WU.
It wouldn't be using the a1 core if it was running a uniprocessor unit though, which means that now it's back to SMP.
 
Back
Top