• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Multi-GPU help

nitrobass24

[H]ard|DCer of the Month - December 2009
2FA
Joined
Apr 7, 2006
Messages
10,477
I have followed the guide and have the first instance setup and working properly.

When I start the second instance it starts but no progress is made.

Heres the log:

[06:34:35] Project: 5756 (Run 14, Clone 284, Gen 20)
[06:34:35]
[06:34:35] Assembly optimizations on if available.
[06:34:35] Entering M.D.
[06:34:42] Working on Protein
[06:34:45] Client config found, loading data.
[06:34:45] Starting GUI Server
[06:35:45] Completed 1%
[06:35:45] mdrun_gpu returned
[06:35:45] NANs detected on GPU
[06:35:45]
[06:35:45] Folding@home Core Shutdown: UNSTABLE_MACHINE
[06:35:49] CoreStatus = 7A (122)
[06:35:49] Sending work to server
[06:35:49] Project: 5756 (Run 14, Clone 284, Gen 20)
[06:35:49] - Read packet limit of 540015616... Set to 524286976.
[06:35:49] - Error: Could not get length of results file work/wuresults_05.dat
[06:35:49] - Error: Could not read unit 05 file. Removing from queue.
[06:35:49] EUE limit exceeded. Pausing 24 hours.
 
I have a GTX 295
SLI has been disabled
I already have the desktop extended.
 
In the copied folder for the second GPU, did you delete the "work" subdirectory and the queue.dat file? I don't think that's mentioned in the [H] guide, but if you already had work for the first GPU when you copied the directory, you don't want the second GPU to be working on the same WU.
 
This snippet :

[06:35:45] Completed 1%
[06:35:45] mdrun_gpu returned

Mean that it's working fine and you probably got a bad unit. Try to flush queue.dat and the work folder then check if you get a good wu.

A bad configuration will not even show 1% and fail immediatly. This often happen when the driver is not installed properly or there is different version for each GPU. In that case, nuke the drivers from orbit, reboot, use Driver Cleaner, reboot then install the right one (don't ever let windows pick the driver for you).
 
It keeps screwing up

I copied the folder before I ran for the first time.


Code:
Launch directory: C:\Users\Stephen\AppData\Roaming\FAH2
Arguments: -gpu 1 

[15:54:50] - Ask before connecting: No
[15:54:50] - User name: Nitrobass24 (Team 33)
[15:54:50] - User ID: 5196C407075E61D4
[15:54:50] - Machine ID: 5
[15:54:50] 
[15:54:50] Work directory not found. Creating...
[15:54:50] Could not open work queue, generating new queue...
[15:54:50] Initialization complete
[15:54:50] - Preparing to get new work unit...
[15:54:50] + Attempting to get work packet
[15:54:50] - Connecting to assignment server
[15:54:50] - Successful: assigned to (171.67.108.11).
[15:54:50] + News From Folding@Home: GPU folding beta
[15:54:50] Loaded queue successfully.
[15:54:51] + Closed connections
[15:54:51] 
[15:54:51] + Processing work unit
[15:54:51] Core required: FahCore_11.exe
[15:54:51] Core not found.
[15:54:51] - Core is not present or corrupted.
[15:54:51] - Attempting to download new core...
[15:54:51] + Downloading new core: FahCore_11.exe
[15:54:52] + 10240 bytes downloaded
[15:54:52] + 20480 bytes downloaded
[15:54:52] + 30720 bytes downloaded
[15:54:52] + 40960 bytes downloaded
[15:54:52] + 51200 bytes downloaded
[15:54:52] + 61440 bytes downloaded
[15:54:52] + 71680 bytes downloaded
[15:54:52] + 81920 bytes downloaded
[15:54:52] + 92160 bytes downloaded
[15:54:52] + 102400 bytes downloaded
[15:54:52] + 112640 bytes downloaded
[15:54:52] + 122880 bytes downloaded
[15:54:52] + 133120 bytes downloaded
[15:54:52] + 143360 bytes downloaded
[15:54:52] + 153600 bytes downloaded
[15:54:52] + 163840 bytes downloaded
[15:54:52] + 174080 bytes downloaded
[15:54:52] + 184320 bytes downloaded
[15:54:52] + 194560 bytes downloaded
[15:54:52] + 204800 bytes downloaded
[15:54:52] + 215040 bytes downloaded
[15:54:52] + 225280 bytes downloaded
[15:54:52] + 235520 bytes downloaded
[15:54:52] + 245760 bytes downloaded
[15:54:52] + 256000 bytes downloaded
[15:54:52] + 266240 bytes downloaded
[15:54:52] + 276480 bytes downloaded
[15:54:52] + 286720 bytes downloaded
[15:54:52] + 296960 bytes downloaded
[15:54:52] + 307200 bytes downloaded
[15:54:52] + 317440 bytes downloaded
[15:54:52] + 327680 bytes downloaded
[15:54:52] + 337920 bytes downloaded
[15:54:52] + 348160 bytes downloaded
[15:54:52] + 358400 bytes downloaded
[15:54:52] + 368640 bytes downloaded
[15:54:52] + 378880 bytes downloaded
[15:54:52] + 389120 bytes downloaded
[15:54:52] + 399360 bytes downloaded
[15:54:52] + 409600 bytes downloaded
[15:54:52] + 419840 bytes downloaded
[15:54:52] + 430080 bytes downloaded
[15:54:52] + 440320 bytes downloaded
[15:54:52] + 450560 bytes downloaded
[15:54:52] + 460800 bytes downloaded
[15:54:52] + 471040 bytes downloaded
[15:54:52] + 481280 bytes downloaded
[15:54:52] + 491520 bytes downloaded
[15:54:52] + 501760 bytes downloaded
[15:54:52] + 512000 bytes downloaded
[15:54:52] + 522240 bytes downloaded
[15:54:52] + 532480 bytes downloaded
[15:54:52] + 542720 bytes downloaded
[15:54:52] + 552960 bytes downloaded
[15:54:52] + 563200 bytes downloaded
[15:54:52] + 573440 bytes downloaded
[15:54:52] + 583680 bytes downloaded
[15:54:52] + 593920 bytes downloaded
[15:54:52] + 604160 bytes downloaded
[15:54:52] + 614400 bytes downloaded
[15:54:52] + 624640 bytes downloaded
[15:54:52] + 634880 bytes downloaded
[15:54:52] + 642475 bytes downloaded
[15:54:52] Verifying core Core_11.fah...
[15:54:52] Signature is VALID
[15:54:52] 
[15:54:52] Trying to unzip core FahCore_11.exe
[15:54:52] Decompressed FahCore_11.exe (1843200 bytes) successfully
[15:54:57] + Core successfully engaged
[15:55:03] 
[15:55:03] + Processing work unit
[15:55:03] Core required: FahCore_11.exe
[15:55:03] Core found.
[15:55:03] Working on queue slot 01 [January 26 15:55:03 UTC]
[15:55:03] + Working ...
[15:55:03] 
[15:55:03] *------------------------------*
[15:55:03] Folding@Home GPU Core - Beta
[15:55:03] Version 1.19 (Mon Nov 3 09:34:13 PST 2008)
[15:55:03] 
[15:55:03] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[15:55:03] Build host: amoeba
[15:55:03] Board Type: Nvidia
[15:55:03] Core      : 
[15:55:03] Preparing to commence simulation
[15:55:03] - Looking at optimizations...
[15:55:03] - Created dyn
[15:55:03] - Files status OK
[15:55:03] - Expanded 46659 -> 252912 (decompressed 542.0 percent)
[15:55:03] Called DecompressByteArray: compressed_data_size=46659 data_size=252912, decompressed_data_size=252912 diff=0
[15:55:03] - Digital signature verified
[15:55:03] 
[15:55:03] Project: 5767 (Run 9, Clone 192, Gen 32)
[15:55:03] 
[15:55:03] Assembly optimizations on if available.
[15:55:03] Entering M.D.
[15:55:10] Working on Protein
[15:55:11] Client config found, loading data.
[15:55:11] Starting GUI Server
[15:55:43] Completed 1%
[15:55:43] mdrun_gpu returned 
[15:55:43] NANs detected on GPU
[15:55:43] 
[15:55:43] Folding@home Core Shutdown: UNSTABLE_MACHINE
[15:55:47] CoreStatus = 7A (122)
[15:55:47] Sending work to server
[15:55:47] Project: 5767 (Run 9, Clone 192, Gen 32)
[15:55:47] - Read packet limit of 540015616... Set to 524286976.
[15:55:47] - Error: Could not get length of results file work/wuresults_01.dat
[15:55:47] - Error: Could not read unit 01 file. Removing from queue.
[15:55:47] - Preparing to get new work unit...
[15:55:47] + Attempting to get work packet
[15:55:47] - Connecting to assignment server
[15:55:47] - Successful: assigned to (171.67.108.11).
[15:55:47] + News From Folding@Home: GPU folding beta
[15:55:47] Loaded queue successfully.
[15:55:48] - Attempt #1  to get work failed, and no other work to do.
Waiting before retry.
[15:55:59] + Attempting to get work packet
[15:55:59] - Connecting to assignment server
[15:55:59] - Successful: assigned to (171.67.108.11).
[15:55:59] + News From Folding@Home: GPU folding beta
[15:56:00] Loaded queue successfully.
[15:56:01] + Closed connections
[15:56:06] 
[15:56:06] + Processing work unit
[15:56:06] Core required: FahCore_11.exe
[15:56:06] Core found.
[15:56:06] Working on queue slot 02 [January 26 15:56:06 UTC]
[15:56:06] + Working ...
[15:56:06] 
[15:56:06] *------------------------------*
[15:56:06] Folding@Home GPU Core - Beta
[15:56:06] Version 1.19 (Mon Nov 3 09:34:13 PST 2008)
[15:56:06] 
[15:56:06] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[15:56:06] Build host: amoeba
[15:56:06] Board Type: Nvidia
[15:56:06] Core      : 
[15:56:06] Preparing to commence simulation
[15:56:06] - Looking at optimizations...
[15:56:06] - Created dyn
[15:56:06] - Files status OK
[15:56:06] - Expanded 96544 -> 489240 (decompressed 506.7 percent)
[15:56:06] Called DecompressByteArray: compressed_data_size=96544 data_size=489240, decompressed_data_size=489240 diff=0
[15:56:06] - Digital signature verified
[15:56:06] 
[15:56:06] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:56:06] 
[15:56:06] Assembly optimizations on if available.
[15:56:06] Entering M.D.
[15:56:13] Working on Protein
[15:56:16] Client config found, loading data.
[15:56:17] Starting GUI Server
[15:57:12] Completed 1%
[15:57:12] mdrun_gpu returned 
[15:57:12] NANs detected on GPU
[15:57:12] 
[15:57:12] Folding@home Core Shutdown: UNSTABLE_MACHINE
[15:57:14] CoreStatus = 7A (122)
[15:57:14] Sending work to server
[15:57:14] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:57:14] - Read packet limit of 540015616... Set to 524286976.
[15:57:14] - Error: Could not get length of results file work/wuresults_02.dat
[15:57:14] - Error: Could not read unit 02 file. Removing from queue.
[15:57:14] - Preparing to get new work unit...
[15:57:14] + Attempting to get work packet
[15:57:14] - Connecting to assignment server
[15:57:14] - Successful: assigned to (171.67.108.11).
[15:57:14] + News From Folding@Home: GPU folding beta
[15:57:14] Loaded queue successfully.
[15:57:15] + Closed connections
[15:57:20] 
[15:57:20] + Processing work unit
[15:57:20] Core required: FahCore_11.exe
[15:57:20] Core found.
[15:57:20] Working on queue slot 03 [January 26 15:57:20 UTC]
[15:57:20] + Working ...
[15:57:20] 
[15:57:20] *------------------------------*
[15:57:20] Folding@Home GPU Core - Beta
[15:57:20] Version 1.19 (Mon Nov 3 09:34:13 PST 2008)
[15:57:20] 
[15:57:20] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[15:57:20] Build host: amoeba
[15:57:20] Board Type: Nvidia
[15:57:20] Core      : 
[15:57:20] Preparing to commence simulation
[15:57:20] - Looking at optimizations...
[15:57:20] - Created dyn
[15:57:20] - Files status OK
[15:57:20] - Expanded 96544 -> 489240 (decompressed 506.7 percent)
[15:57:20] Called DecompressByteArray: compressed_data_size=96544 data_size=489240, decompressed_data_size=489240 diff=0
[15:57:20] - Digital signature verified
[15:57:20] 
[15:57:20] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:57:20] 
[15:57:20] Assembly optimizations on if available.
[15:57:20] Entering M.D.
[15:57:27] Working on Protein
[15:57:31] Client config found, loading data.
[15:57:31] Starting GUI Server
[15:58:26] Completed 1%
[15:58:26] mdrun_gpu returned 
[15:58:26] NANs detected on GPU
[15:58:26] 
[15:58:26] Folding@home Core Shutdown: UNSTABLE_MACHINE
[15:58:30] CoreStatus = 7A (122)
[15:58:30] Sending work to server
[15:58:30] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:58:30] - Read packet limit of 540015616... Set to 524286976.
[15:58:30] - Error: Could not get length of results file work/wuresults_03.dat
[15:58:30] - Error: Could not read unit 03 file. Removing from queue.
[15:58:30] - Preparing to get new work unit...
[15:58:30] + Attempting to get work packet
[15:58:30] - Connecting to assignment server
[15:58:31] - Successful: assigned to (171.67.108.11).
[15:58:31] + News From Folding@Home: GPU folding beta
[15:58:31] Loaded queue successfully.
[15:58:32] + Closed connections
[15:58:37] 
[15:58:37] + Processing work unit
[15:58:37] Core required: FahCore_11.exe
[15:58:37] Core found.
[15:58:37] Working on queue slot 04 [January 26 15:58:37 UTC]
[15:58:37] + Working ...
[15:58:37] 
[15:58:37] *------------------------------*
[15:58:37] Folding@Home GPU Core - Beta
[15:58:37] Version 1.19 (Mon Nov 3 09:34:13 PST 2008)
[15:58:37] 
[15:58:37] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[15:58:37] Build host: amoeba
[15:58:37] Board Type: Nvidia
[15:58:37] Core      : 
[15:58:37] Preparing to commence simulation
[15:58:37] - Looking at optimizations...
[15:58:37] - Created dyn
[15:58:37] - Files status OK
[15:58:37] - Expanded 96544 -> 489240 (decompressed 506.7 percent)
[15:58:37] Called DecompressByteArray: compressed_data_size=96544 data_size=489240, decompressed_data_size=489240 diff=0
[15:58:37] - Digital signature verified
[15:58:37] 
[15:58:37] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:58:37] 
[15:58:37] Assembly optimizations on if available.
[15:58:37] Entering M.D.
[15:58:44] Working on Protein
[15:58:48] Client config found, loading data.
[15:58:48] Starting GUI Server
[15:59:46] Completed 1%
[15:59:46] mdrun_gpu returned 
[15:59:46] NANs detected on GPU
[15:59:46] 
[15:59:46] Folding@home Core Shutdown: UNSTABLE_MACHINE
[15:59:49] CoreStatus = 7A (122)
[15:59:49] Sending work to server
[15:59:49] Project: 5755 (Run 13, Clone 209, Gen 21)
[15:59:49] - Read packet limit of 540015616... Set to 524286976.
[15:59:49] - Error: Could not get length of results file work/wuresults_04.dat
[15:59:49] - Error: Could not read unit 04 file. Removing from queue.
[15:59:49] - Preparing to get new work unit...
[15:59:49] + Attempting to get work packet
[15:59:49] - Connecting to assignment server
[15:59:50] - Successful: assigned to (171.67.108.11).
[15:59:50] + News From Folding@Home: GPU folding beta
[15:59:50] Loaded queue successfully.
[15:59:51] + Closed connections
[15:59:56] 
[15:59:56] + Processing work unit
[16:00:00] Core required: FahCore_11.exe
[16:00:00] Core found.
[16:00:00] Working on queue slot 05 [January 26 16:00:00 UTC]
[16:00:00] + Working ...
[16:00:00] 
[16:00:00] *------------------------------*
[16:00:00] Folding@Home GPU Core - Beta
[16:00:00] Version 1.19 (Mon Nov 3 09:34:13 PST 2008)
[16:00:00] 
[16:00:00] Compiler  : Microsoft (R) 32-bit C/C++ Optimizing Compiler Version 14.00.50727.762 for 80x86 
[16:00:00] Build host: amoeba
[16:00:00] Board Type: Nvidia
[16:00:00] Core      : 
[16:00:00] Preparing to commence simulation
[16:00:00] - Looking at optimizations...
[16:00:00] - Created dyn
[16:00:00] - Files status OK
[16:00:00] - Expanded 96544 -> 489240 (decompressed 506.7 percent)
[16:00:00] Called DecompressByteArray: compressed_data_size=96544 data_size=489240, decompressed_data_size=489240 diff=0
[16:00:00] - Digital signature verified
[16:00:00] 
[16:00:00] Project: 5755 (Run 13, Clone 209, Gen 21)
[16:00:00] 
[16:00:00] Assembly optimizations on if available.
[16:00:00] Entering M.D.
[16:00:13] Working on Protein
[16:00:17] Client config found, loading data.
[16:00:17] Starting GUI Server
[16:01:12] Completed 1%
[16:01:12] mdrun_gpu returned 
[16:01:12] NANs detected on GPU
[16:01:12] 
[16:01:12] Folding@home Core Shutdown: UNSTABLE_MACHINE
[16:01:16] CoreStatus = 7A (122)
[16:01:16] Sending work to server
[16:01:16] Project: 5755 (Run 13, Clone 209, Gen 21)
[16:01:16] - Read packet limit of 540015616... Set to 524286976.
[16:01:16] - Error: Could not get length of results file work/wuresults_05.dat
[16:01:16] - Error: Could not read unit 05 file. Removing from queue.
[16:01:16] EUE limit exceeded. Pausing 24 hours.
 
Does the second instance work properly if the first one is not running? If thats the case you might want to make sure you arent running the same machine id for both
 
Does the second instance work properly if the first one is not running? If thats the case you might want to make sure you arent running the same machine id for both

I'd try that.

Also what happens if you swop the -gpu 0 and gpu 1 flags around ??
Will the working copy run on the second GPU core or does it fail ??
Does the non-working copy now run on the first GPU core.

What's the fan set at.
On my GX2's I've got to set the fan at 100% otherwise the cards heat up and fail faster than the onboard logic can spin the fan up to keep the cores cool enough to fold.

Ps. I'd only try swopping the gpu flag just after you've started a new work-unit.
There's no point in loosing 90% of a crunched one if it fails.

Luck ........... :D
 
Well the Second instance does not run even with the first instance not running.
When I switched the flags for the second instance to run on the first core instead of the second it works so it seems that for some reason of another the second core does not want to work.
 
I would try increasing the fan speed at least as a test.
 
Its at 100%. my temps for both cores and PCB's are ~43 Celsius

I have tried everthing....what could keep it from running on the 2nd GPU?
 
even nuking the driver and reinstalling ?
 
Yea I nuked my whole Windows installation before I posted this topic.
 
Are you only useing one 12v rail to feed the card ??
It could be that the 12v rail is dropping a little low at full load.
Try altering how you power the card and see if that makes any difference.

What happens if you underclock the card ??

Can you alter which PCI-E slot it's in ??

Just a couple of ideas

Luck ............ :D
 
Man I have never been so damn confused.

I have tried all the above suggestions. except moving PCIe slots...I only have one.

You think I would have better luck with a different OS?
 
Have you tried gaming with SLI enabled to see if there's any lockups or slow downs?
Maybe try 3DMark Vantage with AA/AF enabled to stress the cores/memory to see if you have a problem with the second core or memory feeding the second...
 
Hey Nitro.

I responded to your other thread titled "GTX 295"

You may try adding the -forcegpu nvidia_g80 flag in your command line for each client with the -gpu 0 for client 1 and -gpu 1 for client 2, also ensuring the machine ID is unique (as mentioned).

If that doesn't work, you may have a bad second GPU. I have two GTX 295 cards folding in one box successfully with a 1000w PSU. Try gaming or running a benchmark with sli enabled and see if you have any issues. Your temps are lower than mine, so that shouldn't be an issue.

You might want to check GPU-z to ensure all speeds are listed properly. IF they are, this takes out the possibility of a driver problem.

Usually "NANs detected" is a bad GPU or a mis-detected GPU, or an underpowered GPU. I suggest having a 750W PSU for one GTX 295 card.
 
ok that sounds like some things I need to try out.

Yea I upgraded to 1000w cause my box is just loaded to the max.

I dont think I have a bad GPU cause I can run badaboom video encoder on the second one just fine and it runs just as fast and the first core. But I guess its not out of the question.
 
Back
Top