• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Computer crash when processing large amounts of data

KrazeyKami

n00b
Joined
Dec 31, 2010
Messages
61
Hi all,

For a short introduction of my tech level; I’ve been a system manager for the past 8 years, and been playing around with computer for nearly 20 years.
Until recently, I’ve never faced a problem with my own computer that I couldn’t solve, or figure out what was giving issues, until now… And that’s why I need your help, and a fresh look on things.

First, my setup:

Motherboard: ASUS P6T SE
Motherboard BIOS: v.0808
CPU: Intel i920 @ 2,6 GHz (Stock)
CPU Heatsink: ProlimaTech MegaHalems + 2 Cooler Master120mm fans
CPU Idle Temp: 30-35 degrees per core
CPU Load Temp: 45-50 degrees per core (Prime95)
Memory Part Number: 6 GB OCZ Gold PC3-8500U + 6 GB OCZ Gold PC3-10700U (Both OCZ3G1333LV6GK @ 1066 MHz, Stock)
Memory Voltage: 1.65
Video Card(s): Asus nVidia GTX295
Sound Card: Sound Blaster Fatal1ty X-Fi
PSU Model Number: Cooler Master 1000W Real PowerPro
Hard Drive(s): Intel X-25 M SSD 80G (OS/Boot), 4x 2TB HD204UI Samsung Spinpoint F4 (all on ICH10 / Sata2)
Optical Drive(s): GGW-H20L BlueRay RW
Other Cooling: Cooler Master Stacker 831 case with 6x Cooler Master 120mm fans
Operating System: Windows 7 Professional x64

This machine is running 24x7, rock solid. I never had any issues with performance or stability.

Now, for the problem:

Until a few weeks ago, I didn’t have the 4x 2TB drives in it.

I bought the drives, made a RAID5 array, using the ICH10R, and that’s when the problems began:
When copying (or downloading) large amounts of data to the array, thus creating a high I/O on any of the 2TB drives, my computer suddenly reboots. No blue screens (although that option is checked ON), and the only remarks in Event Viewer are: System suddenly rebooted unexpectedly, Event ID 6008, and possible cause: Power failure, Event ID 41).
No error codes, no nothing.

I decided to break up the array, and see what happens when I copy 1.5 TB of data from 1 drive to the next; The computer reboots again. The drives are now single SATA2 drives, with a 2 TB partition. After 10-80 minutes, the system reboots without notice or error.

So, I started to systematically remove drives, and test with the other drives.
Regardless of what drive is the source, or destination (I tried all combo’s, and all directions), the system reboots when a high amount of data is generated. Not just with copying the existing data, but also when downloading (and at the same time repairing files (.PAR), and extracting).

I tried the following:

  • Check for overheating: All values are well below 40 degrees C; (also checked drives);
  • Even put an active cooler on my southbridge (The ICH10);
  • Memory checks. Ran 3 different programs to check / test my memory, ran overnight for hours and hours, multiple passes, 0 errors.
  • Remove all other hardware, except for the absolute minimal necessary;
  • Swap / replace powercords, SATA cables, even rotate drive position on the SATA connectors;
  • Reset BIOS settings;
  • Reinstall Windows 7 (delete the entire 80GB partition on the SSD and reinstall, no other tweaks, but right after install, start the copy transactions) to rule out the possibility of faulty software and / or drivers;
  • Reformat / create the drives / partitions: Tried both MBR and GPT partition; Different block sizes;
  • Turn off Write Back Cache, to even further rule out a problem with my RAM;
  • Calculate the PSU needs; I tried multiple programs, even a paid one, and counted manually: Granted, on a full Direct3D load (games on high etc), my GFX card needs around 450 watts. This makes a grant total of 950 Watts. However, the problems occur while idle in Windows, so the consumption for my GFX is max. 100 Watts, making a total of (roughly) 600 Watts, well within the limits of my 1000W PSU;
  • CPU check / Prime95; runs for days, stable, without a single error;

The facts:
  • the problem only (and only) occurs when copying / downloading a large amount of data;
  • The system runs flawlessly under high load (playing games, watching movies, running programs etc);
  • I never had this problem before, but then again, I never had the space to start downloading 250 GB of data, or copying 1.5 TB ( I didn’t even have 1,5 TB ^^) data to other drives.
I am able to reproduce a “fast” reboot error:
I created 3 separate batch files, which basicly tell Robocopy to copy data from:
  • Drive D: to Drive E:
  • Drive D: to Drive F:
  • Drive D: to Drive G:
  • When I run these scripts separate it runs for a while, but also reboots / crashes after an hour or so.
  • When I start these scripts all at once, it reboots within 5 minutes.
  • Remember, that I already cloned the drives, so I could rotate the source drive, and systematically removed / switched a destination drive, thus trying all different combo’s (and to check whether one of the drives might be faulty).
  • Further this also makes me doubt if it’s the shear size of data that causes the problem, cause within 5 minutes, not even 100 MB is being addressed, and still, it reboots.
  • This is making me think, that the PSU might be the problem. As soon as these drives are actively called upon, i can imagine a sudden increase in the 12V+ rail, can cause to overload my 12V rail... altho my PSU has 6x 12V rails, i'm not to convinced this might work as well as people say...there are many discussions on the web about the use of 6x 12V rails. Could it be, that my 12V rail is maxed? (considering it's giving power to: The Mobo, The Cpu (4/6 pins, can't remember), The GFX card (both 6 and 8 pins), 5 drives (SSD + 4x 2 TB), an Optical BD-RW, and ofcourse the onboard devices (Soundcard) and USB devices (headset, webcam).

I am all out of ideas. If there is something that I haven’t checked / tried, please tell me. I think I wrote down everything I tried thus far; maybe I missed something, but I’ve been testing and trying for 3 weeks now.


For now my conclusion / suspects are:
  • The motherboard. Either a chip in the ICH10 was fried, or the SMBus got a dent;
  • The motherboard (or ICH10 / SMBus) is just not capable of processing such large amounts of data.
  • My PSU. Mainly, the 12V+ rail. It could be (maybe), that my 12V rail is max.loaded, and when kicking in the extra drive operations, it fluctuates, and tilts it a bit above it maximum, thus crashing my computer. Looking at the symptoms (sudden reboots without any errors) it might be a more plausible cause then all the other things I tried. And yet, if you look at my hardware setup, I cannot imagine I reached the 12V max.
    However, I will be testing this week, by taking another PSU on a second desktop, and connect my hard drives to that power supply. Or maybe even, take an el cheepo GFX card and remove my GTX295, and see if the computer stays stable during copying…

I’ll post the results shortly. In the mean time, if any1 has seen this problem before, or has other ideas / solutions to try, please let me know here and I’ll try them.

P.S.:
According to the PSU calculator, this is the recommended Watts / Amperage for my setup with a full 100% load; Below that is a table of power my PSU can handle:

psu.JPG


Afaik, i can add the 12V rails together, so my PSU can handle max. 128 A on the 12V rail?
This is my PSU: http://www.coolermaster.com/product.php?product_id=2519

Does this mean my PSU should have more then enough power?
I'm still gonna try with a different PSU or GFX card to be sure, but according to the above i think all should be covered... Any thoughts?

Many thanks in advance for thinking with me.

Kind regards,
Kami.
 
Last edited:
hello, i have similar problems with my PC.
http://forums.storagereview.com/index.php/topic/27704-ich10r-raid5-silent-data-corruption/page__p__256873__fromsearch__1#entry256873
i haven't resolve my problem, i just use the other SATA controller on my mobo (jmicron 36x).

I just read your post, but your problem is totally different from mine;
  • The data copied in my case is not corrupted; I'm suffering from reboots. The data that does get copied, is fine.
  • It applies to all kind of load, not only large files.
  • I'm no longer using RAID5

Sorry, but i can't help you with your problem...
 
Hi all,

The problem is solved.

The X58 mobo can't handle 12 GB of RAM.
I removed 6 GB, and the machine is rock-solid again. It seems that in Auto settings on the BIOS, the QPI voltage is too low for 12 GB of RAM and needs to be entered manually.

I'm still waiting for a reply from OCZ on what the best settings would be to have both types of RAM run on the right voltage and SPD's.

Kind regards,
Kami.
 
Good to see you've found a source for your issue.
Afaik, i can add the 12V rails together, so my PSU can handle max. 128 A on the 12V rail?
This is my PSU: http://www.coolermaster.com/product.php?product_id=2519

Does this mean my PSU should have more then enough power?

You cannot add the 12V rails. It does not make sense: 128A would mean that your 1000W PSU would have 1536W of power on the +12V rail. You determine the amperage on the +12V rails by first finding out what's the total combined, max load, combined or max wattage set aside for the +12V rails/section alone. Then divide that total by 12 and you get how much amps the PSU has on the +12V rail.

In your case, the PSU label clearly states that it has 936W available on the +12V rail. So dividing that by 12 would get us the total amperage on the +12V rail for your PSU: 78A.
 
Good to see you've found a source for your issue.


You cannot add the 12V rails. It does not make sense: 128A would mean that your 1000W PSU would have 1536W of power on the +12V rail. You determine the amperage on the +12V rails by first finding out what's the total combined, max load, combined or max wattage set aside for the +12V rails/section alone. Then divide that total by 12 and you get how much amps the PSU has on the +12V rail.

In your case, the PSU label clearly states that it has 936W available on the +12V rail. So dividing that by 12 would get us the total amperage on the +12V rail for your PSU: 78A.

Hey Danny,

Thank you for explaining that to me. I wasn't to sure about how to read these values.
Im glad the system works again, so that means i finally can rebuild my RAID5 :D

Btw,
Did you get my PM / Post on the CPU behaviour when Reading / Writing @ the ICH10R RAID5 Array? http://hardforum.com/showthread.php?t=1573816
If you want me to do some other / additional testing, just let me know.

Kind regards,
Kami.
 
Hi all,

The problem is solved.

The X58 mobo can't handle 12 GB of RAM.
I removed 6 GB, and the machine is rock-solid again. It seems that in Auto settings on the BIOS, the QPI voltage is too low for 12 GB of RAM and needs to be entered manually.

I'm still waiting for a reply from OCZ on what the best settings would be to have both types of RAM run on the right voltage and SPD's.

Kind regards,
Kami.

Sweet, I was wondering what the issue could be with that setup.

btw, OCZ Gold was designed to have higher voltages running with it, so if you increase the voltages, your memory should be ok, just make sure your mobo can handle it. ;)
 
Sweet, I was wondering what the issue could be with that setup.

btw, OCZ Gold was designed to have higher voltages running with it, so if you increase the voltages, your memory should be ok, just make sure your mobo can handle it. ;)

Thanks :) right now im waiting for both Asus and OCZ, to see if they know if this issue and second, what the right settings would be to align the different types of memory.

My mobo has a big red flashy warning sticker on the DIMM sockets on purchase, saying not to use memory that has a higher voltage then 1.65. This is probably just a guarantee value, altho i'd rather not OC my system, nor make to much standard changes.

I'll check it out and wait for OCZ to come up with the right values.

Thanks for your input!

Kind regards,
Kami.
 
Hi all,

The problem is solved.

The X58 mobo can't handle 12 GB of RAM.
I removed 6 GB, and the machine is rock-solid again. It seems that in Auto settings on the BIOS, the QPI voltage is too low for 12 GB of RAM and needs to be entered manually.

I'm still waiting for a reply from OCZ on what the best settings would be to have both types of RAM run on the right voltage and SPD's.

Kind regards,
Kami.

Yikes! I have 24GB in two X58 systems right now (6x 4GB) and another system still with 12GB (6x2GB). Very odd that your board cannot handle 12GB.
 
so guys, do you thing it is good idea to have so much ram (not ECC) and also not reliable hardware (intel ICH) without any kind of checksuming against silent data corruptions?
 
Well, silent data corruption was never the issue here. I've managed to create a solid stable RAID5 construction with a 240 mb/s Read and a 260 mb/s Write (http://hardforum.com/showthread.php?t=1573816). The ICH10(R) is very good and have never given any problems, even with this Fakeraid setup.

The problem here is the combo of the OCZ memory and the X58, and more precise, it seems the be the *mixed* / 10700U modules i.c.w. the Auto Memory settings in the BIOS.

OCZ has yet to respond, but if you take a look on their forums (and also Asus'), there are plenty people that experience problems with this setup due to the Auto SPD / Voltage settings.

As soon i have a response from OCZ and some stable, reliable settings, i'll post 'em here.

Oh ps, ECC memory on my desktop for games and movies? :p

If i wanted a professional setup, i'd get me 3x 15 fiber scsi disk enclosure SAN on a high speed fiber network, combined with ESX clusters for high availability access ;)
Only the SAN alone is around 150.000 euros per year in service contract :)
 
Very odd that your board cannot handle 12GB.
It can.

The OP is afraid to raise the voltage to make the RAM stable.

Buy RAM that's rated for 1.5V or follow this....
The reason for the 1.65v "limit" is b/c the default/standard Vtt is 1.15v, and 1.15v + 0.5v = 1.65v. So, if you wanted to run the RAM voltage at 1.7v for some reason, you would need to increase the Vtt to at least 1.2v to stay within the 0.5v difference.

There are many posts on this subject.
 
yeah...12GB ram for desktop, ot 24GB...or 10TB hdd for desktop...do you have ever think what will be if you lose your 10TB data ? :D
 
it is your data...not mine :cool:

I have no idea what you're talking about.

Many people this day and age have 50TB+ data arrays for home or server setups.

2 and 3TB drives have come way down in price and are very affordable now.

btw, I think everyone in this thread understands that RAID does not equal a backup. ;)
 
Anyway, what I wanted to share is that in the presence of such large amounts of data it is reasonable to consider using some method for detecting and maybe recovering silent corruptions. because I had exactly this problem and I found it a little bit late, the thing is that I can no longer count on today's cheap hardware.
 
Anyway, what I wanted to share is that in the presence of such large amounts of data it is reasonable to consider using some method for detecting and maybe recovering silent corruptions. because I had exactly this problem and I found it a little bit late, the thing is that I can no longer count on today's cheap hardware.
I have an idea!

If you have a problem why not start a new thread instead of making a nonsensical jumble of words and negative comments alluding to a "phantom problem" you may have had?

Make any sense to you or would you like a mind-meld? :D
 
Anyway, what I wanted to share is that in the presence of such large amounts of data it is reasonable to consider using some method for detecting and maybe recovering silent corruptions. because I had exactly this problem and I found it a little bit late, the thing is that I can no longer count on today's cheap hardware.

I have no clue what this has to do with this thread.

The OP did not have a HDD problem, it was a voltage problem with his mobo and the amount/type of memory being used.

Yeah, there is a method for detecting and fixing silent corruptions, it's called ZFS, search these forums and you'll find a ton of threads for it. ;)
 
I have no clue what this has to do with this thread.
LOL!

I'm thinking trackersoft can only be contacted thru tin-foil hats.

Mine must be outta commision because I can't contact him. :D
 
I used that same board, same ram, with 12GB. You have to crank the voltage (1.7V iirc, raise it slowly.. read forums) and lower the bus (you're already running at 1066, I was at 1600) to get it stable, and then even so I would get memory errors with the HCI Memtest boot cd. Out of about 200GB, only half would meet their specs, ended up switching to corsair with almost 100% pass rate.

It should be a hint that your ram test wasn't thorough enough if it ended up being ram and you say you tested it.. Which did you use?
 
It should be a hint that your ram test wasn't thorough enough if it ended up being ram and you say you tested it.. Which did you use?

He used OCZ Gold, which should support higher voltages.
 
He used OCZ Gold, which should support higher voltages.

Yeah, but how does that matter if he's running the voltage too low to be stable? Or if he's not testing well enough to determine if it's stable? I guess I don't know what you're getting at..
 
Yeah, but how does that matter if he's running the voltage too low to be stable? Or if he's not testing well enough to determine if it's stable? I guess I don't know what you're getting at..

What I mean is, is that if the OP increases his voltages, his memory will easily handle it without any harm.
 
What I mean is, is that if the OP increases his voltages, his memory will easily handle it without any harm.

Okay, then you missed the two points which my comment actually addresses, so I'm not sure why you replied to me.

1. If his current stability test did not discover the instability at a lower voltage, why would they perform any better at a higher voltage? Even if the ram could accept 5V, he wouldn't be able to say whether it is stable or not.

2. This particular RAM and motherboard using 12GB is tricky to get truly stable, even at higher voltages. I built about 30 of them for scientific computing.

Going back through, it seems you misread my comment. I said which ram test did you use, and you thought I said which ram.
 
Ah, yes, I though we were talking about memory, not tests, haha.
 
Hey all :)

I'm reading all your replies with lots of interest!

Well, i have an update on the issue:

I've tried the 6GB sets seperate from each other. Seems with 6GB, i can stress my PC all i want (high IO, Downloads, etc) and it runs stable without crashing.
So, this rules out the possibility of faulty RAM.

It could very well be, that (one of) the other 3 RAM Sockets is broken / damaged. I'm just not sure if i can start filling the A2, B2, C3 slots. I read that (and afaik it always was like this) you should start with A1, B1, C1 in case of triple memory. Anyway, lets assume the slots are fine. I'm tending to test this tonight tho.

I made a post on the OCZ forum a few days ago, no reply yet: http://www.ocztechnologyforum.com/f...3-10700U-(Both-OCZ3G1333LV6GK)-on-ASUS-P6T-SE

I tested the RAM sets seperately and checked the settings / voltages etc with CPU-Z. This is the result:

3x 2GB PC3-8500F:
PC3-8500F.jpg



3x 2GB PC3-10700F:
PC3-10700F.jpg



I notice that, indeed, the voltage is Auto set to 1.5 volts.
Furthermore, i notice that they run on different timings; 7-7-7 vs 8-8-8.

If i look in the Timings Table, i see that, when comparing e.g. the 8-8-8 settings, they both have different Frequency speeds, different tRAS, different tRC and different tRFC's.

I'm not sure if this would be in indicator to the problem of combining these DIMMS;
Does this mean they'd run on different out-of-sync settings, or are they forced to the same settings? Also, on the sticker of my memory, it says 9-9-9@1,65v.

I hope OCZ will respond soon to these questions. Basicly i'd like to know what the right settings should be to enter manually in the BIOS, and if this won't give any problems with the 2 sets combined.

Oh, i also got a call from an Asus support agent last night; He confirmed that OCZ memory is known to get the wrong / too low Voltages from the Auto Settings. He recommends to manually set the Voltage to 1.65v. Still, i'd like for OCZ to confirm this and also give me the rest of the settings.

I'll keep you all posted, and if some1 has some facts on this, i'm all ears / eyes! :)

Kind regards,
Kami.

p.s.:
Just for the record: seems this problem overall has nothing to do with my HD's, Silent Data Corruption, ICH10(R) or any RAID5 constructions. It was just where i started my search to ID the problem.
 
Well, i tried to change the settings manually, as suggested by OCZ:

http://www.ocztechnologyforum.com/fo...on-ASUS-P6T-SE

Nothing worked. Im seriously doubting that the memory types are compatible.
The funny thing is, OCZ is still claiming this is exactly the same type; it's even sold at the local store under the same Product number. However, both SiSoft and CPU-Z claim these memory types are different, and both have their own specific speeds and SPD's. Funny fact tho, on the sticker on the modules, it both says PC-10666 - 9-9-9@1,65v.
If you look at my previous post you see the screenshots as how the memory is being detected by my computer.

Im still waiting for a reply from OCZ and i wonder as to how they react. Personally, i'd like to try with 6 DIMMS of the PC3-10700U memory. I have a gut feeling that this is the key to all the problems.
 
12GB of RAM is pretty much overkill for most desktop users.

6GB is just fine unless you are doing something really stressful like maybe video editing.

If you really wanted to use all 12GB of your RAM I'd use the timings of the lowest RAM (7, 7, 7), start the voltage at 1.55 and keep increasing it until you get stable.

With the different specs on these RAM sets I'm betting you'll have to experiment with the timings and voltage to get it stable.

You're really making this "much ado about nothing". ;)

The simple solution is to buy a matched 12GB set.
 
12GB of RAM is pretty much overkill for most desktop users.
~
The simple solution is to buy a matched 12GB set.

Yepz, and that's the problem. I bought the same type (in the store, same productnumber etc), but it showed up as different memory. I made a post on OCZ's forum about it a year back, and they claimed it was the same type. I'm planning to prove otherwise.

Tonight i'll test what will happen when i fill each DIMM_Channel with the same type of memory: A1 + A2 = 10700, B1 + B2 = 8500. Then i'll fill C1 with 10700 and if everything is still running smooth and stable, i'll fill C2 with 8500. If the system crashes again at that point, i think it's proven that the memory is of different types / speeds and when mixed in a channel gives an error.

Because OCZ and my local reseller all claim its the same type (altho a newer revision according to OCZ, and a completely different type of memory according to CPU-Z, SiSoft Sandra and Everest), i hope / think it's fair they let me trade in the 8500 memory for 3 DIMMS of 10700 memory. It's the same in their book anyway.

---
Btw, the reason i need 12 GB, is that i'm running several servers with VMWare for test lab / education purposes. With 6GB, i can run a max of 4 servers, but it's getting slughish. With 12Gb, it runs alot better and i can add some more servers with some memory ballooning. I could lower the virtual RAM per server, but for some of the servers (SCCM, SQL) 2GB really is a must have.


Anyway, it's not my intention to make a big fuss out of this ;) I started wondering myself since a few posts back if this thread should belong in the Memory section, rather then Data Storage systems. However, maybe if some1 is experiencing similair problems when shifting large amounts of data, this could help them to find the problem.

Kind regards,
Kami.
 
Btw, the reason i need 12 GB, is that i'm running several servers with VMWare for test lab / education purposes
Ah Ha, I see your reasoning and I agree.

Even if you bought a 12GB RAM kit you'd probably have to up to voltage. That's just the way it goes.

It doesn't seem like you're bold enough to set the voltages yourself so I guess the waiting game for a specific answer, which may never come or even work if it does, is your only answer.

Good Luck and let us know what the official response is. :)
 
Ah Ha, I see your reasoning and I agree.

Even if you bought a 12GB RAM kit you'd probably have to up to voltage. That's just the way it goes.

It doesn't seem like you're bold enough to set the voltages yourself so I guess the waiting game for a specific answer, which may never come or even work if it does, is your only answer.

Good Luck and let us know what the official response is. :)

Hey m8,

Yeah, i got a response and changed / upped the voltages, SPD's, and whatnot. Seems it doesn't work. I feel strongly the problem is with the 2 different types of memory.

And yea, it's on now:
http://www.ocztechnologyforum.com/f...3-10700U-(Both-OCZ3G1333LV6GK)-on-ASUS-P6T-SE

I'm simply stunned with every1 (OCZ and my local reseller) simply denying that there is a hardware conflict / problem.

I'm never gonna buy their crap again, that's for sure.
 
its shouldnt be new news that packing twice the amount of high-frequency signals in the same area (12gb vs 6gb of ram / 3 sticks vs 6 sticks) increases noise and interference in the electrical signals. its not an OCZ problem specifically.

the ways to get around it are...
-increase voltage, which acts as an amplifier to the signal (has obvious downsides)
-properly shield the traces or have better trace routing (the job of the motherboard designer)
-buy RAM rated at a higher speed than what you intend to run (aka, buy 1600 RAM and run at 1333 speed)
-loosen RAM timings until you get stable

all of which have positives and negatives you have to balance.
 
Back
Top