• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Data Corruption with Tyan board?

eno-on

[H]ard|Gawd
Joined
Feb 20, 2005
Messages
1,114
Hey folks.

I have an issue. A client of mine is getting Exchange Database corruption about every 4 months.
Needless to say, they are unhappy with the situation, as it costs them large sums of money every time we have to rebuild it.

The 4th time it's happened now, and I'm starting to suspect some mobo weirdness.

The motherboard is a Tyan B2891G24S4-LC, running 2 opteron 246's, Server 2003 (32 bit), 4gb of corsair EEC/Reg/Buf ram, and a couple sata 150 320gb drives in raid 1.

It has passed every cpu/ram/hd test I've thrown at it, with flying colors. I've moved 2 gb files back and forth 4 or 5 times with no data corruption. Windows itself is not corrupted. It has been rock stable.

Has anyone heard of data corruption from, possibly, the NIC card(s) on this mobo? One of the nic cards went bad... in like a week after we built it.
I remember some weird data corruption issues with nvidia storage hardware a while back on coonsumer level boards. Anyone heard of this on their server chipsets?
 
OK folks, I have replaced the mobo with the same model, changed the ram, used a different raid controller, and different harddrives. It still starts giving me ntfs errors and data corruption.
Any ideas?
 
Personally, I'd start looking at the data itself for the issue, along with what the users are doing to the system. For example, does the data always occur in a similar section of the data, is it a couple of users data that seems to be damaged?
 
The other thing I think it could be is dirty power.
Is the box on an UPS that will condition the power in ?
How close to spec are the voltages and is there any ripple in the lines ?
The 5v line may be suspect as that powers the PCI slots and the hard drives software.
If you cannot test then try swopping the PSU and see if that makes a difference.

Luck ......... :D
 
I'll take a look at the PSU, but it's a server class PSU.

Let me elaborate on this.

Machine was corrupting the OS, and slowly destroying the exchange database.

Turned off the onboard nVidia sata raid controller.

Installed a PCI adaptec sata raid controller. Installed 2 new WD 250gb raid-tested harddrives.

Installed server 2003. Started receiving ntfs errors.

Changed out the motherboard. Installed server 2003...... started receiving ntfs errors...

Changed the Ram (from 4gb corsair ecc/reg to some wintec ecc/reg), installed server 2003, started receiving ntfs errors...

Ran prime 95 on each cpu for an hour..... no errors.....

Changed CDRom cable... installed server 2003.... started receiving ntfs errors.....

quickly built a core2duo machine on a cheap gigabyte 965 board, threw a couple gigs of ram at it, installed server2003 on it to use as a standby while I figure out these damned errors...

I mean, seriously, whats left here? Besides, I guess, cpus and PSU? I'm also going ot try and install from a different CDROM.

But yeah, working on it for 13 hours yesterday, not so much fun.... NOTHING shows up as instable in ANY hardware test I throw at it. No errors... I'm going to install XP on it... see what happens.

I'm not ruling out a chipset issue. I'll try a bios update was well. I think this is really an I/O issue with some part of the chipset. Again,I remember a bad data corruption problem with the NF4 chipset on consumer level boards. I wonder if I have the same issue.
 
Well professional and consumer level NVRAID solutions are identical. They won't vary from board to board too much either as the NVRAID BIOS, drivers and chipsets are all the same. For anything enterprise level, it's nothing but a solid Intel or LSI MegaRAID controller for me.

This probably stems from Tyan's issues in general. Many of their boards are faulty in a number of ways as I've learned with my constant discussions with AreEss. (Validated on both of the Tyan boards I've had.)

Basically, the Tyan boards suck. That's about the end of it. Even if you use nothing but recommended hardware from their lists, you will still have problems in certain situations. Basically Tyan has slipped big time in recent years. My $500 Tyan K8WE was one of the better boards out there, but I got lucky. Anyone else who had one had nothing but issues with it. My advice to you is to switch to Supermicro and don't look back. Not to say I didn't have issues. BTW all my issues were with the LAN controllers and storage controllers, and support of storage controllers used in PCI-X slots.

You also want to look into your power supply very carefully as well.
 
Well, I disabled the NV SATA/RAID shit, and installed a PCI adaptec SATA/RADI controller.
The LAN on the board is NOT NV, its broadcom(?) GB LAN.

Still happens...
 
Well, I disabled the NV SATA/RAID shit, and installed a PCI adaptec SATA/RADI controller.
The LAN on the board is NOT NV, its broadcom(?) GB LAN.

Still happens...

I wasn't saying the LAN port was NVIDIA. I was saying that Tyan's are notorious for issues with the LAN ports. They also have a number of issues with their expansion slots too.
 
You need to check on reports of ECC errors. If you OS is not set up to log them, you can access them in the BIOS.

Superpi is a good test here, mprime is useless for these things.

Most likely a software problem, or maybe the harddrive.
 
At this point, I'd guess software problem. Have you looked through the MSKB for possible ntfs/exchange corruption in certain situations?
 
*sighs*
This is why I refused to sell the 2891.

The symptoms you're describing are consistent with transient bus errors.

But here's the OTHER part of the question that nobody here bothered to ask, and SHOULD have; is this part of an Exchange Cluster? If yes; watch. I'd bet money you're hitting split-brain on a regular basis.

Otherwise; quit wasting your time. The 2891 is a joke; the Transport doubly so. The backplanes in the cases they sourced (Chenbro) are known to cause innumerable failures, because they are such GARBAGE. Among them? Data corruption, data corruption, and data corruption, along with drives going offline for no reason then reappearing, and even floating SCSI IDs on their SCA backplanes.
 
Back
Top