Excessionoz
Weaksauce
- Joined
- Jun 28, 2011
- Messages
- 111
UPDATE: SOLVED
Disabled on-board NIC
Installed Intel NIC
System behaves perfectly, no checksum errors,no crashes on large network transfers.
CASE CLOSED.
Thankyou to all for your various suggestions.
FreeNAS, P5E-VM-HDMI and headaches galore.
Longwinded versions already posted over on Ars Technica:
The Plan: Keen Plan
The Test: The Test
The FAIL: The Fail
The Original Hardware:
The 'Upgraded' Hardware
First step was testing FreeNAS 8.0 (Release) on the original hardware. It seemed to work but NOTE WELL, after transferring about 200GB of files to my test RAIDZ1 pool, I never did a "zpool scrub <pool>" command, so I don't know if there were any of these Checksum errors that my later testing came up with on different drives.
Second test, new configuration, put together
After about 30 minutes the SSD threw a fit during one of the operations (it was already a highly suspect SSD, but that is a different tale), more importantly, 10 minutes after that, Kernel Panic (trap 9) and the machine stopped dead.
Next day,
Restarted FreeNAS 8.0 (Release) from scratch, subsequently had many crashes, kernel panic trap 9's kernel panic trap 12's, with uptimes averaging 2.5->3 hours.
detailed post on ars
Updated to a nightly build of FreeNAS 8.0.1 (BETA3). Similar results, crashes and data corruption on the data pool (checksum errors on scrub, unrecoverable data errors revealed in pool).
Still getting data corruption running 5 drives, and a kernel panic.
On Tuesday I had tried NexentaStor, OS 'silently' crashed after several hours of use (could enter credentials, but console became unresponsive), with both Windows boxes rebooting overnight due to MICROSOFT PATCH TUESDAY. Since the Windows boxes rebooted, I had no idea how much they had transferred to the Nexenta box before they had the rug pulled out from under them.
Subsequently I couldn't reboot the USB stick into Nexenta due to some 'device out of space' error during boot (I used a Sandisk 4GB usb key). So I have no idea if the Zpool drives were getting errors.
Kernel panic screens seem to have vanished (but read on), Zpools are still suffering checksum errors when doing zpool scrub commands (NOTE WELL: All Scrubs were performed whilst the NAS was having data transferred to it)
Last night:
Slowly (using Crashplan) copied ~20GB of data to NAS,
ZPool Scrub reported no errors.
Subsequently hit NAS with as much as I could from two Windows 7 computers with straight file copies of large files, getting transfer rates of 600-700+Mbits/second over gigabit LAN.
Some 20 minutes and 60GB later, did a Zpool scrub, and I got 8 or so checksum errors on all three drives
Continued to hit NAS until transfers complete. Ended up with ~250GB data, and hundreds of Checksum errors on scrub (50-80 per drive) and 11 unrecoverable errors on large (monolithic 211GB Acronis TIB file) backup test file.
:-(
Left both Windows boxes hammering at NAS for an uptime test, it was still functioning 12 hours later.
Again, the Windows boxes both rebooted overnight (thanks for PATCH WEDNESDAY following on from PATCH TUESDAY, Microsoft!). Normally my Windows box is up for weeks at a time
, so I had no idea if they had problems 'writing' to the NAS at any time.
After 12 hours 15 minutes, FreeNAS console reports a Kernel Panic, 12, Page fault in Kernel Mode, no process name apparent.
The box was (essentially) doing nothing at the time of the crash.
So now I'm writing to you folk, hoping that someone will read the story without cringing, without suffering ADD from reading such a lot, and withstanding the urge to shout out already-tried-suggestions:
This hardware was stable for -months- running Kubuntu in 2009 using the original 0301 BIOS firmware. Box has been turned off since mid 2010.
More information (same stuff, as-it-happened reports)
here
and here
and here
Next step -- swap 3 * 2TB Hitachis for 3 * 1TB Western Digital drives in drive cage and retest. Although I can't see how faulty -data only- hard disks could be sending FreeBSD into Kernel Panics.
TLDR version: My FreeBSD/FreeNAS computer keeps getting Kernel Panics and corrupting my ZFS drives constantly, help!
Apologies for vast first post to forum
Hope it's in the 'right' place.
Disabled on-board NIC
Installed Intel NIC
System behaves perfectly, no checksum errors,no crashes on large network transfers.
CASE CLOSED.
Thankyou to all for your various suggestions.
FreeNAS, P5E-VM-HDMI and headaches galore.
Longwinded versions already posted over on Ars Technica:
The Plan: Keen Plan
The Test: The Test
The FAIL: The Fail
The Original Hardware:
Code:
Motherboard: ASUS P5E-VM-HDMI (6 * SATA) GbitE
CPU: Intel A-E8200 Core 2 Duo @ 2.66Ghz
RAM: Corsair XMS2 TWIN2X2048-6400C4 * 2 ( 4 * 1GB DIMM) 4GB
HD: 3 * Western Digital WDEACS 1TB drives
DVDROM: SATA DVD reader.
HDDCase: SUPERMICRO CSE-M35T-1B Black 5x 3.5" Hot-swap SATA
USB Key Crucial 4GB stick (for FreeNAS boot).
The 'Upgraded' Hardware
Code:
RAM: Corsair VS4GBKIT4096-6400C5 * 2 ( 4 * 2GB DIMM) 8GB
HD: 5 * Hitachi 5K3000 2TB Deskstar drives
SSD: 64GB G.Skill SSD
First step was testing FreeNAS 8.0 (Release) on the original hardware. It seemed to work but NOTE WELL, after transferring about 200GB of files to my test RAIDZ1 pool, I never did a "zpool scrub <pool>" command, so I don't know if there were any of these Checksum errors that my later testing came up with on different drives.
Second test, new configuration, put together
Code:
5 * 2TB Hitachi drives
1 * 64GB 'G.Skill' SSD I had lying around, as a CACHE drive in the RaidZ1 array.
After about 30 minutes the SSD threw a fit during one of the operations (it was already a highly suspect SSD, but that is a different tale), more importantly, 10 minutes after that, Kernel Panic (trap 9) and the machine stopped dead.
Next day,
Code:
* Replaced the 4GB of PC6400C4 (4 * 1GB) Corsair RAM with 8GB of PC6400C5 (4 * 2GB) Corsair RAM.
* Changed drives in BIOS from IDE to AHCI
Restarted FreeNAS 8.0 (Release) from scratch, subsequently had many crashes, kernel panic trap 9's kernel panic trap 12's, with uptimes averaging 2.5->3 hours.
detailed post on ars
Updated to a nightly build of FreeNAS 8.0.1 (BETA3). Similar results, crashes and data corruption on the data pool (checksum errors on scrub, unrecoverable data errors revealed in pool).
Code:
* Updated P5E-VM-HDMI BIOS from 2007 (0301) to latest available (0709).
Still getting data corruption running 5 drives, and a kernel panic.
On Tuesday I had tried NexentaStor, OS 'silently' crashed after several hours of use (could enter credentials, but console became unresponsive), with both Windows boxes rebooting overnight due to MICROSOFT PATCH TUESDAY. Since the Windows boxes rebooted, I had no idea how much they had transferred to the Nexenta box before they had the rug pulled out from under them.
Subsequently I couldn't reboot the USB stick into Nexenta due to some 'device out of space' error during boot (I used a Sandisk 4GB usb key). So I have no idea if the Zpool drives were getting errors.
Code:
* Replaced 'no name' 500 watt power supply with branded 550 watt power supply.
Kernel panic screens seem to have vanished (but read on), Zpools are still suffering checksum errors when doing zpool scrub commands (NOTE WELL: All Scrubs were performed whilst the NAS was having data transferred to it)
Last night:
Code:
* Changed 'AUTO' settings on Memory, [.code]
set SPD manually to recommended values (5,5,5,15) and DRAM voltage to 2.0, underclocked RAM from 800Mhz to 667 Mhz. From advice read on Corsair forums.
[code]* reset FreeNAS 8.0.1 BETA 3 to Factory Defaults from Web GUI.
* Created new RAIDZ POOL of only three * 2TB drives, with remaining 2 set as 'SPARE'.
Slowly (using Crashplan) copied ~20GB of data to NAS,
ZPool Scrub reported no errors.
Subsequently hit NAS with as much as I could from two Windows 7 computers with straight file copies of large files, getting transfer rates of 600-700+Mbits/second over gigabit LAN.
Some 20 minutes and 60GB later, did a Zpool scrub, and I got 8 or so checksum errors on all three drives
Continued to hit NAS until transfers complete. Ended up with ~250GB data, and hundreds of Checksum errors on scrub (50-80 per drive) and 11 unrecoverable errors on large (monolithic 211GB Acronis TIB file) backup test file.
:-(
Left both Windows boxes hammering at NAS for an uptime test, it was still functioning 12 hours later.
Again, the Windows boxes both rebooted overnight (thanks for PATCH WEDNESDAY following on from PATCH TUESDAY, Microsoft!). Normally my Windows box is up for weeks at a time
After 12 hours 15 minutes, FreeNAS console reports a Kernel Panic, 12, Page fault in Kernel Mode, no process name apparent.
The box was (essentially) doing nothing at the time of the crash.
So now I'm writing to you folk, hoping that someone will read the story without cringing, without suffering ADD from reading such a lot, and withstanding the urge to shout out already-tried-suggestions:
This hardware was stable for -months- running Kubuntu in 2009 using the original 0301 BIOS firmware. Box has been turned off since mid 2010.
- ran MEMTEST86 for several hours, no errors (pre RAM underclock).
- motherboard is clean (no dust anywhere)
- connectors are correctly plugged in, to drives and motherboard
- memory is in pairs (no warnings about dual sided DIMMS or using all four slots, from Asus site)
- changed to 'brand name' power-supply with actual guts in it.
- underclocked memory and slightly bumped up voltage.
- reduced number of active drives in pool from 5 to 3, to avoid overloading the SATA controller/bus with data
More information (same stuff, as-it-happened reports)
here
and here
and here
Next step -- swap 3 * 2TB Hitachis for 3 * 1TB Western Digital drives in drive cage and retest. Although I can't see how faulty -data only- hard disks could be sending FreeBSD into Kernel Panics.
TLDR version: My FreeBSD/FreeNAS computer keeps getting Kernel Panics and corrupting my ZFS drives constantly, help!
Apologies for vast first post to forum
Last edited: