FreeNAS build problems -- kernel panics

Excessionoz

Weaksauce
Joined
Jun 28, 2011
Messages
111
UPDATE: SOLVED

Disabled on-board NIC
Installed Intel NIC

System behaves perfectly, no checksum errors,no crashes on large network transfers.

CASE CLOSED.

Thankyou to all for your various suggestions.

FreeNAS, P5E-VM-HDMI and headaches galore.

Longwinded versions already posted over on Ars Technica:

The Plan: Keen Plan
The Test: The Test
The FAIL: The Fail

The Original Hardware:

Code:
Motherboard:  ASUS P5E-VM-HDMI (6 * SATA) GbitE 
CPU:          Intel A-E8200 Core 2 Duo @ 2.66Ghz
RAM:          Corsair XMS2 TWIN2X2048-6400C4 * 2 ( 4 * 1GB DIMM) 4GB
HD:           3 * Western Digital WDEACS 1TB drives
DVDROM:       SATA DVD reader.
HDDCase:      SUPERMICRO CSE-M35T-1B Black 5x 3.5" Hot-swap SATA 
USB Key       Crucial 4GB stick (for FreeNAS boot).

The 'Upgraded' Hardware
Code:
RAM:          Corsair VS4GBKIT4096-6400C5 * 2 ( 4 * 2GB DIMM) 8GB
HD:           5 * Hitachi 5K3000 2TB Deskstar drives
SSD:          64GB G.Skill SSD

First step was testing FreeNAS 8.0 (Release) on the original hardware. It seemed to work but NOTE WELL, after transferring about 200GB of files to my test RAIDZ1 pool, I never did a "zpool scrub <pool>" command, so I don't know if there were any of these Checksum errors that my later testing came up with on different drives.

Second test, new configuration, put together

Code:
5 * 2TB Hitachi drives 
1 * 64GB 'G.Skill' SSD I had lying around, as a CACHE drive in the RaidZ1 array.

After about 30 minutes the SSD threw a fit during one of the operations (it was already a highly suspect SSD, but that is a different tale), more importantly, 10 minutes after that, Kernel Panic (trap 9) and the machine stopped dead.

Next day,

Code:
* Replaced the 4GB of PC6400C4 (4 * 1GB) Corsair RAM with 8GB of PC6400C5 (4 * 2GB) Corsair RAM.

* Changed drives in BIOS from IDE to AHCI

Restarted FreeNAS 8.0 (Release) from scratch, subsequently had many crashes, kernel panic trap 9's kernel panic trap 12's, with uptimes averaging 2.5->3 hours.

detailed post on ars

Updated to a nightly build of FreeNAS 8.0.1 (BETA3). Similar results, crashes and data corruption on the data pool (checksum errors on scrub, unrecoverable data errors revealed in pool).

Code:
* Updated P5E-VM-HDMI BIOS from 2007 (0301) to latest available (0709).

Still getting data corruption running 5 drives, and a kernel panic.

On Tuesday I had tried NexentaStor, OS 'silently' crashed after several hours of use (could enter credentials, but console became unresponsive), with both Windows boxes rebooting overnight due to MICROSOFT PATCH TUESDAY. Since the Windows boxes rebooted, I had no idea how much they had transferred to the Nexenta box before they had the rug pulled out from under them.

Subsequently I couldn't reboot the USB stick into Nexenta due to some 'device out of space' error during boot (I used a Sandisk 4GB usb key). So I have no idea if the Zpool drives were getting errors.

Code:
* Replaced 'no name' 500 watt power supply with branded 550 watt power supply.

Kernel panic screens seem to have vanished (but read on), Zpools are still suffering checksum errors when doing zpool scrub commands (NOTE WELL: All Scrubs were performed whilst the NAS was having data transferred to it)

Last night:

Code:
* Changed 'AUTO' settings on Memory, [.code]

set SPD manually to recommended values (5,5,5,15) and DRAM voltage to 2.0, underclocked RAM from 800Mhz to 667 Mhz.  From advice read on Corsair forums.

[code]* reset FreeNAS 8.0.1 BETA 3 to Factory Defaults from Web GUI.

* Created new RAIDZ POOL of only three * 2TB drives, with remaining 2 set as 'SPARE'.

Slowly (using Crashplan) copied ~20GB of data to NAS,

ZPool Scrub reported no errors.

Subsequently hit NAS with as much as I could from two Windows 7 computers with straight file copies of large files, getting transfer rates of 600-700+Mbits/second over gigabit LAN.

Some 20 minutes and 60GB later, did a Zpool scrub, and I got 8 or so checksum errors on all three drives :(

Continued to hit NAS until transfers complete. Ended up with ~250GB data, and hundreds of Checksum errors on scrub (50-80 per drive) and 11 unrecoverable errors on large (monolithic 211GB Acronis TIB file) backup test file.

:-(

Left both Windows boxes hammering at NAS for an uptime test, it was still functioning 12 hours later.

Again, the Windows boxes both rebooted overnight (thanks for PATCH WEDNESDAY following on from PATCH TUESDAY, Microsoft!). Normally my Windows box is up for weeks at a time :(, so I had no idea if they had problems 'writing' to the NAS at any time.

After 12 hours 15 minutes, FreeNAS console reports a Kernel Panic, 12, Page fault in Kernel Mode, no process name apparent.

The box was (essentially) doing nothing at the time of the crash.

So now I'm writing to you folk, hoping that someone will read the story without cringing, without suffering ADD from reading such a lot, and withstanding the urge to shout out already-tried-suggestions:

This hardware was stable for -months- running Kubuntu in 2009 using the original 0301 BIOS firmware. Box has been turned off since mid 2010.

  • ran MEMTEST86 for several hours, no errors (pre RAM underclock).
  • motherboard is clean (no dust anywhere)
  • connectors are correctly plugged in, to drives and motherboard
  • memory is in pairs (no warnings about dual sided DIMMS or using all four slots, from Asus site)
  • changed to 'brand name' power-supply with actual guts in it.
  • underclocked memory and slightly bumped up voltage.
  • reduced number of active drives in pool from 5 to 3, to avoid overloading the SATA controller/bus with data

More information (same stuff, as-it-happened reports)

here
and here
and here

Next step -- swap 3 * 2TB Hitachis for 3 * 1TB Western Digital drives in drive cage and retest. Although I can't see how faulty -data only- hard disks could be sending FreeBSD into Kernel Panics.

TLDR version: My FreeBSD/FreeNAS computer keeps getting Kernel Panics and corrupting my ZFS drives constantly, help!

Apologies for vast first post to forum :) Hope it's in the 'right' place.
 
Last edited:
I was (am) running v9 Freebsd, with a similar-sized MSI board, and putting in an aftermarket pcie video card AFAIK was the key to evaporating seemingly-hourly kernel panics. Should not be, or "could not be", but that is the first thing I'd try in this case; OTOH I've never used ZFS.
 
What is the current setup? Are you still using the SSD for ZFS cache. I suspect some flaky hardware somewhere, kernel panics and ZFS corruption means something is going on somewhere.
 
One observation not related to the crashes. A 3-disk raidz with 2 spares doesn't make a lot of sense. If you are really paranoid, you'd be better off with a 4-disk raidz2 and 1 spare. Same amount of usable data and less exposure to a double failure. I'd actually rather see a 2x2 raid10 with 1 spare.
 
One observation not related to the crashes. A 3-disk raidz with 2 spares doesn't make a lot of sense. If you are really paranoid, you'd be better off with a 4-disk raidz2 and 1 spare. Same amount of usable data and less exposure to a double failure. I'd actually rather see a 2x2 raid10 with 1 spare.

He started out initially with a 5 disk raid-z1 and was paring it down to take drives/ports out of the equations (that said, your point is of course totally valid)
 
Last night I did a test of three Western digital drives and two Hitachi drives.

The purpose of the test was to determine if the Hitachi hard disks I had just purchase were defective, by actively comparing them to the Western Digital HDs which were known to be 'ok'.

At the same time, by splitting up the drives into individual pools, I was testing if individual SATA cables/drives were possibly responsible for the reported checksum errors I had been getting on the 5 disk RAIDZ array.

The P5E-VM-HDMI (BIOS 0709) motherboard had the disk controller set to AHCI.

I created four pools, one mirrored on two WD drives, and three single drives.

Copied a few tens of GBytes of data across, spread across the four pools, then performed a zpool scrub operation on each pool.

Every one of the pools, and every drive, had a number of checksum errors, the more data, the larger the number of checksum errors.

* I changed the BIOS setting from AHCI to IDE.

Created a three drive WD pool (Raidz1), and a two drive Mirror with the Hitachi's, and repeated the copying of files.

This time zpool scrub did not flag any checksum error on either pool.

So the ongoing data corruption issues seem to have been an artifact of running the on board controller in AHCI mode vs IDE mode.


Performance, measured in transfer rates across the Gigabit link, was 'about the same' (except with zero errors this time).

Unfortunately, the stability of the system is still low -- in two hours I got two Kernel Panic Trap 9's (general protection fault in kernel mode) crashes.

Previously I had upped the voltage of my memory from 1.92V to 2.0V (as per Corsair forum suggestions when running 4 DIMMs, here (The Ram Guy, Corsair forum, and I was also running the 6400C5 memory at 667Mhz instead of the nominal 800Mhz, manually setting the clock speed values in the BIOS.

The C2D E8200 CPU is running at 1.825V (2.66Ghz, not overclocked). It's running at about 40C according to the BIOS hardware monitor.

Even though I'm quite happy at locating the source of the 'mysterious' HD data corruption issue (AHCI vs IDE), I'm at a loss as to where to turn next due to the constant kernel panics with FreeBSD.

The PCI-E video card suggestion is perhaps pertinent, but the only replacement cards I have on hand, are ATI HD4850's requiring additional power supply oomph, generating, heat; I planned to run the box headless, so having another heat-generating, fan-whirring behemoth video card in the case is counterproductive. The on-board video (VGA) works fine.

Next plan of attack is to reduce the four DIMMs to two DIMMs, and see if that improves stability.

What I wanted was a reliable 'minimal maintenance' NAS. What I have is an unusable machine, one that had performed fine for -months- at a time, as a Kubuntu Apache Web Cache.
 
I think what you have is a faulty motherboard that you never knew had problems until ZFS rooted it out for you. That board should be supported fine by FreeBSD (FreeNAS) and if ZFS shows corruption, you have hardware issues somewhere. I have always used AHCI mode for SATA drives with ZFS and never seen anything like you're experiencing, switching to IDE mode is probably just putting a band-aid over whatever the problem is. A few more things to test: swap SATA cables, and if you're familiar with FreeBSD by itself try running 9.0-CURRENT, there are a lot of hardware support changes, new methods, and it has ZFS v28.
 
Now that my ZFS volumes aren't reporting checksum errors (after the AHCI->IDE change), I can't see the benefit of swapping out SATA cables. Can you "talk me into it?" with some more information? The basis of my doubt is that I can't fathom the link between a flaky SATA drive/port, and Kernel Panics.

edit spoke too soon: zpool errors galore

Thankyou for your suggestions though.

Someone else has suggested that 'tidying / aligning the cables, like a neat-freak, can lead to signal induction in the SATA cables', and mine are neat-freaked, but again, so what, if my SATA drives are now not reporting any checksum errors, why should I worry about the cables being aligned?

The builds I see people running are mostly rather tidy, rather than random cables snaking all over the place.
 
Last edited:
If you actually read you'll find in the text that he also tried NexentaStor and also experienced corruption with it which uses Solaris/OpenSolaris so it isn't a FreeBSD issue. ZFS on 8.0 isn't that great to be honest, 8-STABLE works a lot better although I haven't run into any issues on 9-CURRENT either. I can confirm though that 3-4Gb can at least on FreeBSD be a bit on the low side without any tuning esp if you're running 8.1 and older and will result in panics after a while. While the P5E-VM HDMI isn't ideal for FreeBSD at least it should work in general, I can confirm that the P5E-VM DO works great but it has different NIC, southbridge and BIOS. AHCI will always be prefered over IDE emulation so don't switch. So in conclusion, I (also) think your motherboard is flaky and needs to be replaced possibly damaged by your first PSU.

If you want recommendations I can confirm that Intel DG45ID and DQ45CB works fine and very well with your other hardware (memory etc) and with FreeBSD. A nice little bonus is that both card also have Intel NICs which boosts performance a bit. Have in mind though that both these cards only have 5 internal SATA connectors. Both these boards are available over at amazon.com and if you prefer Newegg the ASUS P5Q-EM DO should work fine but I cannot guarantee it. The Asus mobo also have 6 internal SATAs.

Intel DG45ID, 4 * 2Gb Corsair Value Select PC6400 RAM, IBM M1015 RAID Adapter (flashed to a LSI 9210-8i HBA otherwise it wont work at all in this mobo), 2 ASMedia ASM1061 PCIe Controller cards, One (old) Intel PCI NIC (100mbit) running FreeBSD 9-CURRENT about a month old snapshot and two RAIDZ arrays.

It also sounds really strange that you would need to bump the voltage up 0.2V from stock 1.8V on stock RAM...

//Danne
 
Last edited:
Another thread suggests "used ati hd 3870" == $30-40. *IF* that fixes the kernel panics, it may be worth the cost; if not, not, but it probably would not use that much power. ( Less power/less heat/cheaper, the g70-g71 pcie chipset boards from circa 2008 ?)
 
Why would a video card help unless the motherboard itself is faulty?
//Danne
 
Frees up more memory for the Os/Zpool etc? Some snippet of code writing to/from the video memory space? (Just guessing!). I mean not to write that it *would* help, but anecdotally *might*... in my case it was a great relief for that problem, though not zfs(zpool etc) related.
 
The story continues.

After it was revealed that the on board controller state (AHCI vs IDE) was a furphy, I went back to scratch.

Cleared out BIOS to be absolutely stock (AUTO on everything).

Ran MEMTEST86 for 14 hours (it eventually stops testing all by itself) reporting 0 errors in my 8GB memory.

Booted up FreeNAS 8.0.1 Beta 3 from a USB key, and configured my array as follows

mixed1 : Hitachi 1.8 + WD .9 ZFS Stripe, giving 2.9TB space (no redundancy)
mixed2 : Hitachi 1.8 + WD .9 ZFS Stripe, giving 2.9TB space (no redundancy)
single1 : WD .9TB ZFS, giving .9TB space (no redundancy)

Did some tests with about 300GB of files copied to each pool, just using dd and cp (via an SSH shell on the FreeNAS box)

zpool scrub performed on all three arrays, with zero errors reported.

Then I created a 1.1TB file on each of mixed1 and mixed2 volumes over about 3 hours, and a further 750GB file on single1 array.

Uptime has been 6 hours, zero errors reported so far. zpool scrub operation is occurring, and is estimated to take another 2.5 hours on the two 'mixed' arrays.

Currently zpool iostat is reporting as follows

Code:
mixed1      1.30T  1.42T  1.29K      0   165M      0
mixed2      1.31T  1.41T  1.32K      0   168M      0
single1      761G   167G      0    457      0  57.2M

With mixed1 and mixed2 performing the scrub, and single1 still creating the 750GB file:

Code:
freenas# dd if=/dev/zero of=/mnt/single1/zfrandom7 bs=10m count=71999

The MAJOR difference is that I'm doing all data creation/copying on the machine itself, without copying data across the NETWORK.

The on-board network card is an Atheros chipset, and that just might be the problem, but I don't know. Certainly without any -network- component, it's going to be useless to me :). I've got a 100Mbit/sec NetGear PCI network card I can plop into the machine, and see if network file transfers clobber the machine, however, 1/10th the network speed (!00Mbit vs 1Gbit) is hardly going to 'stress' the machine. < shrug >
 
I missed the part about trying nexenta, sorry. Yeah, this all sounds really weird. If you are running with vanilla bios settings, maybe you could try again with network xfer and see?
 
The on-board network card is an Atheros chipset, and that just might be the problem, but I don't know. Certainly without any -network- component, it's going to be useless to me :). I've got a 100Mbit/sec NetGear PCI network card I can plop into the machine, and see if network file transfers clobber the machine, however, 1/10th the network speed (!00Mbit vs 1Gbit) is hardly going to 'stress' the machine. < shrug >

Worth a shot. I've personally had Broadcom gigE onboard NIC's go bad and FreeBSD kernel panics, they start spewing garbage all over the bus and things go downhill fast. Now I just put Intel PRO1000's in all my machines and never had any NIC related problems.
 
I reconfirmed that network traffic would kill the box, by mount-smbfs from the FreeNAS box to my windows machine, then pulling data across the network (instead of having windows shovelling data down). After ~35Gbytes, FreeBSD kernel panic trap 12 occurred, yay! Doing a zpool scrub on the array in question was fine for 95% of the scrub, and only in the last couple of minutes did any zpool status errors show up -- all on files in the Wincopy directory I was copying to.

Tried to disable the on-board NIC and put an ancient NetGear 100Mbps card in, but the motherboard wouldn't complete POST (Couldn't get into BIOS), so I've given up for the weekend.

Ordered a PCI-E Intel NIC, and will test during the week.

tdg, thanks for the confirmation that NIC's going bad cause random kernel panics. It helps heaps to have other people acknowledge such things.
 
Not just NICs, I think. It may be that a hardware fault accessing device registers and such causes the random panics, but yeah, in your case, I think that's the next move...
 
Case Closed.

Disabled P5E-VM-HDMI Atheros on-board NIC.

Installed Intel PCI-E Gigabit NIC.

Booted up, created pool, transferred 550GB with no zpool scrub errors. Transmitted 200GB back to other windows machine, with no errors.

Uptime now three hours.

Will be doing an acid-test of file transfers from both windows systems overnight (unless Sir Microsoft throws a reboot at us,because, you know, it's Wednesday and Updates Must Be Done). Zpool scrub in morning will tell me if there are any new errors after a few hundred thousand files have moved across.

Performance is great -- 70-80Mbytes/sec write, 85Mbytes/sec read, over single gigabit NIC through a cheapie Dlink switch.

Colour me relieved.
 
Back
Top