• A Great friend to the HardForum with a great kid that he is trying to get a scholorship to continue his schooling. Please give hime a vote! Only 24 hours left! Thanks.
    If you have an VOTE FOR KEENAN!

ZFS getting insufficient replicas trying to replace drive

ahnooie

n00b
Joined
Dec 10, 2011
Messages
23
I'm running a VMware Napp-It All-In-One on ESXi and OmniOS. Don't know if it matters but my setup is a Gen8 HP Microserver, using an IBM Serverraid M1015 in IT mode. I'm not passing the controller to OmniOS, instead I created a vmdk for OmniOS on each drive.

I had a drive failure, zpool status said that "c2t2d0 FAULTED 0 191 0 too many errors.

I ordered a new and just replaced the bad drive with it last night... I removed the vmdk on the bad drive from OmniOS and created a vmdk identical in size on the new drive added it to OmniOS. In VMware all the drives are up and accessible. I ran a zpool replace tank c2t0d0 ... it started to resilver. This morning I checked on it and I'm seeing this error:

pool: tank
state: DEGRADED
scan: scrub repaired 0 in 2h49m with 0 errors on Sat Mar 29 05:46:24 2014
config:

NAME STATE READ WRITE CKSUM
tank DEGRADED 0 0 0
raidz1-0 DEGRADED 0 0 0
c2t1d0 ONLINE 0 0 0
replacing-1 UNAVAIL 0 0 0 insufficient replicas
c2t2d0/old FAULTED 0 0 0 too many errors
c2t2d0 FAULTED 0 0 0 too many errors
c2t3d0 ONLINE 0 0 0
logs
c2t4d0 ONLINE 0 0 0
cache
c2t5d0 ONLINE 0 0 0

I tried running the replace command again but the same thing, starts to resilver then results in this screen.

I've also tried rebooting the physical server as well as the OmniOS VM to no avail.

As far as I can tell the pool is functioning normally, I have other VMs running fine and my CIFS shares are still working.

I also tried to see if I could do anything in the Napp-It gui but it just hangs when I click on "pools" or "disks". It's stuck at the "processing, please wait..."

So what do I need to do to? I have a full local backup so if wiping the pool and re-creating it is the only way I could do that... but it seems like that shouldn't be necessary.

Thanks.
 
You have a faulted disk that you replaced - I suppose in the same slot.
During replace the new disk had problems as well.

Now you must check if the new disk is bad as well or if you have another problem like bad backplane, power or cablings.

As napp-it stalls with menu disk or pool where it requests a disklist with commands like iostat or format, the faulted disk is blocking the bus or controller.

Remove the affected disks and restart and check if the other disks are fine and napp-it displayes them. (Pool is offline due to missing disks). Now add the first faulted disks to check behaviours. Pool goes to online or degraded if enough disks come back.

Other options:
-You can clear the too many errors with menu pools clear errors (zpool clear)
-If possible, insert new new disk in a new slot and do a replace faulted -> new
-If possible, move to a pass-through config, where ZFS has full control of disks, disk cache, smartvalues or full hotplug capability and use unique WWN id's.
- If pass-through is not possible, use single disks via RDM - avoid ZFS over vmdk beside the disk for OmniOS
- use Raid 10 instead of a 4 disk Z1. Performance is better and up to two disks in different vdevs can fail.
- Check the disks with a manufacurers lowlevel tool (mostly Windows)
 
Last edited:
I think you are right that their is a problem with the replacement disk or cabling, while I was looking at the storage in VMware this morning the capacity went blank and then the drive vanished.

I'll switch over to my mirror and when I rebuild go back to pass-through.

Thanks, Gea!
 
Back
Top