• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

mdadm RAID 5 problem

sonicz

n00b
Joined
May 21, 2013
Messages
28
Hi There

I'm running a EXT4 raid with 4x3TB disks in RAID 5. I wanted to grow with one more 3TB disk. After I resized everything seemed fine, until a made a reboot.

One of the old disks now stands with F (failure) 40 bad sectors and the new one stands as S a spare disk.

I tried to force the disks online again and rebuild the raid, but it fails around 75-80% and now it stands the same way.

Just to make sure, I'm out of luck here right? - I can't really find any work arounds.

So if you clever guys think it is a dead array too, should I go for EXT4 again or something else? - I'm planning to keep growing untill I hit 10x 3TB and as far I can see I will hit a wall @16TB with EXT4 right?

Help will be greately appreciated.
 
try booting a debian install disk in recovery mode, it will see that the disks are soft raided and may correctly detect, and repair or reconfigure if needed, sometimes a clean OS works better to repair a soft array.

what distro are you using? how are you monitoring the softraid? EXT4 and MDADM are great, for software RAID.
 
try booting a debian install disk in recovery mode, it will see that the disks are soft raided and may correctly detect, and repair or reconfigure if needed, sometimes a clean OS works better to repair a soft array.

what distro are you using? how are you monitoring the softraid? EXT4 and MDADM are great, for software RAID.

I'm running a Ubuntu 12.10. Not sure what you mean by monitoring?

Yeah I'm also happy with the idea of EXT4, but just been reading that you can't resize a array to over 16TB, because of some bug in e2fsprogs/resizefs?
 
do you just run mdadm --detail /dev/md0 occasionally to check on the array? or do you use webmin plugins or something?

im not sure if it is a bug or more of a limitation of the combination but yes about 16TB.

im not a fan of ubuntu for political reasons, but the debian install disk will work on it fine, a ubuntu install might work, i dunno how it handles the installer though.
 
What were the commands you used to grow the array? The new drive should not be a spare if you added it correctly and the reshape completed succesfully. Check if you can examine the new drive with mdadm -E /dev/sd? does it give you information?

An EXT4 volume can be created to expand over 16TB when using the "64bit" feature. But, the feature has to be used during creation of the file system.
 
What were the commands you used to grow the array? The new drive should not be a spare if you added it correctly and the reshape completed succesfully. Check if you can examine the new drive with mdadm -E /dev/sd? does it give you information?

An EXT4 volume can be created to expand over 16TB when using the "64bit" feature. But, the feature has to be used during creation of the file system.

I will get back to you on this, I''m just trying to recover the raid, from recovery mode.

It is just a last shot before giving up on this array.
 
If the expansion really did complete you should be able to recover the array. Please post a 'mdadm --examine' of all member devices, a 'mdadm --detail' of the RAID device (if possible) and a 'cat /proc/mdstat'. If you already forced a reassemble the recovery mode may not help at this point.

For arrays with these sizes I use XFS or now BTRFS.
 
Well it seems the recovery worked, sort of.

Personalities : [linear] [multipath] [raid0] [raid1] [raid6] [raid5] [raid4] [ra
id10]
md0 : active raid5 sdc1[5] sdg1[2] sdd1[0] sde1[1] sdf[3]
11720534016 blocks super 1.2 level 5, 512k chunk, algorithm 2 [5/5] [UUUUU
]

unused devices: <none>

But the filesystem seems to be gone?

Here is the examine:

/dev/sdg1:
Magic : a92b4efc
Version : 1.2
Feature Map : 0x0
Array UUID : 177b9582:09bfa016:4b5cbb51:d8df5c86
Name : (none):0
Creation Time : Tue May 21 17:32:21 2013
Raid Level : raid5
Raid Devices : 5

Avail Dev Size : 5860268032 (2794.39 GiB 3000.46 GB)
Array Size : 11720534016 (11177.57 GiB 12001.83 GB)
Used Dev Size : 5860267008 (2794.39 GiB 3000.46 GB)
Data Offset : 262144 sectors
Super Offset : 8 sectors
State : clean
Device UUID : b15519ba:5519bbfb:3220a6d8:0a95903c

Update Time : Tue May 21 23:47:13 2013
Checksum : e1cf3068 - correct
Events : 18

Layout : left-symmetric
Chunk Size : 512K

Device Role : Active device 2
Array State : AAAAA ('A' == active, '.' == missing)

/dev/sdc1:
Magic : a92b4efc
Version : 1.2
Feature Map : 0x0
Array UUID : 177b9582:09bfa016:4b5cbb51:d8df5c86
Name : (none):0
Creation Time : Tue May 21 17:32:21 2013
Raid Level : raid5
Raid Devices : 5

Avail Dev Size : 5860268032 (2794.39 GiB 3000.46 GB)
Array Size : 11720534016 (11177.57 GiB 12001.83 GB)
Used Dev Size : 5860267008 (2794.39 GiB 3000.46 GB)
Data Offset : 262144 sectors
Super Offset : 8 sectors
State : clean
Device UUID : 2c2a2036:b319d51c:af203b09:b7bd2e5d

Update Time : Tue May 21 23:47:13 2013
Checksum : d0232e6c - correct
Events : 18

Layout : left-symmetric
Chunk Size : 512K

Device Role : Active device 4
Array State : AAAAA ('A' == active, '.' == missing)

/dev/sdd1:
Magic : a92b4efc
Version : 1.2
Feature Map : 0x0
Array UUID : 177b9582:09bfa016:4b5cbb51:d8df5c86
Name : (none):0
Creation Time : Tue May 21 17:32:21 2013
Raid Level : raid5
Raid Devices : 5

Avail Dev Size : 5860268032 (2794.39 GiB 3000.46 GB)
Array Size : 11720534016 (11177.57 GiB 12001.83 GB)
Used Dev Size : 5860267008 (2794.39 GiB 3000.46 GB)
Data Offset : 262144 sectors
Super Offset : 8 sectors
State : clean
Device UUID : bdec3cea:c86c3dd3:877d7de0:a3fdac9e

Update Time : Tue May 21 23:47:13 2013
Checksum : 5368e0d4 - correct
Events : 18

Layout : left-symmetric
Chunk Size : 512K

Device Role : Active device 0
Array State : AAAAA ('A' == active, '.' == missing)

/dev/sde1:
Magic : a92b4efc
Version : 1.2
Feature Map : 0x0
Array UUID : 177b9582:09bfa016:4b5cbb51:d8df5c86
Name : (none):0
Creation Time : Tue May 21 17:32:21 2013
Raid Level : raid5
Raid Devices : 5

Avail Dev Size : 5860268032 (2794.39 GiB 3000.46 GB)
Array Size : 11720534016 (11177.57 GiB 12001.83 GB)
Used Dev Size : 5860267008 (2794.39 GiB 3000.46 GB)
Data Offset : 262144 sectors
Super Offset : 8 sectors
State : clean
Device UUID : 97760fb3:8abce147:1f22456f:f9ce47d0

Update Time : Tue May 21 23:47:13 2013
Checksum : 5142305e - correct
Events : 18

Layout : left-symmetric
Chunk Size : 512K

Device Role : Active device 1
Array State : AAAAA ('A' == active, '.' == missing)

/dev/sdf:
Magic : a92b4efc
Version : 1.2
Feature Map : 0x0
Array UUID : 177b9582:09bfa016:4b5cbb51:d8df5c86
Name : (none):0
Creation Time : Tue May 21 17:32:21 2013
Raid Level : raid5
Raid Devices : 5

Avail Dev Size : 5860271024 (2794.40 GiB 3000.46 GB)
Array Size : 11720534016 (11177.57 GiB 12001.83 GB)
Used Dev Size : 5860267008 (2794.39 GiB 3000.46 GB)
Data Offset : 262144 sectors
Super Offset : 8 sectors
State : clean
Device UUID : 79fdb09e:529cfd0a:2302337b:44450ced

Update Time : Tue May 21 23:47:13 2013
Checksum : 28b1f909 - correct
Events : 18

Layout : left-symmetric
Chunk Size : 512K

Device Role : Active device 3
Array State : AAAAA ('A' == active, '.' == missing)


sudo mdadm --detail
mdadm: No devices given.
 
sudo mdadm --detail
mdadm: No devices given.

You wanted to do sudo mdadm --detail /dev/md0

Also did you really want /dev/sdf in your array and not /dev/sdf1?

Not that using raw devices is a problem but its best to stay consistent.

But the filesystem seems to be gone?

Is it that it did not automount or did you try to mount it or you used fsck?
 
You wanted to do sudo mdadm --detail /dev/md0

Also did you really want /dev/sdf in your array and not /dev/sdf1?

Not that using raw devices is a problem but its best to stay consistent.



Is it that it did not automount or did you try to mount it or you used fsck?

yeah I wanted sdf, by mistake I added a raw device, but sholdn't be a problem.

When I try to mount it tell me that it does not exist.

I think its time for me to say farewell to my data and start reading up on xfs and make a new raid.
 
I think its time for me to say farewell to my data and start reading up on xfs and make a new raid.

I would try fsck before giving up.

sudo fsck.ext4 -v -f /dev/md0

When I try to mount it tell me that it does not exist.

Are you sure you are mounting with the correct device /dev/md0
 
Should I try the suggestion?


e2fsck 1.42.5 (29-Jul-2012)
ext2fs_open2: Bad magic number in super-block
fsck.ext4: Superblock invalid, trying backup blocks...
fsck.ext4: Bad magic number in super-block while trying to open /dev/md0

The superblock could not be read or does not describe a correct ext2
filesystem. If the device is valid and it really contains an ext2
filesystem (and not swap or ufs or something else), then the superblock
is corrupt, and you might try running e2fsck with an alternate superblock:
e2fsck -b 8193 <device>
 
I found a post where people had good results with mdadm --create with the exact same disks and options as the old raid. It then comes and ask If you wanna recover, because It can see the existing raid.
 
Ohh it seem BTRFS does not support RAID 5 yet.

btrfs actually does have builtin raid 5 and raid6 in the 3.9.0 kernel and above (I am testing that on unimportant data). Although I have used btrfs on top of mdadm raid6 for over 1 year and > 20 to 30 TB of data with no problems at all.
 
I found a post where people had good results with mdadm --create with the exact same disks and options as the old raid. It then comes and ask If you wanna recover, because It can see the existing raid.

Did you use the --assume-clean parameter? If not you likely trashed your array.
 
btrfs actually does have builtin raid 5 and raid6 in the 3.9.0 kernel and above (I am testing that on unimportant data). Although I have used btrfs on top of mdadm raid6 for over 1 year and > 20 to 30 TB of data with no problems at all.

Okay not sure that is something I should be giving myself into.. As you proberly have guessed I'm no shark in Linux. Trying to learn my way though :)


Must see if I can find some guides about how to do, else I guess xfs is a good alternative?
 
I guess xfs is a good alternative?

I believe it is. There used to be some performance issues on xfs with small files and folders with thousands of files however I believe these are all solved by now.

As you proberly have guessed I'm no shark in Linux. Trying to learn my way though

I recommend that you test whatever mdadm + filesystem you choose before putting real data onto the system. Expand an array, hot pull a disk by pulling out the sata cable on a raid member when the raid is up to learn about recovery. This has helped me a lot when I have had problems. I have been able to successfully recover from a raid 6 array having 3 disk kicked out of the array at work (I did have a backup but I use backups as a last resort).


Also I highly recommend going to raid6 especially when you increase the # of disks in your raid.
 
Last edited:
think of these things as layers. actual disk, MDADM (software RAID), filesystem, data.

then like they said, PRACTICE. grab some old/small drives and build and destroy the raid, attempt recovery, add a drive, remove a drive, test a hotspare.

i like raid 5 for about 4-8 drives, raid 6 for 8+ drives. but then performance can change with duty, drives, filesystems, etc. so the best thing to do is to test it in a scenario as similar to your real world use as possible.


at one of my previous jobs i managed a server that had 24 1tb drives in a SINGLE raid 5 array. with the OS and storage all on that one array. not even seperate pools, partitions, nothing. just a lump of roughly 20tb with the OS and everything, whoever set that thing up should have been slapped.
 
at one of my previous jobs i managed a server that had 24 1tb drives in a SINGLE raid 5 array. with the OS and storage all on that one array. not even seperate pools, partitions, nothing. just a lump of roughly 20tb with the OS and everything
That is terrible on so many levels. Well at lest they did not use raid0 or a large lvm span..

whoever set that thing up should have been slapped.

I would wish for more drastic punishment. :D
 
I'm currently preparing for some tests and I found this command, that should form a ext4 64bit system.

But I wanted to here if someone in here can help with a bit of understanding.

# mke2fs -O 64bit,has_journal,extents,huge_file,flex_bg,uninit_bg,dir_nlink,extra_isize -i 4194304 /dev/vg0/lv_data

4194304 is this a variable number or should it always be this? - the reason I ask this is because I got a error when I tried to execute it on a single disk, just to try it out.

Not sure if this is the right way to go, but I really would like to stay with ext4 if possible, because I've been a fan of the performance so far.

Found this about that option, but it doesn't make much sense to me:

-i bytes-per-inode
Specify the bytes/inode ratio. mke2fs creates an inode for every bytes-per-inode bytes of space on the disk. The larger the bytes-per-inode ratio, the fewer inodes will be created. This value generally shouldn't be smaller than the blocksize of the filesystem, since in that case more inodes would be made than can ever be used. Be warned that it is not possible to expand the number of inodes on a filesystem after it is created, so be careful deciding the correct value for this parameter.
 
Last edited:
Ohh well after some more reading, I'm not sure I will be able to grow the array later on >16TB with ext4 64bit.
 
Heres a update:

I've decided to go with xfs. Yesterday I made a raid 5 array with 3 disks. Afterwards I made a grow with one disk, just to check if it would go well. And it did.

But.. Everything looks fine except when I try to run a check on the filesystem:

xfs_check -f /dev/datastore/datastore
ERROR: The filesystem has valuable metadata changes in a log which needs to
be replayed. Mount the filesystem to replay the log, and unmount it before
re-running xfs_check. If you are unable to mount the filesystem, then use
the xfs_repair -L option to destroy the log and attempt a repair.
Note that destroying the log may cause corruption -- please attempt a mount
of the filesystem before doing this.

Any ideas? - I can with no problem mount the system and read/write from it.
 
Did you unmount the XFS FS before you run xfs_check? Is is supposed to work on unmounted volumes.
 
can someone please tell me what I'm doring wrong here :)

I added the fourth disk and did go well. I made sure to update the mdadm.conf and double checked that it had updated.

When I booted it had thrown the new added disk out of the raid? - Can someone tell me what I am forgetting here? - I have tried it before. Here is what I do from what I have found in diffrent guides:

sudo -i

parted /dev/sdx

mklabel gpt

unit TB

mkpart primary 0.00TB 3.00TB

print

mkfs.xfs /dev/sdxx

mdadm --add /dev/md0 /dev/sdxx

echo 50000 > /proc/sys/dev/raid/speed_limit_min
echo 200000 > /proc/sys/dev/raid/speed_limit_max

watch cat /proc/mdstat


mdadm --grow /dev/md127 --raid-devices=x?

umount /dev/md0 | umount /dev/datastore/datastore

pvresize /dev/md127

lvresize -l 100%VG /dev/datastore/datastore

lvdisplay

xfs_growfs /dev/datastore/datastore

check diskspace

df -h


Hope for some input, so I don't keep making the same mistakes, it is not well for my raid, which is currently rebuilding.
 
I would like to share a lesson that I learned with XFS on logical volumes. Make sure you follow the xfs.org faq and manually set the agcount, sunit and swidth option at creation, when using it with a logical volume. While its true the defaults are fine when using the mdadm device directly, the default settings are real bad when using the lvm logical volume. I went with the defaults and it creates a filing system with 25954 agcount, which takes roughly two hours to file system check. Should of been around 16.

http://xfs.org/index.php/XFS_FAQ#Q:_How_to_calculate_the_correct_sunit.2Cswidth_values_for_optimal_performance
 
I would like to share a lesson that I learned with XFS on logical volumes. Make sure you follow the xfs.org faq and manually set the agcount, sunit and swidth option at creation, when using it with a logical volume. While its true the defaults are fine when using the mdadm device directly, the default settings are real bad when using the lvm logical volume. I went with the defaults and it creates a filing system with 25954 agcount, which takes roughly two hours to file system check. Should of been around 16.

http://xfs.org/index.php/XFS_FAQ#Q:_How_to_calculate_the_correct_sunit.2Cswidth_values_for_optimal_performance


Sounds to me like xfs is not very raid capable then? - or can you change these values without wiping the disks? - Because I'm currently using 5 disks, but will grow 1 by 1 until around 10 disks and then the values will be off again?
 
Back
Top