• A Great friend to the HardForum with a great kid that he is trying to get a scholorship to continue his schooling. Please give hime a vote! Only 24 hours left! Thanks.
    If you have an VOTE FOR KEENAN!

WHS file corruption and RAID arrays

LionelW

n00b
Joined
Dec 31, 2009
Messages
9
Hi everyone,

I have been a WHS user for a few years and have been largely using the WGS forums but I have only recently stumbled across this one.

I have two questions:

  • File corruption on multiple drives - is this resolved with PP3?
  • How to get past the 2TB limit on single drives?

The second question is really in relation to the first. I really hope you can help as I have seen some large scale systems in here and you must have come across the excruciating flaw in WHS that causes file corruption on multiple drives. I have built systems with excess of 10 x 1TB drives only to find the same problem that has plagued me for the last few years; corruption. It tells me there is a problem with "this" file and that it is unrecoverable and the only solution is to nuke the file with prejudice!

Bah! Just as well I have a full copy of the data on other systems but why isn't it as reliable as the standard file system?

This fault makes WHS nigh on unsueable to me but there is so much going for it, I am tempted to build a normal Win2k8 R2 next to it and deploy DFS and see if WHS can cope with that as useable space - probably not knowing the 2TB limit (trying my best not to swear and cause GBH to WHS at this time). BTW the file corruption occurred under WHS PP2 - I thought it had been solved by this release - clearly not. I have high transactional throughput of the data, will that contribute to WHS's issues?

I run enterprise RAID arrays at work and figured I'd build a 12TB array using modern hardware, only to find that WHS is capped at 2TB per logical volume. I have seen other systems here with much larger arrays such as the one built by sonofander:

http://www.hardforum.com/showpost.php?p=1034216365&postcount=222

He clearly shows arrays of around 9TB in size, how did he get past the 2TB volume limit? I tried converting to GPT but WHS ignores me and repartitions it down to 2TB and formats it as MBR. The result is an unexpandable partition which wastes the remainder of the disk space.

I have only chosen this option to get past the horrible file corruption problems. The RAID array also delivers higher performance than that achievable using the standard SATA array alone.

Hardware:
  • Asus P6T6 Workstation
  • Intel i7 920
  • 12GB RAM
  • 12x 1TB Samsung F1 HDDs
  • nVidia G210
  • Areca 1680ix-24

Notes:
  • I know WHS only supports 4GB but it was cheap and here’s hoping a new WHS will come out for 64bit.
  • The Samsung drives are replacements for the first batch which just seemed to fail all the time - none of these appear to have had the same problems
  • I read that there have been issues with the Areca 1680i's is this still the case with the latest generation firmware (read this thread)? I ask because I have two - the second is in my main PC driving 6x OCZ Agility EX SLC drives.

Apologies if this is the wrong place to post but thank you in advance for any help.

Many thanks,


Lionel
 
Last edited:
I also experienced file corruption for a long time with WHS (including PP3 beta) before finally giving up. The file corruption for me was alway file system issues so having the base store as RAID or non-RAID didn't make any difference. I was able to reproduce the corruption - it was tied to very large number of files and timing for the DE agent.

I moved to 2K8R2 and am VERY happy. The new fully automatic Server Backup functionality is fantastic and having "Previous Versions" available is a god-send. I installed WHS in a Hyper-V container and only use it for workstation backups. Once I did this the amount of time I waste on "system admin" functions have virtually disappeared.
 
i had a corrupt file system
the backup service would crash every time i tried to back-up a particular computer
when i tried to repair the system WHS would crash requiring a re-boot of win2003

i ran chkdsk in repair mode on the affected disk (chkdsk took ~10hrs to complete on a 2tb drive)
i then ran repair

all was then well :)
 
i don't think i have had a corrupt file... but then again I am not sure I would know until i went to use that file... right?
 
I also have a large whs and i have not had any corruption so i cannot speak to that.

OP you said your data is highly transactional, and i think that may be part of your problem just because of the way that DE moves things around.
Since you have a nice Areca, i think you can create "Logical Volumes" on the RAID controller to pass to the OS.
So you could make two 4+2P RAID6 arrays and create 4 2TB volumes to pass to the OS for WHS.


Ravton, you said you were able to reproduce the corruption? What did you do to create it, what was causing it?
 
I am giving up on trusting WHS now TBH, it has corrupted on me 4 times in total now and I cannot see how MS can release a product this bad. It occured on a HP MediaSmart server, a Dell PC converted to a WHS build, and two custom builds now.

Problem is I like all the add-ons especially MyMovies which I stream all my movies too and centrally store music and video in such a way that My Win7 Media Centre plays them seamlessly.

I guess I'll just have to put critical data on a Win 2008 R2 server and use RichCopy/SyncToy to replicate folders to the WHS toy server.

BTW why RAID 6? I can see no benefit over a RAID 5 array with multiple drives and a HS or two? RAID 6 is known to be slower, unless there is something new that is not common knowledge yet?
 
Last edited:
I also experienced file corruption for a long time with WHS (including PP3 beta) before finally giving up. The file corruption for me was alway file system issues so having the base store as RAID or non-RAID didn't make any difference. I was able to reproduce the corruption - it was tied to very large number of files and timing for the DE agent.

I moved to 2K8R2 and am VERY happy. The new fully automatic Server Backup functionality is fantastic and having "Previous Versions" available is a god-send. I installed WHS in a Hyper-V container and only use it for workstation backups. Once I did this the amount of time I waste on "system admin" functions have virtually disappeared.

+1 on loving Win2008 Server R2. The new Windows 7 functionality in the interface is cool, and besides that I'm also one of those people that's revisited WHS constantly ever since the early days, trying to be open minded because SO many people out there swear "well it works for me." But each time I revisit WHS and run into the same old stumbling blocks, I realize once again its just a toy. I should qualify that maybe I'm not the average WHS user since I have over 100Tb worth of data spread across 5 servers.

Microsoft needs to do two things with WHS v2 if they don't want to hold the product back from broader adoption: besides bringing it out of the Win2003 kernel stone age and onto the newer server kernel, Drive Extender needs to gain RAID4 type parity ability. Also, it needs to gain the ability back up the system (boot) drive and maybe let people create a "restore disc". The fact that if your boot drive fails and you have to sit reinstalling from scratch is totally unacceptable, and I don't care about "well tombstones is why you can't image the boot drive" - then they just need to rethink it.

Until that's done it's a toy and I'm staying far away, and I haven't even touched on these scary corruption issues people are having. :headdesk: I'd love to be able to get away from expensive raid arrays and be able to mix different sized drives that can spin up and down individually on demand, but nothing else has proven as rock solid for me as arrays in the last 2 years.
 
BTW why RAID 6? I can see no benefit over a RAID 5 array with multiple drives and a HS or two? RAID 6 is known to be slower, unless there is something new that is not common knowledge yet?

To me RAID6 is the way to go. The price of the extra drive is cheap insurance for an extra parity stripe. Raid5 can suffer from "bit rot" (look it up) which means if one drive fails, and you haven't set your array to auto-scrub (meaning scheduled checking for errors and healing the data from parity for any bad sectors it encounters) then your parity is potentially compromised when trying to rebuild after a drive failure.

Additionally, RAID6 only trails RAID5 in write speed, and on newer raid controllers they're practically even due to newer dualcore IOP chips. On both my Adaptec 5 series and Areca 1680 controllers, Raid5 vs Raid6 write speed is pretty close in benchmarks.

The only RAID5 array I run is to hold duplicate data of my RAID6 arrays, and I've got it set to auto-scrub every 2 weeks. The last auto-scrub performed encountered 23 errors, so I'd hate to have lost data in a worst case scenario where I had to rely on that data and a drive failure occurred prior.
 
Last edited:
I am sorry but I have to disagree odditory, I am however happy to change my opinion if you have better info on this.

RAID 10 is the way to go.

RAID 6 is still slower than 10. Hardware is cheap data is not.I feel that bit rot does not affect the data in any way close to what you are suggesting.

We handle data across 100's of TB and over the last 15 years of using RAID 5 across 1000's of spindles and I have never experienced data loss due to fault in the array, only through other symptoms.

Where RAID 6 kicks in is in the scenario of a double disk failure. I'm not saying it wont happen but this is a very rare scenario. It is more likely with modern SATA drives with large capacities than with enterprise HDDs admittedly. To put it in perspective, in my 25 years in IT I have not had bitrot occur to a point that causes the loss of data. At least not that I know of.

If a bit flips, the drive contains enough redundant data that it can and will be corrected the next time that sector is read. You can see this if you check the SMART stats on the drive, as the 'Correctable error rate'.

Depending on the details of the drive, it should even be able to recover from more than one flipped bit in a sector. There will be a limit to the number of flipped bits that can be silently corrected, and probably another limit to the number of flipped bits that can be detected as an error (even if there is no longer enough reliable data to correct it)

This all adds up to the fact that hard drives can automatically correct most errors as they happen, and can reliably detect most of the rest. You would have to be have a large number of bit errors in a single sector, that all occurred before that sector was read again, and the errors would have to be such that the internal error detection codes see it as valid data again, before you would ever have a silent failure. It's not impossible, and I'm sure that companies operating very large data centres do see it happen (or rather, it occurs and they don't see it happen), but it's certainly not as big a problem as you might think.

RAID 5 is already pretty good compared to single drives, RAID 6 is better but compromises both space and performance and RAID 10 achieves a far higher quality standard and good performance at the cost of space. Personally I run RAID 10 for critical or performance arrays, RAID 5 for space with a second RAID 5 backing up the first RAID 5 which is offset in time allowing D2D2T.

The real question is "how likely are we to really see the effects of bit rot on modern hardware?" Although this would derail this thread, is there a thread already running that this could be picked up on?

I do run RAID 0 arrays but these are for pure throughput and these are backed by RAID 10 units.
 
I don't disagree Lionel, and you make good points. Its unlikely that we'll see the effects of bit rot. However I'm talking within the scope of media storage and archiving, not enterprirse. In other words in this day of $300 20-bay server cases, and 2Tb drives to be had for $125, it seems silly *not* to go RAID6 for the extra insurance against a worst-case-scenario, like summer when it gets hot and chances for multi-drive failure are multiplied. People are storing their servers in the garage or in the house where its anything but a climate controlled server room. As well, some of us have spent years painstakingly transferring our optical media to harddisk and so $125 or $150 for an extra parity stripe is no-brainer insurance against having to repeat transferring from our optical media. People often forget to place a value on their own time but it's more valuable than anything.

not to mention a lot of people are going to build a media storage server and NOT bother to duplicate the files somewhere else, I know because for a long time *I* didn't until drives got cheaper, since I figured "well my optical media is my backup" and so it's for those people that I know won't bother (or be able to afford) file duplication as part of their storage strategy that I recommend as much redundancy as possible. Lastly, early last year I had a situation where had I been running R5 instead of R6, I'd have lost my array- granted it was due to a confirmed bug in Adaptec Storage Manager and not any drive's fault, but still. 20 minutes into a raid level migration/expansion, 2 drives got kicked out and array status changed to "0 tb". Card wouldn't rebuild array or do anything after new drives inserted, Adaptec said I'm S.O.L. and have to delete and recreate array. Luckily NTFS partition was still intact so I could transfer data off temporarily, albeit at about 60MB/s so let's just say 22Tb took A WHILE.
 
Last edited:
Back
Top