• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

MoDER RAID ERC Matrix for Error Recover

WynnSmith

n00b
Joined
Jul 17, 2012
Messages
2
I'm tempted to start this article, like so many others you've seen, by reporting how much time I've invested into researching RAID error recovery, and therefore you should trust me to have the ultimate answer. However, in truth, I have no answers; just an idea, and more questions.

Let me begin with what I have concluded. The solution we need does not yet exist because ERC is clouded by such a complex matrix of perspectives that even the OEMs haven't yet thought about it correctly. The good news is, the solution could be implemented through firmware fixes and made available quickly.

Before I explain the ERC matrix, (or atleast the part I see), and explain why the OEMs are confused, let me quickly cover some terminology. I'm talking about the need for a RAID controller and disk drive to cooperate over Error Recovery Control (ERC ), Time-Limited Error Control (TLER), or Command Completion Time Limit (CCTL). More importantly, I'm talking about a need to change their level of cooperation based on a change in the mode of operation.

Did you notice the word "change"? Later, I'll explain this new idea of mode change, which I call MoDER, and the matrix of issues that reveal its need. MoDER stands for Mode Dependent Error Recovery. If it turns out this idea is new, then dear OEM, we now have 12 months to get our patent ducks in a row.

You've no-doubt already heard both sides of the ERC debate. In my opinion, both sides miss the picture. Now, while I advocate "Enterprise" class drives for RAID, (Intel publishes an excellent whitepaper called "Enterprise-class versus desktop-class hard drives comparison whitepaper"), I believe the RAID controller OEMs and disk drive OEMs are shooting-themselves in the foot by not providing the best operation possible with whatever they find themselves plugged in to. As a result, the market is not as big as it could be.

So let's quickly review the disk part of the ERC matrix. With consumer grade disk drives, also called desktop grade or stand-alone grade, the disk drive itself is smart enough to attempt to recover bad data read from a disk platter. While attempting to recover, the drive may go offline for two minutes or more as it tries over and over to read a good version of the data. Sadly, many users, lacking an appreciation for this opportunity to refill their coffee cup, mistakenly think their computer is locked up, and turn off their computer before recovery succeeds; but that's a side issue.

The good news here, lacking an error-tolerant data-redundancy scheme, is that disks are smart enough to attempt a recovery from the only place a recovery can be made. In other words, with stand-alone disks, longer time-limits provide better reliability. This is the reason disk OEMs of consumer disk drives, purposely default to long ERC time-limits. The gotcha for MoDER, as I'll explain in a minute, is that some OEMs purposely disable the ability of drivers and utilities to set a lower ERC time-limit.

With enterprise-class disks, the goal is to optimize "system" performance. Therefore, OEMs assume the disk will become part of an array of data-redundant disks. When a disk detects an error, it will only attempt recovery for a few seconds, and if the recovery fails, report the failure to a controller so that data can be recovered from different disks.

The good news here, lacking an actual drive failure, is that the system of disks remain responsive and relatively efficient and the potential for data loss remains low.

Now let's quickly review the controller part of the ERC matrix. Desktop disk controllers are relatively dumb. They receive a disk request, issue the request to the disk, and wait for a response. In the event of a request failure, assuming the user enjoys coffee and avoids the 'Off' button, the user will eventually see an error message.

An Enterprise disk controller, or RAID controller, is smarter and has multiple goals. It strives to maintain performance (bandwidth), responsiveness (latency), and reliability of the disk system and of the entire server system. RAID controllers generally manage an array of disks, ensuring that data is stored redundantly. In the event of a request failure, the RAID controller must perform multiple tasks. It must determine where the data can be recovered from elsewhere in the system, recover the data from different disks, remove the bad sector from future use, and then rebuild a redundant copy of the data on a good sector.

Now suppose it's not just a bad sector of data. Suppose enough thresholds have been crossed that the RAID controller determines an entire disk drive has failed. A drive failure means the system enters a mode of degraded performance as it seeks alternate data sources, and then uses an algorithm to produce good versions of the data. This is independent of additional issues such as hot-spare, hot-swap, logs, user notifications, and such. The good and bad news is that no disk is so important, the RAID controller isn't willing to sacrifice it when enough thresholds have been crossed.

In the ERC debate, the key threshold being discussed is the issue of ERC time-limits. So far our matrix is relatively simple, so let's examine it in relation to ERC time-limits.

When a desktop controller is paired with desktop disks, the controller is dumb and any hope of recovering data from a bad sector is left to the disk drive. Therefore longer ERC time-limits are good for consumer grade disks and controllers.

When a desktop controller is paired with Enterprise disks, a short ERC time-limit reduces the potential of recovering data from a bad sector. With many of these disk drives, it's possible to use tools to turn up the time-limit. Nevertheless, most users lack such knowledge and therefore the less expensive, non-enterprise disk drive is both financially and technically a better choice for desktop systems; particularly for coffee drinkers.

When a RAID controller is paired with Enterprise disks, a short ERC time-limit allows the controller to respond more quickly to disk errors. The controller quickly grasps control and calls on other disks to deliver reliable data. But notice, with all the talk of RAID reliability, the bias is toward performance of the system rather than reliability of the disk.

The fourth quadrant in our simple matrix is when we pair RAID controllers with consumer disks. This is where debates over ERC time-limits flair up. As you read through various passionate on-line commentaries, you can just imagine red angry faces and the racket of keyboards as fingers dance madly across key tops. In some cases, utilities can be used to turn down the ERC time-limit on consumer-grade disk drives. In other cases, OEMs made extra effort to disable such an ability. In the best cases, the system builder must arrange for a method to reset the short time-limit after each power cycle.

Let's ask why. Why the passion? Why the extra effort by OEMs to take away a feature they made efforts to give us? For the OEM it's about Enterprise sales. For the rest of us, it's all about the threshold that might be crossed in which a RAID controller determines it's not just a bad sector, but a bad drive. In a RAID system, with consumer disks, a disk initiated error recovery may exceed the time limits of the controller. One too many slow responses could be enough to knock the entire disk offline and permanently out of commission.

From here, the matrix grows more complex, and so I won't revisit each intersection of each axis.

Next we'll consider contrasts. In a desktop system with a consumer grade disk controller, no matter what grade drives are used, there is never a time when the system gives up on an entire disk. Yes, the drive could fail completely, fail to boot, and never again work. The point with consumer grade controllers is, if the car didn't start, it's because "the car wouldn't start" as opposed to, "we decided not to start the car".

In contrast, RAID controllers follow every operation by a line of code that asks, "can I fail you now?" RAID always checks if the disk crossed enough threshholds to justify removing it from operation. The point with RAID controllers is, if the car didn't start, it's because the controller chose not to start the car, even if the car can be started.

It's this contrast that causes people to believe they can seek a "correct" ERC time-limit. They know an otherwise good disk is too often taken offline by a RAID system where a disk's attempt to self-recover exceeded the controller's time-limit. They know a consumer-grade controller will fail to receive potentially good data when the ERC time-limit of an Enterprise disk is too short.

Have you noticed this debate focuses on the "normal" mode of operation? I think we need another axis on our matrix.

Let's begin the next axis on our matrix by comparing similarities. In a desktop system with a consumer grade disk controller, there is only one mode of operation. We might call it "do your best Mr. disk drive." The data is not redundant. Any hope of receiving good data depends on a disk drive and its built-in error-recovery mechanism. The disk is truly stand-alone.

In normal operation on a RAID based system, data is redundant. If a disk slows down to perform built-in data recovery, the RAID controller can call on other disks to get the data. However, if the disk hesitates too long and the controller gives up and takes the disk offline, the disk system enters a degraded mode of operation. At this point the data is no longer redundant. Any hope of receiving good data depends on a disk drive and its built-in mechanism to deliver reliable data, even if that data is a checksum used for calculating resultant data.

When a RAID controller enters a degraded error-recovery mode of operation, reliability of a disk drive becomes just as critical as with the desktop consumer-grade controller. Unfortunately, with current RAID controller designs that I'm aware of, the bias stays toward system performance rather than disk reliability. The same ERC time-limits that were used under the normal mode of operation, continue to be used under the error-recovery mode of operation.

This is disasterous. A RAID system in the degraded error-recovery mode of operation can be much less reliable than a consumer-grade disk system because RAID ERC time-limits are low, server disk usage is high, and an error on any one of multiple disks is possible and becomes fatal because redundant data is no longer available.

Making things worse, disk sizes have vastly increased, while disk error rates and interface speeds have not vastly improved. It doesn't matter if hot-spares are available. The system faces many hours or days of degraded, at-risk performance while the entire disk is rebuilt. During the rebuild, the RAID system is pushed to its highest level of utilization reading and writing huge volumes of data. There is a high likely-hood of another disk error before the RAID system can return to normal operation. During this time, the data is truly at risk.

What is MoDER and how would it change things? MoDER stands for Mode Dependent Error Recovery. It endorses two basic ideas. The first idea is for the disk controller and disk drive to cooperate in selecting ERC time-limits during operation, and to change those time-limits to match circumstances. The second idea is for the controller to interpret the circumstances and change its modes of operation, biased to performance when conditions are ideal, but allowing the bias to change toward disk reliability, leveraging a disk's built-in error recovery mechanisms, when circumstances dictate.

The controller would effect ERC time-limits through at least three different modes of operation. The first mode is "normal" operation in which ERC time-limits are low. As disks report errors, the controller performs its normal function of recovering data from different disks.

When various types of disk errors reach an appropriate threshhold, the controller will enter the second mode of operation, which might be called "guarded". Two important tasks take place in the guarded state. First, the controller should set a higher ERC time-limit, and become tolerant of the time-limit, allowing a disk to recover data when necessary. Second, the controller should invoke a hot-spare or notify the user to install a new disk.

An important idea here is to install the new disk without removing the old. Barring catastrophic failure, it is likely the old drive can aid data redundancy, enhance error recovery, and improve fault tolerance during the rebuild process.

The third mode of operation is the "rebuild" mode of operation. During the rebuild, the controller should remain tolerant of higher ERC time-limits. When the failing drive is unable to deliver data, and the data is not redundant, higher ERC time-limits will allow data to be recovered from the drives in the most reliable manner possible. In other cases, a higher ERC time-limit allows the failing drive to enhance data reliability when necessary.

Only after a successful rebuild on the replacement drive, or in the case of a catastrophic failure, will the original failing drive be taken offline.

By leveraging a disk's built-in error recovery mechanisms, and allowing a controller to change its bias from performance to disk reliability, RAID systems can be made more beneficial and less vulnerable. Disk OEMs who disable the ability to set ERC time-limits will discover a shrinking demand, while other OEMs who support MoDER features, will help build a larger market.

The market for Enterprise disks as stand-alone disks could be opened too if MoDER features were implemented in consumer-grade controllers. By allowing the controller to set an appropriate ERC time-limit, the increased performance and reliability of Enterprise disks could be leveraged for stand-alone use in Home Servers, Engineering workstations, and Graphic systems. If you read Intel's whitepaper mentioned above, you'll find yourself wanting Enterprise features in a stand-alone system.

Now having thought through these issues, I'm wondering about existing controller features. Most forum threads I've seen, discuss disk features. Do RAID controllers exist that tolerate high ERC time-limits? Are these limits settable? Are controllers capable of setting a low ERC time-limit on disks which allow it?

Wynn Smith
nwhelp.com
 
Here is why MoDER substantially increases the market size for hard drives.

As customers become aware of the contrast between computers which support "expandable" disk systems and computers that do not, all computer makers will be pushed toward a RAID implementation because the included disk can be considered as simply the first disk of an expandable set.

By implementing MoDER, the market for Enterprise disks, used as stand-alone disks, could be opened. High-end buyers and suppliers are more likely to choose an Enterprise disk for stand-alone use because hope is maintained that additional disks could be added to the system later, with significant benefit, and without the cost penalty of replacing the original drive. Such markets include Small Business Servers, Home Servers, Engineering workstations, and Graphic systems.

In consumer computers with "expandable" disk systems, adding a drive becomes easier to accomplish, and the benefit of doing so becomes easier to recognize. The pain of migrating data and hiring experts is eliminated.
 
Back
Top