• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Reliable controller for 8 SATA ports?

Cerulean

[H]F Junkie
2FA
Joined
Jul 27, 2006
Messages
9,484
I don't know exactly what it is yet, but for the past 1-2 weeks I've been trying to diagnose this problem. At first, I thought my brand new Western Digital harddrive was dying. I did a chkdsk /f /r on it which lasted about 5 days. It found a few bad sectors, made some repairs, did what it needed to to make things operational. Before I did the chkdsk, it was capping at at no more than 1 MB/s on read and write transfer rates. After the chkdsk + a reboot, I decided the first thing I will try to copy over are my new unpublished photos. It started out at 10 MB/s and about 9 hours remaining. Within a couple hours it went from 9 hours remaining to around 10 hours. The next morning, about 8-10 hours later, it was about 50-55% complete with the process, but the remaining time went to 10.5-11 hours. 8 hours later, it is at 12 hours. According to Performance Monitor, it has slowed in write speeds to the OTHER harddisk that I am writing to. It reads at like 5-8 MB/s from my new WD drive.

Cliffs:
* New drive has read and write speeds no greater than 1 MB/s
* After chkdsk, has read and write speeds of 20-30 MB/s; normal speeds are supposed to be around 90-130 MB/s
* I begin copying a folder with around 250 GB of unpublished data to another drive that is already mounted beginning at a rate of 10 MB/s and approximately 9 hours remaining
* Over time, I notice the transfer rate decrease until 16-18 hours later it is writing at speeds of around 2-3 MB/s with 12 hours remaining
* I watched throughout the day the time remaining act like it was going down but each time I checked it it was a little higher than the last time I checked
* I decide to plug in two of my other drives (one a 3TB and the other a 1TB)
* Computer hangs up during boot time with "B4" in the bottom right corner of the screen
* I remove power from all HDDs except SSD, and boot is successful this time

Note:
* Going back about 3-5 months ago, I encountered the same or similar boot issue, and eventually concluded that my 750W power supply must be showing its age and very slow death
* I got everything back up by adding a second daisy chained 500W brand new power supply, which only reinforced my thinking about my 750W power supply (which I bought in the mid 00s)

Continued cliffs:
* I pull out all my HDDs (except SSD) from my machine, hookup my StarTech docking station to a computer at work, and do a test copy from the HDD I was trying to copy off of in the first place: 20-21 MB/s solid, right at USB2's typical limits (this was a HUGE relief, I don't think my HDD is dying)

I am beginning to lose my mind.

I have not tried replacing SATA cables, or hooking HDDs up to different SATA ports (such as the one my SSD is hooked up to, which does not show any signs of issues).

I do not know if it is my motherboard (chipset/controller) going bad.

I do not know if I need to get a RAID card (to use merely as a controller) to solve this problem. If I do, I'm looking for suggestions for a reliable card that can support 8 SATA ports.
 
OK, apologies for OP. I've been very stressed out by the situation. I'm hoping that getting a RAID card will solve the issue, but to do this would require me to resolve my PSU issue, which would also require me to replace my chassis with one that can take a physically longer PSU.


Total: $1153.28 shipped

Then I would need to begin replacing my HDDs with five identical 2TB or 3TB drives (shoot for reliability), do a RAID6, and get either a APC Smart-UPS SMT1500RM2U 1000W/1440VA 2U Rackmount UPS ($599.99 shipped) or a APC Smart-UPS SUA2200RM2U 1980W/2200VA 2U Rackmount UPS ($810.25 shipped). Initially I was thinking about getting a Synology DiskStation DS1512+ 5-Bay NAS ($799.99 shipped) with like a APC Back UPS Pro 1500 865W/1500VA ($189.60 shipped), but with the tier of work that I've gotten into having a good hardware RAID card could potentially go a long ways and provide the availability that I need.

I understand that the bigger the HDD capacity that there is more to lose / the disaster is bigger, and that errors are more likely to pop-up and with as many multi-TB drives that I have the likelihood of incidents are significantly higher, etc, --> need to put together a RAID and establish a backup scheme. I don't want to rely on my motherboard's controller should it die or go back and I lose my stuff. :(

EDIT: I haven't pulled any trigger on this yet. I am contemplating if I can drag this out somehow for one week.
 
As an Amazon Associate, HardForum may earn from qualifying purchases.
Your issue might be just the PSU dying in your current setup. Try replacing the PSU first and see if the problem goes away.
 
Placed an order for the chassis and PSU and will see about that. If that is all it is, I might be able to hold off on the RAID controller until I can get the RAID controller + UPS + 5x 2TB (or 3TB) NAS-grade HDDs in one shot.
 
I don't know exactly what it is yet, but for the past 1-2 weeks I've been trying to diagnose this problem. At first, I thought my brand new Western Digital harddrive was dying. I did a chkdsk /f /r on it which lasted about 5 days. It found a few bad sectors, made some repairs, did what it needed to to make things operational.

chkdsk is a filesystem check. Not a disk check (although it can perform a surface scan, you're better leaving such things to the drive to do; yeah, I know what the name is... it's just not particularly appropriate In My Opinion).

Use a tool for reading the SMART statistics and check whether the disc itself is failing. For WD discs, any sign of failure from SMART is useful as an indication (and invariably is grounds enough for RMA). Any entries in the log for commands which failed indicate that it is definitely failing. Counters which are particularly important on the SMART statistics for WD are:

* Offline uncorrectable - when the disc is idle it spins over the contents trying to read them. If it finds any it cannot read, they are flagged here. This way you can find out that there are problems even when you're not actively accessing the disc.
* Reallocated sector count - when it fails to read a sector (possbily as part of the offline uncorrectable checks), the disc may reallocate it internally elsewhere in the disc. Any value > 0 here indicates that the disc itself (or the head) is failing.
* Seek error rate - usually means that the head is failing, or the drive that moves it is failing, although it could be the disc itself if the index are becoming corrupt. Any value > 0 is a very likely failure.
* Current Pending Sector - usually mirrors 'offline uncorrectable', and means 'this sector is the one that's broken and will be reallocated.
* Spin/Calibration retry - may indicate a bad power supply if it's >0. Check the leads to the drive
* Spin up time - if the raw value of this is high, it's likely a problem. That said, the green power drives often have values around 5000 here, so take it with a pinch of salt.
* Power-on hours - how long the disc has been running for. 1TB WD Green discs will run for about 3-4 years. 2TB WD Green discs will run for about 1-3 years (my experience is that they suck) - I've even had them fail within months. Because of my experiences with the 2TBs, I've not bought anything larger so have no experience to impart there.
* Load-cycle count - expect this to be high. The raw value may be 10x that of the power-on hours count, or higher. No real assumptions on the failure rate - some people talk about using the WD tools to reduce the unload rate on their drives.
* UDMA_CRC_Error_Count - this is the most important other than the offline-uncorrectable figure. This tells you 'the on board controller detected problems in the messages from the disc controller' - in other words, the signal isn't getting through. If you see this increasing (or >0 if it's the first time you've checked) then it's almost a certainty that the cable between the controller and the drive is loose, faulty or experiencing interference). When such problems occur they may be detected by the operating system and cause it to switch to a slower mode in the hope that this will help - usually switching down the DMA modes until it finally switches to PIO mode. The visible effect of this is that disc accesses turn to treacle.

If you are seeing slow access and the UDMA CRC Error count is >0, check your cables. Reseat them at first. Re-route them if you think 'that's going a bit close the foo in the computer'. This may stop the number of faults increasing. The OS may, however, remember that you've had problems and be stuck using a poor transfer mode - you'll have to find out how to increase the mode from PIO to UDMA again. My limited experience with this problem tells me that this happens more with CD/DVD drives than HDs; on Windows versions I've used, the system has restored itself to UDMA without intervention. On Linux (at least ones I've seen) the modes are dealt with on every system start, so there's nothing to restore.

If you see reallocated sectors, or offline uncorrectable, run a disc check - the disc, not the filing system. All modern discs have the ability to run diagnostics themselves, just by having a command issued to them. The SMART tools can trigger this. WD 'short' test takes 2 minutes, and will report errors. You schedule the test and then leave it for 2 minutes. When you return, you query the SMART logs again and you'll see that the test passed. Or that it failed - a failure is an instant RMAable state. If you see that the drive itself reports that it is failing that should be enough, every time, to get the disc RMA'd without any further questions asked. The fact that you did the test and the results is recorded on the disc itself, so you don't even need to make a record of it - when returned to the manufacturer, the evidence that the test was run and its results is right there for them.

If nothing appears to be reporting failures, you need to look at other causes - preferably ones that don't cost you money, and take moments to do. Reseat any PCI/PCIe/PCI-X cards that you have in the system - spurious interrupts from cards, even those that aren't the ones handling the disc, can be a problem. This is less of an issue these days, but at least checking the card helps.

Isolate the disc that's having problems - connect it, alone, to the motherboard and boot a rescue system that can run the SMART diagnostics (preferably from CD/DVD/USB), so that you can see what's going on without anything interfering. This is particularly useful when the disc that's a problem is the boot/system disc, as booting from a different device will allow you to do tests on the failing disc without being affected by the OS loaded from it.

Exchange SATA cables (or IDE cables if you're using a system that's that old). Just swapping the cables between drives is a very simple and easy way of eliminating the cables as a cause of problems - or implicating them. If the problem goes away when you swap the cables (and doesn't appear on another device) then it was most likely loose or electrical noise interfering. If the problem moves to the device you swapped with, then the problem is the cable. SATA cables are cheap - bin the dodgy one.

If the problem remains on the device whatever the cable, consider whether the socket on the motherboard/controller card is faulty (or the controller itself), and move that end of the cable elsewhere. If things improve, then maybe that socket or controller's unhappy. Stick masking tape over the socket. Consider replacement card, or thinking about another MB, but if you moved it somewhere that worked and you're not short of sockets... it doesn't really matter that that port's not working.

Consider the size of your PSU. A 750W PSU can happily drive a dual-core 2.6GHz processor with a basic graphics card and 8 WDC green discs, even if they all spin up at once - and probably much more. If the PSU is failing you'll see other problems as well, typically including CPU clock rate reductions. Graphics card corruption may be visible, and it's possible that fans may be fitful. With a cheap PSU, a bad line input may cause power fluctuations which affect the system badly - you can get little inline plugs that sit between your plug and the power socket that measure the line rate (power draw, voltage, frequency), although these can also vary in quality. If nothing else, they're interesting because you can see just how many amps your kit is drawing. A UPS between the power socket and the machine will often save you from those problems - which happen in some areas, although getting your supplier to admit that there is a problem is sometimes tricky, even when you're measuring voltages well out of spec. If the PSU is failing you'll probably be able to see fluctuations in the measurements that are shown on the BIOS screen (many BIOSs have the ability to show you the measured voltage on a number of inputs, which can be useful in locating issues).

If you're concerned that the PSU is failing, reduce the load on the system and try it in the limited state - with less load, the PSU should be able to supply power (assuming it's not failing because the line input is noisy or out of spec). Removing drives, graphics cards, and other devices from the system will reduce the load, and should let you get to the point at which things work. If on a bare system you still cannot get things to work, maybe it's not the PSU (or it's so utterly broken that it just doesn't work). Depending on your setup, swapping out the PSU might be easy - in which case you can confirm or refute that the PSU is the problem by doing so. Easier, though, is to move the disc to another machine and boot it there. If it all works then you've shown that the drive is fine, so it may be the motherboard, controller or PSU (or cables, if you didn't already discount them - move them with the drive if you like).

To address specific things which you've said:

* Yes, you should try switching around the cables and the drives. This should be an early thing to try.
* Thinking you need a RAID card is probably a bad assumption, based on the evidence that you've given. It's useful to have RAID and it might help you retain the data you've got but since we don't know whether the ports or cables are working as you've said that you haven't tried that.

You said that you got slow disc speeds after use. Did you try checking the error logs to see if there were any issues being reported. On Windows I have no idea where these would be, but on linux you would see reports in /var/log/kern.log (or similar depending on distribution/configuration) which told you that disc reads were failing, or that processes were hung waiting on IO.

Don't spend any money on guesses about the components that are failing (unless you don't have spares lying around) - actually test whether they are a problem. Otherwise you're just going to flail around replacing stuff randomly. Eventually you'll probably hit something that works, or something that works because of something incidental that you changed (eg routing a cable a different way when you fit a PSU), without understanding what the problem actually is. Be systematic.

Although your comments about connecting the additional drives imply that the system may be more heavily loaded when it starts, there could be other reasons for there being a problem, so eliminating them by doing the cable swaps and reseats will help. I've seen occasions in the past where, on connecting devices the system fails, and then on removing them it works fine - implying that the component that was connected caused the problem. In fact the problem was that I had knocked a PCI card slightly out of its socket, and this completely prevented the system startup... and I knocked it back again as I put removed the device. Stupid, but repeating the test (especially the actions in the same way!) showed me the mistake.

A quick note about your copying the data off the disc reducing the speed of writes to another disc... unless you're doing something special, that's to be expected. If you're pulling data off one drive at 5MB/s and writing that data to a second drive, the maximum speed you will be able to achieve is 5MB/s. You cannot write the data out faster than you can get it in. Well, you can, in bursts, but the average write speed will match the average read speed.

Anyhow, I hope some of that is useful... if not then... well, hopefully I've not wasted too much of your time.
 
Thanks for the advice gerph; I found a good amount of the information you suggested as useful and beneficial, but other information as "not exactly true" based on my knowledge and experience. I have been very tired lately, so that has made me incredibly lazy and careless about utilizing my skills. I greatly appreciate your input, as you've done some of the thinking for me that I would have had to do otherwise (I've been mentally burned out from work). Ex. of information "not exactly true": the 5 MB/s disk-to-disk maximum speed is a very far from the truth (sorry, but it is :( ), unless I am gravely misunderstanding you somehow and totally misreading what you are saying -- please clear this up for me if you can. I would be more inclined to accept that if we were talking about very old PATA HDDs and Socket A (or older) era controllers. Even USB2 and cheap flash drives are faster than that.

But, I did I spend several days running the WG Diagnostic Tool from my UBCD 5.1.1 disk, doing one powered HDD at a time. Two out of five of my HDDs were reported to have too many errors to complete a Full/Extended Test. Of the remaining three HDDs, two reported errors but were repaired successfully. I still have to go back and lookup the error code for each so that I know exactly what the issue is.

I did receive the following items:

I didn't think the chassis would be as huge as it is. D: Still blows my mind to this day; looks like a miniature supercomputer. Need to take some pictures of it sometime, which I will when I have the energy available to focus on windtunnel strategies (with two 252cfm 120mm Delta fans available in my arsenal, HAH HAH HAAA...).
 
Last edited:
As an Amazon Associate, HardForum may earn from qualifying purchases.
(sorry for the long delay; I immediately forgot about this post and didn't come back - until it came up in a search when I was looking for something else)

My example of the transfer rate might not have been explained very well.
Essentially what I was saying was that if you are copying data from one disc to another, and reading data into memory at 5MB per second, that same data can only go out at 5MB per second (or less) - it can't write out more than 5MB per second, because it hasn't got more than that. Now there might be buffering involved so that the OS reads in 5MB in a second, and *doesn't* write it out to the second disc. Instead it waits for a bit until it has a lot more data to write out, and then does it all at once. If it waited 5 seconds before flushing the buffered data then it might get a burst speed of 25MB per second, but that would be after 5 seconds of doing nothing - so still 5MB per second.

If you meant that a completely different device only got 5MB/second writing, whilst you were reading your disc, then ... something is indeed up. It could be that the controller is maxed out - but as you quite rightly say, it'd be pretty astonishing to get a maxed out system on anything but PIO PATA devices at 5MB/second!

Sorry if that wasn't clear the first time around :-(
 
Back
Top