• Some users have recently had their accounts hijacked. It seems that the now defunct EVGA forums might have compromised your password there and seems many are using the same PW here. We would suggest you UPDATE YOUR PASSWORD and TURN ON 2FA for your account here to further secure it. None of the compromised accounts had 2FA turned on.
    Once you have enabled 2FA, your account will be updated soon to show a badge, letting other members know that you use 2FA to protect your account. This should be beneficial for everyone that uses FSFT.

Testing heavy IO affects

Red Squirrel

[H]F Junkie
Joined
Nov 29, 2009
Messages
9,223
I've been strugling trying to figure out WTF is causing me so many issues with my raid. I've replaced the backplanes, no go, I'm waiting for the other backplane to get here so I can start replacing drives (I want to convert to raid 6 first which will require more bays). It could be the controller, cables, particular ports... this micro troubleshooting is driving me nuts.

Is there some kind of test I can do on a single drive that could determine if there is any issue as far as IO goes between that drive and the rest of the system? The issue I have happens randomly, usually overnight, and because it's raid, I don't know which drive is the culpit. I just get tons of memory dumps and other very general IO errors, none pertaining to a particular drive. Though at one point I did have drives dropping left and right out of the array but it was so random and fast I never managed to determine which physical drive so I can trace it to the port.

Smart shoes no errors.

I'm running Linux, so I'm thinking maybe some kind of hdparm command?
 
Break the raid array and try separate overnight writes to each separate drive. That won't test out the raid subsystem, but will test out most of the drive/io/etc. pathways
 
I would start with a dd read test of a few 100GB and see if that causes any problems.
 
I was thinking that too, so just dd it to /dev/null I guess? Guess I'll start with that and see what happens from there.
 
Been doing this for the past few days to see if I can isolate a single drive, but it's always random. Think it's time for a new server, too hard trying to troubleshoot this. I just need this to work so I can move on. The errors are random too.

Originally, it was drives dropping out of the raid. Now it's this:

Code:
INFO: task pdflush:14763 blocked for more than 120 seconds.
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
pdflush       D ffff8801dac8fa80     0 14763      2
 ffff8800048c9da0 0000000000000046 ffff8800048c9d00 ffffffff8101686f
 ffffffff8162a500 ffffffff8162a500 ffff88020f12adc0 ffff8801ff4496e0
 ffff88020f12b108 000000020dc05dc6 ffff880028062570 ffff88020f12b108
Call Trace:
 [<ffffffff8101686f>] ? read_tsc+0xe/0x24
 [<ffffffff81033569>] ? __dequeue_entity+0x61/0x6a
 [<ffffffff8100e80e>] ? __switch_to+0x1b0/0x3e0
 [<ffffffff812c1f41>] __down_read+0xa3/0xbd
 [<ffffffff812c1250>] down_read+0x2a/0x2e
 [<ffffffff810c04ff>] sync_supers+0x4a/0xc4
 [<ffffffff810946e0>] wb_kupdate+0x35/0x119
 [<ffffffff81095163>] pdflush+0x16e/0x231
 [<ffffffff810946ab>] ? wb_kupdate+0x0/0x119
 [<ffffffff81094ff5>] ? pdflush+0x0/0x231
 [<ffffffff81094ff5>] ? pdflush+0x0/0x231
 [<ffffffff810534bf>] kthread+0x49/0x76
 [<ffffffff81011719>] child_rip+0xa/0x11
 [<ffffffff81010a37>] ? restore_args+0x0/0x30
 [<ffffffff81053476>] ? kthread+0x0/0x76
 [<ffffffff8101170f>] ? child_rip+0x0/0x11

At one point I also had just a drive reset (as if it was unplugged and replugged) but it happened fast enough to not cause a raid rebuild. What a pain this is. Now I know why real sans are 60k.. You plug it in, and it works!
 
Dumb question, but what memory do you have? A couple years back I had a system based around a ASUS PC-DL Deluxe with an onboard Promise TX2. I ended up with data corruption on the two drives connected to the TX2 (MD RAID 0) and blamed the controller.

Not long after that the memory died; I'm wondering if the data corruption was actually caused by the memory going bad...
 
You know I was starting to wonder if my issue might be ram related. I don't know off hand what type of ram is in there, I just know it's 4 sticks of DDR2 2GB to make a total of 8GB. I will run a memtest overnight and see how things turn out. Though I suspect a memtest could show everything is ok and still have issues so I may need to do 1 ram at a time and see if these issues continue. The error I posted actually started months ago, but I thought it was a virtualbox issue. I replaced a failing drive, then it went away, and now it's back. But it's so intermittent, it's really hard to tell, could have just been a coincidence that it did not happen for a while after replacing one drive.

I'm also debating on replacing the cpu motherboard and ram. I could go with a Phenom x6 16GB of ram system for about 1k. Same case and all, just a swap out.

Also, could a bad PSU cause issues like this? I tested the voltages and they look good, though but could they somehow fluctuate randomly? I wish I had a way to graph the voltages over time to keep a log. It's a 600w psu and the server + my network stuff + my firewall is pulling 400w out of the wall so it's definitely not overtaxed. Did not test the server alone but I'd imagine it's only like 200w.
 
Bad PSU? Possibly, but they usually show their faults under load, not when they have 30% headroom available.

Memtest is a good idea; I've noticed with modern systems that hitting a bad memory location causes the system to shut down when running memory tests, so hopefully you'll come down in the morning and find your system switched off...at least then you'll know where the problem lies... :)
 
I had a similar issue in the past that ended up being a bad cable.

I would test the memory, then bypass the backplanes and install a new set of cables.
 
Changed the rest of the drives, all WD blacks, I actually see a big performance boost, at least based on benchmarks. 200-300MB/sec writes and 3GB/sec reads.

Then I started running a backup job on a new clean drive so it ran for a good 4 hours. No issues.

Ran another job, and started getting these nasty errors again. seems so random.

Code:
kjournald starting.  Commit interval 5 seconds
EXT3 FS on sdh1, internal journal
EXT3-fs: mounted filesystem with ordered data mode.
ata4: exception Emask 0x50 SAct 0x0 SErr 0x4090800 action 0xe frozen
ata4: irq_stat 0x00400040, connection status changed
ata4: SError: { HostInt PHYRdyChg 10B8B DevExch }
ata4: hard resetting link
ata4: SATA link down (SStatus 0 SControl 300)
ata4: hard resetting link
ata4: SATA link down (SStatus 0 SControl 300)
ata4: limiting SATA link speed to 1.5 Gbps
ata4: hard resetting link
ata4: SATA link down (SStatus 0 SControl 310)
ata4.00: disabled
ata4: EH complete
ata4.00: detaching (SCSI 3:0:0:0)
sd 3:0:0:0: [sdh] Synchronizing SCSI cache
sd 3:0:0:0: [sdh] Result: hostbyte=DID_BAD_TARGET driverbyte=DRIVER_OK,SUGGEST_OK
sd 3:0:0:0: [sdh] Stopping disk
sd 3:0:0:0: [sdh] START_STOP FAILED
sd 3:0:0:0: [sdh] Result: hostbyte=DID_BAD_TARGET driverbyte=DRIVER_OK,SUGGEST_OK
ata4: exception Emask 0x10 SAct 0x0 SErr 0x4050002 action 0xe frozen
ata4: irq_stat 0x00400040, connection status changed
ata4: SError: { RecovComm PHYRdyChg CommWake DevExch }
ata4: hard resetting link
ata4: link is slow to respond, please be patient (ready=0)
ata4: softreset failed (device not ready)
ata4: SATA link up 1.5 Gbps (SStatus 113 SControl 300)
ata4: link online but device misclassified, retrying
ata4: hard resetting link
ata4: SATA link up 1.5 Gbps (SStatus 113 SControl 300)
ata4.00: ATA-7: SAMSUNG HD103UJ, 1AA01110, max UDMA7
ata4.00: 1953525168 sectors, multi 0: LBA48 NCQ (depth 31/32)
ata4.00: configured for UDMA/133
ata4: EH complete
scsi 3:0:0:0: Direct-Access     ATA      SAMSUNG HD103UJ  1AA0 PQ: 0 ANSI: 5
sd 3:0:0:0: [sdh] 1953525168 512-byte hardware sectors (1000205 MB)
sd 3:0:0:0: [sdh] Write Protect is off
sd 3:0:0:0: [sdh] Mode Sense: 00 3a 00 00
sd 3:0:0:0: [sdh] Write cache: enabled, read cache: enabled, doesn't support DPO or FUA
sd 3:0:0:0: [sdh] 1953525168 512-byte hardware sectors (1000205 MB)
sd 3:0:0:0: [sdh] Write Protect is off
sd 3:0:0:0: [sdh] Mode Sense: 00 3a 00 00
sd 3:0:0:0: [sdh] Write cache: enabled, read cache: enabled, doesn't support DPO or FUA
 sdh: sdh1
sd 3:0:0:0: [sdh] Attached SCSI disk
sd 3:0:0:0: Attached scsi generic sg6 type 0
kjournald starting.  Commit interval 5 seconds
EXT3 FS on sdh1, internal journal
EXT3-fs: mounted filesystem with ordered data mode.
INFO: task kjournald:28502 blocked for more than 120 seconds.
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
kjournald     D ffff8801005d3740     0 28502      2
 ffff8801de339ca0 0000000000000046 0000000000000000 ffff88021fc04110
 ffffffff8162a500 ffffffff8162a500 ffff88010d5096e0 ffff8801dad45b80
 ffff88010d509a28 000000001fc04100 ffff8801de339c30 ffff88010d509a28
Call Trace:
 [<ffffffff8101686f>] ? read_tsc+0xe/0x24
 [<ffffffff810590b6>] ? getnstimeofday+0x54/0xb0
 [<ffffffff812c08a3>] io_schedule+0x63/0xa5
 [<ffffffff810e15e1>] sync_buffer+0x3b/0x3f
 [<ffffffff812c0de1>] __wait_on_bit+0x47/0x79
 [<ffffffff810e15a6>] ? sync_buffer+0x0/0x3f
 [<ffffffff810e15a6>] ? sync_buffer+0x0/0x3f
 [<ffffffff812c0e7d>] out_of_line_wait_on_bit+0x6a/0x77
 [<ffffffff81053861>] ? wake_bit_function+0x0/0x2a
 [<ffffffff810e150a>] __wait_on_buffer+0x36/0x3a
 [<ffffffffa0024deb>] wait_on_buffer+0x41/0x45 [jbd]
 [<ffffffffa002540e>] journal_commit_transaction+0x55d/0xf2f [jbd]
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffff81049a0d>] ? try_to_del_timer_sync+0x58/0x63
 [<ffffffffa0028ad8>] kjournald+0xe3/0x23a [jbd]
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffffa00289f5>] ? kjournald+0x0/0x23a [jbd]
 [<ffffffff810534bf>] kthread+0x49/0x76
 [<ffffffff81011719>] child_rip+0xa/0x11
 [<ffffffff81010a37>] ? restore_args+0x0/0x30
 [<ffffffff81053476>] ? kthread+0x0/0x76
 [<ffffffff8101170f>] ? child_rip+0x0/0x11

INFO: task rsync:28530 blocked for more than 120 seconds.
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
rsync         D ffff8801e09d9040     0 28530  28528
 ffff8801a97dd578 0000000000000082 ffff88021aa11e60 000312008108f44b
 ffffffff8162a500 ffffffff8162a500 ffff8801000b96e0 ffff88013b410000
 ffff8801000b9a28 0000000300000001 ffff880128285800 ffff8801000b9a28
Call Trace:
 [<ffffffff8101686f>] ? read_tsc+0xe/0x24
 [<ffffffff810590b6>] ? getnstimeofday+0x54/0xb0
 [<ffffffff812c08a3>] io_schedule+0x63/0xa5
 [<ffffffff8113dca9>] get_request_wait+0xc1/0x152
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffff8113a64f>] ? elv_merge+0x163/0x185
 [<ffffffff8113e07e>] __make_request+0x344/0x3dd
 [<ffffffff8108f44b>] ? mempool_alloc_slab+0x11/0x13
 [<ffffffff8113ca66>] generic_make_request+0x27f/0x2ba
 [<ffffffff81053824>] ? wake_up_bit+0x1e/0x23
 [<ffffffff8113cb81>] submit_bio+0xe0/0xe9
 [<ffffffff810e0297>] submit_bh+0x10b/0x12f
 [<ffffffff810e2c3a>] __block_write_full_page+0x217/0x32c
 [<ffffffffa0037e85>] ? ext3_get_block+0x0/0xfc [ext3]
 [<ffffffffa0037e85>] ? ext3_get_block+0x0/0xfc [ext3]
 [<ffffffff810e2dc9>] block_write_full_page+0x7a/0x83
 [<ffffffffa0036dc0>] ext3_ordered_writepage+0xe1/0x159 [ext3]
 [<ffffffff810939aa>] __writepage+0x12/0x2f
 [<ffffffff81094462>] write_cache_pages+0x26b/0x3b4
 [<ffffffff81093998>] ? __writepage+0x0/0x2f
 [<ffffffffa002449b>] ? do_get_write_access+0x3c4/0x404 [jbd]
 [<ffffffff810945ca>] generic_writepages+0x1f/0x25
 [<ffffffff810945ff>] do_writepages+0x2f/0x38
 [<ffffffff810dbf95>] __writeback_single_inode+0x1a2/0x332
 [<ffffffff8114ea4c>] ? prop_fraction_single+0x3c/0x5e
 [<ffffffff810dc526>] generic_sync_sb_inodes+0x245/0x390
 [<ffffffff810dc86b>] writeback_inodes+0xa4/0xfd
 [<ffffffff81094e05>] balance_dirty_pages_ratelimited_nr+0x15a/0x285
 [<ffffffff8108e1cb>] generic_file_buffered_write+0x21c/0x643
 [<ffffffff812a57ba>] ? unix_stream_recvmsg+0x5e0/0x612
 [<ffffffff810d623b>] ? mnt_drop_write+0x82/0x143
 [<ffffffff810d453d>] ? mnt_want_write+0x77/0x8d
 [<ffffffff8108e9e7>] __generic_file_aio_write_nolock+0x25e/0x292
 [<ffffffff812342ca>] ? __sock_recvmsg+0x6d/0x7a
 [<ffffffff8108f233>] generic_file_aio_write+0x67/0xc3
 [<ffffffffa003443f>] ext3_file_write+0x1e/0x9f [ext3]
 [<ffffffff810beb88>] do_sync_write+0xe7/0x12d
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffff8103c847>] ? finish_task_switch+0x31/0xc9
 [<ffffffff812c0812>] ? thread_return+0xab/0xd9
 [<ffffffff81120e84>] ? security_file_permission+0x11/0x13
 [<ffffffff810bf444>] vfs_write+0xab/0x105
 [<ffffffff810bf562>] sys_write+0x47/0x6f
 [<ffffffff8101027a>] system_call_fastpath+0x16/0x1b

INFO: task kjournald:28502 blocked for more than 120 seconds.
"echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
kjournald     D ffff8801005d3740     0 28502      2
 ffff8801de339bd0 0000000000000046 ffff88021aa11e60 0003120000000000
 ffffffff8162a500 ffffffff8162a500 ffff88010d5096e0 ffff8801dad45b80
 ffff88010d509a28 0000000000000001 ffff88010399dc00 ffff88010d509a28
Call Trace:
 [<ffffffff8101686f>] ? read_tsc+0xe/0x24
 [<ffffffff810590b6>] ? getnstimeofday+0x54/0xb0
 [<ffffffff812c08a3>] io_schedule+0x63/0xa5
 [<ffffffff8113dca9>] get_request_wait+0xc1/0x152
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffff8113a51e>] ? elv_merge+0x32/0x185
 [<ffffffff8113e07e>] __make_request+0x344/0x3dd
 [<ffffffff8108f44b>] ? mempool_alloc_slab+0x11/0x13
 [<ffffffff8113ca66>] generic_make_request+0x27f/0x2ba
 [<ffffffff8113cb81>] submit_bio+0xe0/0xe9
 [<ffffffff810e0297>] submit_bh+0x10b/0x12f
 [<ffffffffa00252c2>] journal_commit_transaction+0x411/0xf2f [jbd]
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffff81049a0d>] ? try_to_del_timer_sync+0x58/0x63
 [<ffffffffa0028ad8>] kjournald+0xe3/0x23a [jbd]
 [<ffffffff81053829>] ? autoremove_wake_function+0x0/0x38
 [<ffffffffa00289f5>] ? kjournald+0x0/0x23a [jbd]
 [<ffffffff810534bf>] kthread+0x49/0x76
 [<ffffffff81011719>] child_rip+0xa/0x11
 [<ffffffff81010a37>] ? restore_args+0x0/0x30
 [<ffffffff81053476>] ? kthread+0x0/0x76
 [<ffffffff8101170f>] ? child_rip+0x0/0x11


The top part is me inserting the disk to run the backups, I'm not sure if this happened right after or a bit later on (I wish dmesg would have date/time of events).

I still have to do the memory test as suggested, and I'll try changing the cables again.

So next step is a new controller I guess. :eek:

Though I was thinking, I may not need a new motherboard after all, I just need to get a PCI video card. There is a x16 slot, the problem is, it disables the video, but if I get a PCI video card, then I solve that issue. So think that's what I'll do, I'll order this card:
http://www.newegg.ca/Product/Product.aspx?Item=N82E16816101358

And a cheap PCI video card, and fan out cables, and hopefully I've ruled out everything by now. Next thing to try is a new power supply, I guess.
 
Back
Top