You are not logged in.

#1 2021-09-26 22:02:42

Automath
Member
Registered: 2016-05-16
Posts: 115

Self destructing hard drives...

With this it's the 4th hard drive that's agonyzing after being mounted on arch-linux.

I know that disk failure is multifactorial but arch linux is part of the equation and I'm still questioning what I'm doing wrong?

Could it be excessive use by a swap partition? but I also lost two drives with a single ext4 partition and no swap on them!.

An OS phisically breaking a hard drive seems unlikely; but the other option is that all pc parts seller on my city are pricks (which I shouldn't discount).

I mean it's not a question of years but months or even weeks before a recently bought hard drive randomly  starts to die and the only "unusual" thing I do with them is installing in a system that's almost always on but that doesn't mean that they're contantly being used.

By the other side I don't have this trouble under Debian on my other computer, though I don't use new hard drives there.

Btw in case that it matters they are all mechanical disks since I don't use any solid state.

Edit:
Here is the journalctl report:

Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#8 Sense Key : Illegal Request [current] 
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#8 Add. Sense: Unaligned write command
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#8 CDB: Read(10) 28 00 3b d4 ad e8 00 00 10 00
Sep 27 22:59:05 archiso kernel: blk_update_request: I/O error, dev sdb, sector 1003793896 op 0x0:(READ) flags 0x80700 phys_seg 1 prio class 0
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#9 FAILED Result: hostbyte=DID_OK driverbyte=DRIVER_SENSE cmd_age=4s
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#9 Sense Key : Illegal Request [current] 
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#9 Add. Sense: Unaligned write command
Sep 27 22:59:05 archiso kernel: sd 1:0:0:0: [sdb] tag#9 CDB: Read(10) 28 00 5b d4 b2 30 00 00 08 00
Sep 27 22:59:05 archiso kernel: blk_update_request: I/O error, dev sdb, sector 1540665904 op 0x0:(READ) flags 0x80700 phys_seg 1 prio class 0
Sep 27 22:59:05 archiso kernel: ata2: EH complete

And there's more somewhere on my system.
here is the output of smartctl -x

smartctl 7.2 2020-12-30 r5155 [x86_64-linux-5.11.16-arch1-1] (local build)
Copyright (C) 2002-20, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Seagate BarraCuda 3.5
Device Model:     ST1000DM010-2EP102
Serial Number:    ZN1R1X25
LU WWN Device Id: 5 000c50 0dbee2477
Firmware Version: CC46
User Capacity:    1,000,204,886,016 bytes [1.00 TB]
Sector Sizes:     512 bytes logical, 4096 bytes physical
Rotation Rate:    7200 rpm
Form Factor:      3.5 inches
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA8-ACS T13/1699-D revision 4
SATA Version is:  SATA 3.0, 6.0 Gb/s (current: 6.0 Gb/s)
Local Time is:    Mon Sep 27 23:06:25 2021 UTC
SMART support is: Available - device has SMART capability.
SMART support is: Enabled
AAM feature is:   Unavailable
APM level is:     128 (minimum power consumption without standby)
Rd look-ahead is: Enabled
Write cache is:   Enabled
DSN feature is:   Unavailable
ATA Security is:  Disabled, NOT FROZEN [SEC1]
Wt Cache Reorder: Unavailable

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x00)	Offline data collection activity
					was never started.
					Auto Offline Data Collection: Disabled.
Self-test execution status:      (   0)	The previous self-test routine completed
					without error or no self-test has ever 
					been run.
Total time to complete Offline 
data collection: 		(    0) seconds.
Offline data collection
capabilities: 			 (0x73) SMART execute Offline immediate.
					Auto Offline data collection on/off support.
					Suspend Offline collection upon new
					command.
					No Offline surface scan supported.
					Self-test supported.
					Conveyance Self-test supported.
					Selective Self-test supported.
SMART capabilities:            (0x0003)	Saves SMART data before entering
					power-saving mode.
					Supports SMART auto save timer.
Error logging capability:        (0x01)	Error logging supported.
					General Purpose Logging supported.
Short self-test routine 
recommended polling time: 	 (   1) minutes.
Extended self-test routine
recommended polling time: 	 ( 116) minutes.
Conveyance self-test routine
recommended polling time: 	 (   2) minutes.
SCT capabilities: 	       (0x1085)	SCT Status supported.

SMART Attributes Data Structure revision number: 10
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAGS    VALUE WORST THRESH FAIL RAW_VALUE
  1 Raw_Read_Error_Rate     POSR--   069   063   006    -    8617064
  3 Spin_Up_Time            PO----   099   097   000    -    0
  4 Start_Stop_Count        -O--CK   100   100   020    -    343
  5 Reallocated_Sector_Ct   PO--CK   100   100   010    -    0
  7 Seek_Error_Rate         POSR--   061   060   045    -    1413579
  9 Power_On_Hours          -O--CK   099   099   000    -    930
 10 Spin_Retry_Count        PO--C-   100   100   097    -    0
 12 Power_Cycle_Count       -O--CK   100   100   020    -    343
183 Runtime_Bad_Block       -O--CK   100   100   000    -    0
184 End-to-End_Error        -O--CK   100   100   099    -    0
187 Reported_Uncorrect      -O--CK   100   100   000    -    0
188 Command_Timeout         -O--CK   100   100   000    -    0 0 0
189 High_Fly_Writes         -O-RCK   100   100   000    -    0
190 Airflow_Temperature_Cel -O---K   067   063   040    -    33 (Min/Max 30/33)
193 Load_Cycle_Count        -O--CK   100   100   000    -    374
194 Temperature_Celsius     -O---K   033   012   000    -    33 (0 12 0 0 0)
195 Hardware_ECC_Recovered  -O-RC-   004   001   000    -    8617064
197 Current_Pending_Sector  -O--C-   100   100   000    -    0
198 Offline_Uncorrectable   ----C-   100   100   000    -    0
199 UDMA_CRC_Error_Count    -OSRCK   200   200   000    -    0
240 Head_Flying_Hours       ------   100   253   000    -    659h+27m+34.555s
241 Total_LBAs_Written      ------   100   253   000    -    425240194
242 Total_LBAs_Read         ------   100   253   000    -    990859577
                            ||||||_ K auto-keep
                            |||||__ C event count
                            ||||___ R error rate
                            |||____ S speed/performance
                            ||_____ O updated online
                            |______ P prefailure warning

General Purpose Log Directory Version 1
SMART           Log Directory Version 1 [multi-sector log support]
Address    Access  R/W   Size  Description
0x00       GPL,SL  R/O      1  Log Directory
0x01           SL  R/O      1  Summary SMART error log
0x02           SL  R/O      5  Comprehensive SMART error log
0x03       GPL     R/O      5  Ext. Comprehensive SMART error log
0x04       GPL,SL  R/O      8  Device Statistics log
0x06           SL  R/O      1  SMART self-test log
0x07       GPL     R/O      1  Extended self-test log
0x09           SL  R/W      1  Selective self-test log
0x10       GPL     R/O      1  NCQ Command Error log
0x11       GPL     R/O      1  SATA Phy Event Counters log
0x21       GPL     R/O      1  Write stream error log
0x22       GPL     R/O      1  Read stream error log
0x24       GPL     R/O    512  Current Device Internal Status Data log
0x30       GPL,SL  R/O      9  IDENTIFY DEVICE data log
0x80-0x9f  GPL,SL  R/W     16  Host vendor specific log
0xa1       GPL,SL  VS      20  Device vendor specific log
0xa2       GPL     VS    4120  Device vendor specific log
0xa8       GPL,SL  VS     129  Device vendor specific log
0xa9       GPL,SL  VS       1  Device vendor specific log
0xab       GPL     VS       1  Device vendor specific log
0xb0       GPL     VS    4800  Device vendor specific log
0xbe-0xbf  GPL     VS   65535  Device vendor specific log
0xc0       GPL,SL  VS       1  Device vendor specific log
0xc1       GPL,SL  VS      10  Device vendor specific log
0xe0       GPL,SL  R/W      1  SCT Command/Status
0xe1       GPL,SL  R/W      1  SCT Data Transfer

SMART Extended Comprehensive Error Log Version: 1 (5 sectors)
No Errors Logged

SMART Extended Self-test Log Version: 1 (1 sectors)
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Short offline       Interrupted (host reset)      00%       926         -

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

SCT Status Version:                  3
SCT Version (vendor specific):       522 (0x020a)
Device State:                        Active (0)
Current Temperature:                    32 Celsius
Power Cycle Min/Max Temperature:     31/32 Celsius
Lifetime    Min/Max Temperature:     12/36 Celsius
Under/Over Temperature Limit Count:   0/0

SCT Data Table command not supported

SCT Error Recovery Control command not supported

Device Statistics (GP Log 0x04)
Page  Offset Size        Value Flags Description
0x01  =====  =               =  ===  == General Statistics (rev 1) ==
0x01  0x008  4             343  ---  Lifetime Power-On Resets
0x01  0x010  4             930  ---  Power-on Hours
0x01  0x018  6       426097940  ---  Logical Sectors Written
0x01  0x020  6         5275103  ---  Number of Write Commands
0x01  0x028  6      1111543100  ---  Logical Sectors Read
0x01  0x030  6         6138628  ---  Number of Read Commands
0x01  0x038  6               -  ---  Date and Time TimeStamp
0x03  =====  =               =  ===  == Rotating Media Statistics (rev 1) ==
0x03  0x008  4             929  ---  Spindle Motor Power-on Hours
0x03  0x010  4             303  ---  Head Flying Hours
0x03  0x018  4             374  ---  Head Load Events
0x03  0x020  4               0  ---  Number of Reallocated Logical Sectors
0x03  0x028  4               0  ---  Read Recovery Attempts
0x03  0x030  4               0  ---  Number of Mechanical Start Failures
0x03  0x038  4               0  ---  Number of Realloc. Candidate Logical Sectors
0x04  =====  =               =  ===  == General Errors Statistics (rev 1) ==
0x04  0x008  4               0  ---  Number of Reported Uncorrectable Errors
0x04  0x010  4               0  ---  Resets Between Cmd Acceptance and Completion
0x05  =====  =               =  ===  == Temperature Statistics (rev 1) ==
0x05  0x008  1              32  ---  Current Temperature
0x05  0x010  1              31  ---  Average Short Term Temperature
0x05  0x018  1               -  ---  Average Long Term Temperature
0x05  0x020  1              36  ---  Highest Temperature
0x05  0x028  1              20  ---  Lowest Temperature
0x05  0x030  1              31  ---  Highest Average Short Term Temperature
0x05  0x038  1              27  ---  Lowest Average Short Term Temperature
0x05  0x040  1               -  ---  Highest Average Long Term Temperature
0x05  0x048  1               -  ---  Lowest Average Long Term Temperature
0x05  0x050  4               0  ---  Time in Over-Temperature
0x05  0x058  1              55  ---  Specified Maximum Operating Temperature
0x05  0x060  4               0  ---  Time in Under-Temperature
0x05  0x068  1              13  ---  Specified Minimum Operating Temperature
                                |||_ C monitored condition met
                                ||__ D supports DSN
                                |___ N normalized value

Pending Defects log (GP Log 0x0c) not supported

SATA Phy Event Counters (GP Log 0x11)
ID      Size     Value  Description
0x000a  2            1  Device-to-host register FISes sent due to a COMRESET
0x0001  2            0  Command failed due to ICRC error
0x0003  2            0  R_ERR response for device-to-host data FIS
0x0004  2            0  R_ERR response for host-to-device data FIS
0x0006  2            0  R_ERR response for device-to-host non-data FIS
0x0007  2            0  R_ERR response for host-to-device non-data FIS

And no, kernel messages are not gibberish to me, but smart tests are and not only for me since there is no standard to interpret the data and every manufacturer put there what they want, sure there are some mayor guidelines but one have to put faith that they  are met, and moreover it's not that I get intimidated but smart test are overcomplicated for something that should be yes or no, it tires me instead. I still owe you a smartctl -t "long|short|banana|do the monke ride"

For now I'm backing up ll the partitions with ddrescue but there's some chance that I have to install everything again.

The weird thing is that the physical layout of the disks seem to influence the behaviour of this particular one, e.g. unplugging some other random disk, or switching sata cables seems to change the broken state of this disk weather for better or worse... I observed this before, for a moment I thought that it was a sata cable badly connected but it was a false positive after cable swapping.

Some observations: I unplug a hard drive and use that cable to plug the failing disk, turn on computer, boot the live cd, try to read the disk again and errors misteriously dissapeared; then I poweroff and connect a new disk just in case to make my backup and after I boot again on the live I again get read errors, tried many other things but errors remain. I don't know if it's arch linux, probably not, but there's something wrong going on, even a faulty hdd's dealer doing good money in my city. I hope that this information helps in something.

@fukawi2 Now I posted the data but the post aged and nobody replies anymore I can't have the data avaible all the time, mostly when I'm using the live cd.

Edit2: Solved for now, apparently it was a false contact, the power sata plug was unaligned with the socket, weird tough because I checked by pushing it forward -- but never realocated it-- anyway it seemed firmly attached. Good old mollok? may be.

Last edited by Automath (2021-09-29 01:47:21)

Offline

#2 2021-09-26 22:18:09

V1del
Forum Moderator
Registered: 2012-10-16
Posts: 25,302

Re: Self destructing hard drives...

There's not much of use here. Have you checked their SMART stats? Have you checked the amount of writes that happen between debian and Arch, what use cases do you have on Arch or debian? Unless you run the same workload then comparing systems this way doesn't really tell much. In terms of "hardware access" both will generally use a more or less vanilla linux kernel, everything else is going to be up to what you're actually doing.

FWIW no HDD failures here in 10 years of using Arch on one so YMMV

Offline

#3 2021-09-26 22:55:51

mpan
Member
Registered: 2012-08-01
Posts: 1,630
Website

Re: Self destructing hard drives...

Is the operating system the only factor? You have mentioned that the Debian machine is not using new HDDs. Which already indicates very important variables: disk model and age. I would first suspect early mortality and a possible difference between disk models.

Swap use couldn’t cause significant wear on the storage. Unless system is badly configured, it is not even a noticeable portion of the load.

How exactly have the drives failed?

I am on two HDDs with 70kh of nearly 24/7 operation with Arch. No issues.

Offline

#4 2021-09-26 23:24:04

Automath
Member
Registered: 2016-05-16
Posts: 115

Re: Self destructing hard drives...

mpan wrote:

Is the operating system the only factor? You have mentioned that the Debian machine is not using new HDDs. Which already indicates very important variables: disk model and age. I would first suspect early mortality and a possible difference between disk models.

Swap use couldn’t cause significant wear on the storage. Unless system is badly configured, it is not even a noticeable portion of the load.

How exactly have the drives failed?

I am on two HDDs with 70kh of nearly 24/7 operation with Arch. No issues.

All of this is really weird.

How exactly have the drives failed?

It's just like I start having

 linux kernel: atax.000: exception ... freezed 

messages on my journalctl, I should have uploaded some of them but you may know them well.

And then I have to run fsck many times until I notice it's no use. Somertimes the system boots in emergency shell because of this.

disk model and age.

Yeah, ironically older disks (decades old ide) are healthier than the new ones.

Same as you, all of this is unusual for me, there's also much going on and computer market changes fast.

I wouldn't be surprised if I be scammed many times by purportedly reputable pc stores, people have fewer less scruples and more even on some countries.

Offline

#5 2021-09-26 23:39:45

Automath
Member
Registered: 2016-05-16
Posts: 115

Re: Self destructing hard drives...

V1del wrote:

There's not much of use here. Have you checked their SMART stats? Have you checked the amount of writes that happen between debian and Arch, what use cases do you have on Arch or debian? Unless you run the same workload then comparing systems this way doesn't really tell much. In terms of "hardware access" both will generally use a more or less vanilla linux kernel, everything else is going to be up to what you're actually doing.

FWIW no HDD failures here in 10 years of using Arch on one so YMMV

Have you checked their SMART stats?

Of course, though besides my efforts and research it's all gibberish to me (and to many), no

replace as soon as possible or you are going to loose your data

message this time though, but any read attemp on three partititions will trigger

ata[0-9].000: exception ... freezed

warnings, so I must assume bad sectors on disk, the only thing yet to discard is some error on the partition table.

I really lack the knowledge to discard wether those errors indicate physical damage or some other structure problem, but being that these errors are reported for many partitions on the same drive it's surely not the filesystem itself.

I should give you more detailed information and I'll add it later though they're so many cases up to now that it's a never ending report.

About case use of different systems, yeah, arch linux is my main system and I use it for work and entertainment, the other debian is just an accessory computer to use to kill boredom and browse the internet.

Offline

#6 2021-09-27 00:02:30

loqs
Member
Registered: 2014-03-06
Posts: 18,998

Re: Self destructing hard drives...

The full output of smartctl -x for the drive might still be helpful.  Do you need to recover any data stored on the failing/failed drive?

Offline

#7 2021-09-27 00:02:47

mpan
Member
Registered: 2012-08-01
Posts: 1,630
Website

Re: Self destructing hard drives...

The specific errors are important. They are not all the same and have different meanings: from benign to a total failure; and may be caused by different components. It’s not like either of us will recognize them instantly, but there are chances someone else will. Or at least their google-fu is good.

Have you tested the drives in any other computer after you considered them to be broken? Have you eliminated bad cables, broken SATA sockets, possibly the motherboard malfunctioning?

Last edited by mpan (2021-09-27 00:03:31)

Offline

#8 2021-09-27 01:37:47

HalosGhost
Forum Fellow
From: Twin Cities, MN
Registered: 2012-06-22
Posts: 2,097
Website

Re: Self destructing hard drives...

Automath wrote:

Of course, though besides my efforts and research it's all gibberish to me

This is one of many reasons why you need to actually post command output.

All the best,

-HG

Offline

#9 2021-09-27 02:50:43

Automath
Member
Registered: 2016-05-16
Posts: 115

Re: Self destructing hard drives...

HalosGhost wrote:
Automath wrote:

Of course, though besides my efforts and research it's all gibberish to me

This is one of many reasons why you need to actually post command output.

All the best,

-HG

I know; but i also know that smarctl reports are not as usefull as they're could be; surely somebody more trained than me in try and error could tell better than me that all I know that some stats MAY report important data about the disk state and failure according to the smartctl wikipedia page but as that pasge states I also know it's not totally reliable information, I'd wish it was. Later I can upload some reports.

Offline

#10 2021-09-27 03:21:37

Automath
Member
Registered: 2016-05-16
Posts: 115

Re: Self destructing hard drives...

mpan wrote:

The specific errors are important. They are not all the same and have different meanings: from benign to a total failure; and may be caused by different components. It’s not like either of us will recognize them instantly, but there are chances someone else will. Or at least their google-fu is good.

Have you tested the drives in any other computer after you considered them to be broken? Have you eliminated bad cables, broken SATA sockets, possibly the motherboard malfunctioning?

The first time I had errors over sensible data and guess what? I had not a significant backup.

Sso I've just isolated the disks, connectem them in another computer just in case; and ddrescued them, so the failure is not as cathastrophical as could be but yet I have no watrranty about the integrity of the data rescued; it just seems good.

And yes; the drives had errors also on the other computer.

As this hapenned I got enough of it and replaced all the hardware; new mobo, processor; drives EVERYTHING (It was time anyway).

Well this time I was at the process of re-establishing myself (transferring the rescued data and restoring the system) and and better prepared: I bought much more than enough storage  and made backups of everything. And soon I needed it as this hard drive containing the operative system also fails.

So that's the story; the cables are new, the temperature of the drives is 35C, I haven't checked it under stress but it's not like there's an oven inside the case; the only thing I yet not checked and yeah it's the first case scenario is if the cables are correctly plugged, but I double check every time I install a disk so I don't see imminent need to this; and anyway I'll do it as soon as I replace this failed.

If that motherboard fails I'll do computation on weaving machines instead of silicium; it's a 10 gen biostar z490 chipset; unpackaged by myself. I was already suspicious of my other motherboard being too old, but seems not the case.

Offline

#11 2021-09-27 03:34:21

Automath
Member
Registered: 2016-05-16
Posts: 115

Re: Self destructing hard drives...

loqs wrote:

The full output of smartctl -x for the drive might still be helpful.  Do you need to recover any data stored on the failing/failed drive?

Thanks for your preocupation. Yes and no. It was just the disk containing only the operative system and it's backup, luckly my home partition and other data are on separate disks. So better rescue to not go after all the installation and configuration again but I would not miss it so bad.

I just only wish the partition backup be in other drive as to not be so cautious at recovering but well, my user data is still safe, I ran out of sata slots on the motherboard and I'm about to buy a controller board in order to have more drives.

And yes this is why I tried to re invent curl, and no, my service script was taken down in order to prevent unnessary writes due to false positives, so no.

Offline

#12 2021-09-27 05:20:19

fukawi2
Ex-Administratorino
From: .vic.au
Registered: 2007-09-28
Posts: 6,237
Website

Re: Self destructing hard drives...

Automath, stop replying to your own threads. Use the Edit button if you need to add extra information to your last post.

Up to this point, you have provided zero information for anyone to actually help you, just some random thought that Arch is somehow a secret project to destroy hard drives. smartctl output and journal logs may be gibberish to you, but it's the only hard evidence of what's going on for anyone to be able to assist you.

At a random guess, I'm going to say possibly a bad power supply in your computer.

Offline

Board footer

Powered by FluxBB