You are not logged in.

#1 2013-11-15 10:42:14

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

random crash of PC (ata2 errors in journal)

I found this in my journal after my PC froze entirely, bios would not detect any harddrives for a while too.

Nov 14 20:59:17 russell-server kernel: perf samples too long (11101 > 9920), lowering kernel.perf_event_max_sample_rate to 12600
Nov 14 21:00:46 russell-server start_pms[26639]: Error obtaining Plex movie data for 2075315
Nov 14 21:01:01 russell-server crond[2204]: pam_unix(crond:session): session opened for user root by (uid=0)
Nov 14 21:01:01 russell-server CROND[2205]: (root) CMD (run-parts /etc/cron.hourly)
Nov 14 21:01:01 russell-server CROND[2204]: pam_unix(crond:session): session closed for user root
Nov 14 21:01:02 russell-server kernel: ata2: lost interrupt (Status 0x50)
Nov 14 21:01:02 russell-server kernel: ata2.01: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
Nov 14 21:01:02 russell-server kernel: ata2.01: failed command: READ DMA EXT
Nov 14 21:01:02 russell-server kernel: ata2.01: cmd 25/00:88:98:a2:b7/00:00:a9:00:00/f0 tag 0 dma 69632 in
                                                res 40/00:00:00:00:00/00:00:00:00:00/10 Emask 0x4 (timeout)
Nov 14 21:01:02 russell-server kernel: ata2.01: status: { DRDY }
Nov 14 21:01:02 russell-server kernel: ata2: soft resetting link
Nov 14 21:01:03 russell-server kernel: ata2.01: configured for UDMA/133
Nov 14 21:01:03 russell-server kernel: ata2.01: device reported invalid CHS sector 0
Nov 14 21:01:03 russell-server kernel: ata2: EH complete
Nov 14 21:01:31 russell-server start_pms[26639]: Error obtaining Plex movie data for 2302755
Nov 14 21:03:38 russell-server kernel: ata2: lost interrupt (Status 0x50)
Nov 14 21:03:38 russell-server kernel: ata2.01: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
Nov 14 21:03:38 russell-server kernel: ata2.01: failed command: READ DMA EXT
Nov 14 21:03:38 russell-server kernel: ata2.01: cmd 25/00:88:c8:5f:bf/00:00:a9:00:00/f0 tag 0 dma 69632 in
                                                res 40/00:00:00:00:00/00:00:00:00:00/10 Emask 0x4 (timeout)
Nov 14 21:03:38 russell-server kernel: ata2.01: status: { DRDY }
Nov 14 21:03:38 russell-server kernel: ata2: soft resetting link
Nov 14 21:03:39 russell-server kernel: ata2.01: configured for UDMA/133
Nov 14 21:03:39 russell-server kernel: ata2.01: device reported invalid CHS sector 0
Nov 14 21:03:39 russell-server kernel: ata2: EH complete
Nov 14 21:11:58 russell-server kernel: ata2: lost interrupt (Status 0x50)
Nov 14 21:11:58 russell-server kernel: ata2.01: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
Nov 14 21:11:58 russell-server kernel: ata2.01: failed command: READ DMA EXT
Nov 14 21:11:58 russell-server kernel: ata2.01: cmd 25/00:00:28:f6:c5/00:01:a9:00:00/f0 tag 0 dma 131072 in
                                                res 40/00:00:00:00:00/00:00:00:00:00/10 Emask 0x4 (timeout)
Nov 14 21:11:58 russell-server kernel: ata2.01: status: { DRDY }
Nov 14 21:11:58 russell-server kernel: ata2: soft resetting link
Nov 14 21:12:03 russell-server kernel: ata2.01: qc timeout (cmd 0x27)
Nov 14 21:12:03 russell-server kernel: ata2.01: failed to read native max address (err_mask=0x4)
Nov 14 21:12:03 russell-server kernel: ata2.01: HPA support seems broken, skipping HPA handling
Nov 14 21:12:03 russell-server kernel: ata2.01: revalidation failed (errno=-5)
Nov 14 21:12:08 russell-server kernel: ata2: link is slow to respond, please be patient (ready=0)
Nov 14 21:12:10 russell-server kernel: ata2: soft resetting link
Nov 14 21:12:11 russell-server kernel: ata2.01: configured for UDMA/133
Nov 14 21:12:11 russell-server kernel: ata2.01: device reported invalid CHS sector 0
Nov 14 21:12:11 russell-server kernel: ata2: EH complete
Nov 14 21:13:16 russell-server kernel: ata2: drained 7520 bytes to clear DRQ
Nov 14 21:13:16 russell-server kernel: ata2.01: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
Nov 14 21:13:16 russell-server kernel: ata2.01: failed command: READ DMA EXT
Nov 14 21:13:16 russell-server kernel: ata2.01: cmd 25/00:00:50:23:ce/00:01:a9:00:00/f0 tag 0 dma 131072 in
                                                res 40/00:00:00:00:00/00:00:00:00:00/10 Emask 0x4 (timeout)
Nov 14 21:13:16 russell-server kernel: ata2.01: status: { DRDY }
Nov 14 21:13:19 russell-server kernel: ata2: soft resetting link
Nov 14 21:13:26 russell-server kernel: ata2.01: qc timeout (cmd 0xec)
Nov 14 21:13:26 russell-server kernel: ata2.01: failed to IDENTIFY (I/O error, err_mask=0x4)
Nov 14 21:13:26 russell-server kernel: ata2.01: revalidation failed (errno=-5)
Nov 14 21:13:31 russell-server kernel: ata2: link is slow to respond, please be patient (ready=0)
Nov 14 21:13:36 russell-server kernel: ata2: device not ready (errno=-16), forcing hardreset
Nov 14 21:13:36 russell-server kernel: ata2: soft resetting link
Nov 14 21:13:37 russell-server kernel: ata2.01: configured for UDMA/133
Nov 14 21:13:37 russell-server kernel: ata2.01: device reported invalid CHS sector 0
Nov 14 21:13:37 russell-server kernel: ata2: EH complete
Nov 14 21:14:23 russell-server kernel: ata2.01: exception Emask 0x0 SAct 0x0 SErr 0x0 action 0x6 frozen
Nov 14 21:14:23 russell-server kernel: ata2.01: failed command: READ DMA EXT
Nov 14 21:14:23 russell-server kernel: ata2.01: cmd 25/00:00:50:e7:b3/00:01:a9:00:00/f0 tag 0 dma 131072 in
                                                res 40/00:00:00:00:00/00:00:00:00:00/10 Emask 0x4 (timeout)
Nov 14 21:14:23 russell-server kernel: ata2.01: status: { DRDY }
Nov 14 21:14:27 russell-server kernel: ata2: soft resetting link
Nov 14 21:14:34 russell-server kernel: ata2.01: qc timeout (cmd 0xec)
Nov 14 21:14:34 russell-server kernel: ata2.01: failed to IDENTIFY (I/O error, err_mask=0x4)
Nov 14 21:14:34 russell-server kernel: ata2.01: revalidation failed (errno=-5)
Nov 14 21:14:36 russell-server kernel: ata2: soft resetting link
Nov 14 21:14:49 russell-server kernel: ata2.01: qc timeout (cmd 0xec)
Nov 14 21:14:49 russell-server kernel: ata2.01: failed to IDENTIFY (I/O error, err_mask=0x4)
Nov 14 21:14:49 russell-server kernel: ata2.01: revalidation failed (errno=-5)
Nov 14 21:14:54 russell-server kernel: ata2: link is slow to respond, please be patient (ready=0)
Nov 14 21:14:54 russell-server kernel: ata2: soft resetting link
Nov 14 21:14:55 russell-server kernel: ata2.01: configured for UDMA/133
Nov 14 21:14:55 russell-server kernel: ata2.01: device reported invalid CHS sector 0
Nov 14 21:14:55 russell-server kernel: ata2: EH complete
Nov 14 22:01:01 russell-server crond[3852]: pam_unix(crond:session): session opened for user root by (uid=0)

busted motherboard?

edit: found this https://bugzilla.redhat.com/show_bug.cgi?id=549981



Are there any tests I can run?

Last edited by jrussell (2013-11-15 13:56:06)


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

#2 2013-11-15 15:54:44

x33a
Forum Fellow
Registered: 2009-08-15
Posts: 4,587

Re: random crash of PC (ata2 errors in journal)

Try a SMART scan for starters. Either the drive or the cable is going.

Offline

#3 2013-11-15 19:02:40

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

Re: random crash of PC (ata2 errors in journal)

getting a lot of:

Nov 15 16:03:34 russell-server kernel: ata1: SATA max UDMA/133 cmd 0x1f0 ctl 0x3f6 bmdma 0xf000 irq 14
Nov 15 16:03:34 russell-server kernel: ata2: SATA max UDMA/133 cmd 0x170 ctl 0x376 bmdma 0xf008 irq 15
Nov 15 16:03:34 russell-server kernel: ata2.01: ATA-8: WDC WD20EARX-00PASB0, 51.0AB51, max UDMA/133
Nov 15 16:03:34 russell-server kernel: ata2.01: 3907029168 sectors, multi 16: LBA48 NCQ (depth 0/32)
Nov 15 16:03:34 russell-server kernel: ata2.01: configured for UDMA/133
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 2 using ehci-pci
Nov 15 16:03:34 russell-server kernel: ata1.00: HPA detected: current 312579695, native 312581808
Nov 15 16:03:34 russell-server kernel: ata1.00: ATA-7: ST3160811AS, 3.AAE, max UDMA/133
Nov 15 16:03:34 russell-server kernel: ata1.00: 312579695 sectors, multi 16: LBA48 NCQ (depth 0/32)
Nov 15 16:03:34 russell-server kernel: ata1.01: HPA detected: current 1953523055, native 1953525168
Nov 15 16:03:34 russell-server kernel: ata1.01: ATA-8: WDC WD1000FYPS-01ZKB0, 02.01B01, max UDMA/133
Nov 15 16:03:34 russell-server kernel: ata1.01: 1953523055 sectors, multi 16: LBA48 NCQ (depth 0/32)
Nov 15 16:03:34 russell-server kernel: ata1.00: configured for UDMA/133
Nov 15 16:03:34 russell-server kernel: ata1.01: configured for UDMA/133
Nov 15 16:03:34 russell-server kernel: scsi 0:0:0:0: Direct-Access     ATA      ST3160811AS      3.AA PQ: 0 ANSI: 5
Nov 15 16:03:34 russell-server kernel: scsi 0:0:1:0: Direct-Access     ATA      WDC WD1000FYPS-0 02.0 PQ: 0 ANSI: 5
Nov 15 16:03:34 russell-server kernel: scsi 1:0:1:0: Direct-Access     ATA      WDC WD20EARX-00P 51.0 PQ: 0 ANSI: 5
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] 312579695 512-byte logical blocks: (160 GB/149 GiB)
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] 1953523055 512-byte logical blocks: (1.00 TB/931 GiB)
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] Write Protect is off
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] Mode Sense: 00 3a 00 00
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] Write cache: enabled, read cache: enabled, doesn't support DPO or FUA
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] Write Protect is off
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] Mode Sense: 00 3a 00 00
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] Write cache: enabled, read cache: enabled, doesn't support DPO or FUA
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] 3907029168 512-byte logical blocks: (2.00 TB/1.81 TiB)
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] 4096-byte physical blocks
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] Write Protect is off
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] Mode Sense: 00 3a 00 00
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] Write cache: enabled, read cache: enabled, doesn't support DPO or FUA
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel:  sdc: sdc1
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel:  sda: sda1 sda2
Nov 15 16:03:34 russell-server kernel:  sdb: sdb1
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: tsc: Refined TSC clocksource calibration: 1607.727 MHz
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 3 using ehci-pci
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: EXT4-fs (sda2): mounted filesystem with ordered data mode. Opts: (null)
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 4 using ehci-pci
Nov 15 16:03:34 russell-server kernel: usb 5-5: device not accepting address 4, error -71
Nov 15 16:03:34 russell-server kernel: Switched to clocksource tsc
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 5 using ehci-pci
Nov 15 16:03:34 russell-server systemd[1]: systemd 208 running in system mode. (+PAM -LIBWRAP -AUDIT -SELINUX -IMA -SYSVINIT +LIBCRYPTSETUP +GCRYPT +ACL +XZ)
Nov 15 16:03:34 russell-server systemd[1]: Set hostname to <russell-server>.
Nov 15 16:03:34 russell-server kernel: usb 5-5: device not accepting address 5, error -71
Nov 15 16:03:34 russell-server kernel: hub 5-0:1.0: unable to enumerate USB device on port 5
Nov 15 16:03:34 russell-server kernel: usb 3-1: new full-speed USB device number 2 using uhci_hcd
Nov 15 16:03:34 russell-server systemd[1]: Starting Forward Password Requests to Wall Directory Watch.
Nov 15 16:03:34 russell-server systemd[1]: Started Forward Password Requests to Wall Directory Watch.
Nov 15 16:03:34 russell-server systemd[1]: Expecting device sys-subsystem-net-devices-enp2s5.device...
Nov 15 16:03:34 russell-server systemd[1]: Starting Remote File Systems.
Nov 15 16:03:34 russell-server systemd[1]: Reached target Remote File Systems.
Nov 15 16:03:34 russell-server systemd[1]: Starting Device-mapper event daemon FIFOs.
Nov 15 16:03:34 russell-server systemd[1]: Listening on Device-mapper event daemon FIFOs.
Nov 15 16:03:34 russell-server systemd[1]: Starting Delayed Shutdown Socket.
Nov 15 16:03:34 russell-server systemd[1]: Listening on Delayed Shutdown Socket.

More specifically:

Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel:  sdc: sdc1
Nov 15 16:03:34 russell-server kernel: sd 1:0:1:0: [sdc] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel:  sda: sda1 sda2
Nov 15 16:03:34 russell-server kernel:  sdb: sdb1
Nov 15 16:03:34 russell-server kernel: sd 0:0:1:0: [sdb] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel: sd 0:0:0:0: [sda] Attached SCSI disk
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: tsc: Refined TSC clocksource calibration: 1607.727 MHz
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 3 using ehci-pci
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: usb 5-5: device descriptor read/64, error -71
Nov 15 16:03:34 russell-server kernel: EXT4-fs (sda2): mounted filesystem with ordered data mode. Opts: (null)
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 4 using ehci-pci
Nov 15 16:03:34 russell-server kernel: usb 5-5: device not accepting address 4, error -71
Nov 15 16:03:34 russell-server kernel: Switched to clocksource tsc
Nov 15 16:03:34 russell-server kernel: usb 5-5: new high-speed USB device number 5 using ehci-pci
Nov 15 16:03:34 russell-server systemd[1]: systemd 208 running in system mode. (+PAM -LIBWRAP -AUDIT -SELINUX -IMA -SYSVINIT +LIBCRYPTSETUP +GCRYPT +ACL +XZ)
Nov 15 16:03:34 russell-server systemd[1]: Set hostname to <russell-server>.
Nov 15 16:03:34 russell-server kernel: usb 5-5: device not accepting address 5, error -71
Nov 15 16:03:34 russell-server kernel: hub 5-0:1.0: unable to enumerate USB device on port 5

Are these errors related to the ata errors in post one?

Last edited by jrussell (2013-11-15 19:03:41)


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

#4 2013-11-16 05:05:07

x33a
Forum Fellow
Registered: 2009-08-15
Posts: 4,587

Re: random crash of PC (ata2 errors in journal)

As for the usb errors, do you have any external hard drives attached?

Also, as I mentioned earlier, run a SMART scan.

Install smartmontools and first do

# smartctl --all /dev/<device> > smart.out

to generate a log. And post its result here.

and then do

# smartctl --test=short /dev/<device>

to do a quick scan of the hard drive. And see what result it gives.

Offline

#5 2013-11-16 05:52:35

WonderWoofy
Member
From: Los Gatos, CA
Registered: 2012-05-19
Posts: 8,414

Re: random crash of PC (ata2 errors in journal)

If you have important data on the drive, you should probably ensure that your backups of it are current.  In the event that the drive is on the way out, you probably want to do that before anything else.

Offline

#6 2013-11-16 11:31:23

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

Re: random crash of PC (ata2 errors in journal)

Ok thanks for the help so far, USB errors seem to be related to the new USB printer, its the only USB device plugged into the PC.

Everything is backed up.

Here is lsblk:

NAME   MAJ:MIN RM   SIZE RO TYPE MOUNTPOINT
sda      8:0    0 149.1G  0 disk
├─sda1   8:1    0  1007K  0 part
└─sda2   8:2    0   144G  0 part /
sdb      8:16   0 931.5G  0 disk
└─sdb1   8:17   0 931.5G  0 part
sdc      8:32   0   1.8T  0 disk
└─sdc1   8:33   0   1.8T  0 part /mnt/storage

ran

smartctl --all /dev/<device> > smart.out

per harddrive:

sda:

russell-server% sudo smartctl --all /dev/sda > sda.out
russell-server% cat sda.out
smartctl 6.2 2013-07-26 r3841 [x86_64-linux-3.12.0-1-ARCH] (local build)
Copyright (C) 2002-13, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Seagate Barracuda 7200.9
Device Model:     ST3160811AS
Serial Number:    6PT5P4WZ
Firmware Version: 3.AAE
User Capacity:    160,040,803,840 bytes [160 GB]
Sector Size:      512 bytes logical/physical
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA/ATAPI-7 (minor revision not indicated)
Local Time is:    Sat Nov 16 13:32:46 2013 SAST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x82) Offline data collection activity
                                        was completed without error.
                                        Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (  430) seconds.
Offline data collection
capabilities:                    (0x5b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        No Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   1) minutes.
Extended self-test routine
recommended polling time:        (  54) minutes.

SMART Attributes Data Structure revision number: 10
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x000f   117   090   006    Pre-fail  Always       -       163744227
  3 Spin_Up_Time            0x0003   095   095   000    Pre-fail  Always       -       0
  4 Start_Stop_Count        0x0032   100   100   020    Old_age   Always       -       1000
  5 Reallocated_Sector_Ct   0x0033   100   100   036    Pre-fail  Always       -       0
  7 Seek_Error_Rate         0x000f   088   060   030    Pre-fail  Always       -       744116533
  9 Power_On_Hours          0x0032   063   063   000    Old_age   Always       -       32820
 10 Spin_Retry_Count        0x0013   100   100   097    Pre-fail  Always       -       0
 12 Power_Cycle_Count       0x0032   099   099   020    Old_age   Always       -       1096
187 Reported_Uncorrect      0x0032   092   092   000    Old_age   Always       -       8
189 High_Fly_Writes         0x003a   100   100   000    Old_age   Always       -       0
190 Airflow_Temperature_Cel 0x0022   065   051   045    Old_age   Always       -       35 (Min/Max 33/35)
194 Temperature_Celsius     0x0022   035   049   000    Old_age   Always       -       35 (0 16 0 0 0)
195 Hardware_ECC_Recovered  0x001a   069   045   000    Old_age   Always       -       75074513
197 Current_Pending_Sector  0x0012   001   001   000    Old_age   Always       -       4294967295
198 Offline_Uncorrectable   0x0010   001   001   000    Old_age   Offline      -       4294967295
199 UDMA_CRC_Error_Count    0x003e   200   200   000    Old_age   Always       -       0
200 Multi_Zone_Error_Rate   0x0000   100   253   000    Old_age   Offline      -       0
202 Data_Address_Mark_Errs  0x0032   100   253   000    Old_age   Always       -       0

SMART Error Log Version: 1
ATA Error Count: 4859 (device log contains only the most recent five errors)
        CR = Command Register [HEX]
        FR = Features Register [HEX]
        SC = Sector Count Register [HEX]
        SN = Sector Number Register [HEX]
        CL = Cylinder Low Register [HEX]
        CH = Cylinder High Register [HEX]
        DH = Device/Head Register [HEX]
        DC = Device Command Register [HEX]
        ER = Error register [HEX]
        ST = Status register [HEX]
Powered_Up_Time is measured from power on, and printed as
DDd+hh:mm:SS.sss where DD=days, hh=hours, mm=minutes,
SS=sec, and sss=millisec. It "wraps" after 49.710 days.

Error 4859 occurred at disk power-on lifetime: 32819 hours (1367 days + 11 hours)
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  10 51 01 6e 96 a1 e2

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
  -- -- -- -- -- -- -- --  ----------------  --------------------
  37 00 01 6e 96 a1 e2 00      04:37:49.675  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 00 6e 96 a1 e0 00      04:37:49.603  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 00 6e 96 a1 e2 00      04:37:47.702  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 01 6e 96 a1 e0 00      04:37:47.079  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 01 6e 96 a1 e2 00      04:37:47.033  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]

Error 4858 occurred at disk power-on lifetime: 32819 hours (1367 days + 11 hours)
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  10 51 01 6e 96 a1 e2

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
  -- -- -- -- -- -- -- --  ----------------  --------------------
  37 00 01 6e 96 a1 e2 00      04:37:49.675  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 00 6e 96 a1 e0 00      04:37:49.603  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 00 6e 96 a1 e2 00      04:37:47.702  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 01 af 9e a1 e0 00      04:37:47.079  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  29 00 01 af 9e a1 e0 00      04:37:47.033  READ MULTIPLE EXT

Error 4857 occurred at disk power-on lifetime: 32806 hours (1366 days + 22 hours)
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  10 51 01 6e 96 a1 e2

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
  -- -- -- -- -- -- -- --  ----------------  --------------------
  37 00 01 6e 96 a1 e2 00      04:28:07.943  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 00 6e 96 a1 e0 00      04:28:07.718  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 00 6e 96 a1 e2 00      04:28:07.656  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 01 6e 96 a1 e0 00      04:28:05.574  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 01 6e 96 a1 e2 00      04:28:05.515  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]

Error 4856 occurred at disk power-on lifetime: 32806 hours (1366 days + 22 hours)
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  10 51 01 6e 96 a1 e2

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
  -- -- -- -- -- -- -- --  ----------------  --------------------
  37 00 01 6e 96 a1 e2 00      04:28:02.150  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 00 6e 96 a1 e0 00      04:28:01.231  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 00 6e 96 a1 e2 00      04:28:01.231  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 01 af 9e a1 e0 00      04:28:05.574  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  29 00 01 af 9e a1 e0 00      04:28:05.515  READ MULTIPLE EXT

Error 4855 occurred at disk power-on lifetime: 32801 hours (1366 days + 17 hours)
  When the command that caused the error occurred, the device was active or idle.

  After command completion occurred, registers were:
  ER ST SC SN CL CH DH
  -- -- -- -- -- -- --
  10 51 01 6e 96 a1 e2

  Commands leading to the command that caused the error were:
  CR FR SC SN CL CH DH DC   Powered_Up_Time  Command/Feature_Name
  -- -- -- -- -- -- -- --  ----------------  --------------------
  37 00 01 6e 96 a1 e2 00      08:16:52.078  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 00 6e 96 a1 e0 00      08:16:52.010  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 00 6e 96 a1 e2 00      08:16:50.057  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  27 00 01 6e 96 a1 e0 00      08:16:49.290  READ NATIVE MAX ADDRESS EXT [OBS-ACS-3]
  37 00 01 6e 96 a1 e2 00      08:16:49.247  SET NATIVE MAX ADDRESS EXT [OBS-ACS-3]

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Extended offline    Completed without error       00%     32802         -
# 2  Short offline       Completed without error       00%     32801         -
# 3  Short offline       Completed without error       00%     17765         -

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

russell-server%

sdb:

russell-server% sudo smartctl --all /dev/sdb > sdb.out
russell-server% cat sdb.out
smartctl 6.2 2013-07-26 r3841 [x86_64-linux-3.12.0-1-ARCH] (local build)
Copyright (C) 2002-13, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Western Digital RE2-GP
Device Model:     WDC WD1000FYPS-01ZKB0
Serial Number:    WD-WCASJ1172116
LU WWN Device Id: 5 0014ee 2abb9c188
Firmware Version: 02.01B01
User Capacity:    1,000,203,804,160 bytes [1.00 TB]
Sector Size:      512 bytes logical/physical
Rotation Rate:    5400 rpm
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA8-ACS (minor revision not indicated)
SATA Version is:  SATA 2.5, 3.0 Gb/s
Local Time is:    Sat Nov 16 13:33:40 2013 SAST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x84) Offline data collection activity
                                        was suspended by an interrupting command from host.
                                        Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (27960) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        ( 320) minutes.
Conveyance self-test routine
recommended polling time:        (   5) minutes.
SCT capabilities:              (0x303f) SCT Status supported.
                                        SCT Error Recovery Control supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x000f   200   200   051    Pre-fail  Always       -       0
  3 Spin_Up_Time            0x0003   253   174   021    Pre-fail  Always       -       4275
  4 Start_Stop_Count        0x0032   100   100   000    Old_age   Always       -       170
  5 Reallocated_Sector_Ct   0x0033   200   200   140    Pre-fail  Always       -       0
  7 Seek_Error_Rate         0x000e   200   200   000    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   093   093   000    Old_age   Always       -       5778
 10 Spin_Retry_Count        0x0012   100   100   000    Old_age   Always       -       0
 11 Calibration_Retry_Count 0x0012   100   100   000    Old_age   Always       -       0
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       146
192 Power-Off_Retract_Count 0x0032   200   200   000    Old_age   Always       -       105
193 Load_Cycle_Count        0x0032   188   188   000    Old_age   Always       -       38003
194 Temperature_Celsius     0x0022   118   100   000    Old_age   Always       -       34
196 Reallocated_Event_Count 0x0032   200   200   000    Old_age   Always       -       0
197 Current_Pending_Sector  0x0012   200   200   000    Old_age   Always       -       0
198 Offline_Uncorrectable   0x0010   200   200   000    Old_age   Offline      -       0
199 UDMA_CRC_Error_Count    0x003e   200   200   000    Old_age   Always       -       0
200 Multi_Zone_Error_Rate   0x0008   200   200   000    Old_age   Offline      -       0

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Extended offline    Completed without error       00%      5776         -
# 2  Short offline       Completed without error       00%      5771         -

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

russell-server%

sdc:

russell-server% sudo smartctl --all /dev/sdc > sdc.out
russell-server% cat sdc.out
smartctl 6.2 2013-07-26 r3841 [x86_64-linux-3.12.0-1-ARCH] (local build)
Copyright (C) 2002-13, Bruce Allen, Christian Franke, www.smartmontools.org

=== START OF INFORMATION SECTION ===
Model Family:     Western Digital Caviar Green (AF, SATA 6Gb/s)
Device Model:     WDC WD20EARX-00PASB0
Serial Number:    WD-WCAZAD562439
LU WWN Device Id: 5 0014ee 05859ce68
Firmware Version: 51.0AB51
User Capacity:    2,000,398,934,016 bytes [2.00 TB]
Sector Sizes:     512 bytes logical, 4096 bytes physical
Device is:        In smartctl database [for details use: -P show]
ATA Version is:   ATA8-ACS (minor revision not indicated)
SATA Version is:  SATA 3.0, 6.0 Gb/s (current: 3.0 Gb/s)
Local Time is:    Sat Nov 16 13:34:38 2013 SAST
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status:  (0x82) Offline data collection activity
                                        was completed without error.
                                        Auto Offline Data Collection: Enabled.
Self-test execution status:      (   0) The previous self-test routine completed
                                        without error or no self-test has ever
                                        been run.
Total time to complete Offline
data collection:                (38880) seconds.
Offline data collection
capabilities:                    (0x7b) SMART execute Offline immediate.
                                        Auto Offline data collection on/off support.
                                        Suspend Offline collection upon new
                                        command.
                                        Offline surface scan supported.
                                        Self-test supported.
                                        Conveyance Self-test supported.
                                        Selective Self-test supported.
SMART capabilities:            (0x0003) Saves SMART data before entering
                                        power-saving mode.
                                        Supports SMART auto save timer.
Error logging capability:        (0x01) Error logging supported.
                                        General Purpose Logging supported.
Short self-test routine
recommended polling time:        (   2) minutes.
Extended self-test routine
recommended polling time:        ( 375) minutes.
Conveyance self-test routine
recommended polling time:        (   5) minutes.
SCT capabilities:              (0x3035) SCT Status supported.
                                        SCT Feature Control supported.
                                        SCT Data Table supported.

SMART Attributes Data Structure revision number: 16
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME          FLAG     VALUE WORST THRESH TYPE      UPDATED  WHEN_FAILED RAW_VALUE
  1 Raw_Read_Error_Rate     0x002f   200   200   051    Pre-fail  Always       -       3
  3 Spin_Up_Time            0x0027   239   158   021    Pre-fail  Always       -       3016
  4 Start_Stop_Count        0x0032   100   100   000    Old_age   Always       -       826
  5 Reallocated_Sector_Ct   0x0033   200   200   140    Pre-fail  Always       -       0
  7 Seek_Error_Rate         0x002e   200   200   000    Old_age   Always       -       0
  9 Power_On_Hours          0x0032   083   083   000    Old_age   Always       -       12866
 10 Spin_Retry_Count        0x0032   100   100   000    Old_age   Always       -       0
 11 Calibration_Retry_Count 0x0032   100   100   000    Old_age   Always       -       0
 12 Power_Cycle_Count       0x0032   100   100   000    Old_age   Always       -       224
192 Power-Off_Retract_Count 0x0032   200   200   000    Old_age   Always       -       141
193 Load_Cycle_Count        0x0032   146   146   000    Old_age   Always       -       164822
194 Temperature_Celsius     0x0022   117   101   000    Old_age   Always       -       33
196 Reallocated_Event_Count 0x0032   200   200   000    Old_age   Always       -       0
197 Current_Pending_Sector  0x0032   200   200   000    Old_age   Always       -       1
198 Offline_Uncorrectable   0x0030   200   200   000    Old_age   Offline      -       1
199 UDMA_CRC_Error_Count    0x0032   200   200   000    Old_age   Always       -       0
200 Multi_Zone_Error_Rate   0x0008   200   200   000    Old_age   Offline      -       1

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
Num  Test_Description    Status                  Remaining  LifeTime(hours)  LBA_of_first_error
# 1  Extended offline    Completed: read failure       90%     12845         2208103368
# 2  Short offline       Completed: read failure       90%     12845         2208103368

SMART Selective self-test log data structure revision number 1
 SPAN  MIN_LBA  MAX_LBA  CURRENT_TEST_STATUS
    1        0        0  Not_testing
    2        0        0  Not_testing
    3        0        0  Not_testing
    4        0        0  Not_testing
    5        0        0  Not_testing
Selective self-test flags (0x0):
  After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.

russell-server%

Last edited by jrussell (2013-11-16 11:40:13)


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

#7 2013-11-16 18:29:08

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

Re: random crash of PC (ata2 errors in journal)

Here is the journal when it happens

Nov 16 17:52:24 russell-server sshd[3254]: pam_unix(sshd:auth): authentication failure; logname= uid=0 euid=0 tty=ssh ruser= rhost=lt49787.lagserv.info
Nov 16 17:52:27 russell-server sshd[3254]: Failed password for invalid user demo from 88.190.63.53 port 56283 ssh2
Nov 16 17:52:27 russell-server sshd[3254]: Received disconnect from 88.190.63.53: 11: Bye Bye [preauth]
Nov 16 17:52:29 russell-server sshd[3257]: Invalid user oracle from 88.190.63.53
Nov 16 17:52:29 russell-server sshd[3257]: input_userauth_request: invalid user oracle [preauth]
Nov 16 17:52:29 russell-server sshd[3257]: pam_tally(sshd:auth): pam_get_uid; no such user
Nov 16 17:52:29 russell-server sshd[3257]: pam_unix(sshd:auth): check pass; user unknown
Nov 16 17:52:29 russell-server sshd[3257]: pam_unix(sshd:auth): authentication failure; logname= uid=0 euid=0 tty=ssh ruser= rhost=lt49787.lagserv.info
Nov 16 17:52:31 russell-server sshd[3257]: Failed password for invalid user oracle from 88.190.63.53 port 56557 ssh2
Nov 16 17:52:31 russell-server sshd[3257]: Received disconnect from 88.190.63.53: 11: Bye Bye [preauth]
Nov 16 17:52:33 russell-server sshd[3259]: Invalid user postgrest from 88.190.63.53
Nov 16 17:52:33 russell-server sshd[3259]: input_userauth_request: invalid user postgrest [preauth]
Nov 16 17:52:33 russell-server sshd[3259]: pam_tally(sshd:auth): pam_get_uid; no such user
Nov 16 17:52:33 russell-server sshd[3259]: pam_unix(sshd:auth): check pass; user unknown
Nov 16 17:52:33 russell-server sshd[3259]: pam_unix(sshd:auth): authentication failure; logname= uid=0 euid=0 tty=ssh ruser= rhost=lt49787.lagserv.info
Nov 16 17:52:35 russell-server sshd[3259]: Failed password for invalid user postgrest from 88.190.63.53 port 56794 ssh2
Nov 16 17:52:35 russell-server sshd[3259]: Received disconnect from 88.190.63.53: 11: Bye Bye [preauth]
Nov 16 17:55:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Currently unreadable (pending) sectors
Nov 16 17:55:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Offline uncorrectable sectors
Nov 16 18:01:01 russell-server crond[3344]: pam_unix(crond:session): session opened for user root by (uid=0)
Nov 16 18:01:01 russell-server CROND[3345]: (root) CMD (run-parts /etc/cron.hourly)
Nov 16 18:01:01 russell-server CROND[3344]: pam_unix(crond:session): session closed for user root
Nov 16 18:25:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 64 to 63
Nov 16 18:25:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 194 Temperature_Celsius changed from 36 to 37
Nov 16 18:25:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 195 Hardware_ECC_Recovered changed from 55 to 53
Nov 16 18:25:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Currently unreadable (pending) sectors
Nov 16 18:25:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Offline uncorrectable sectors
Nov 16 18:55:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 63 to 64
Nov 16 18:55:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 194 Temperature_Celsius changed from 37 to 36
Nov 16 18:55:07 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Currently unreadable (pending) sectors
Nov 16 18:55:07 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Offline uncorrectable sectors
Nov 16 19:01:01 russell-server crond[3889]: pam_unix(crond:session): session opened for user root by (uid=0)
Nov 16 19:01:01 russell-server CROND[3890]: (root) CMD (run-parts /etc/cron.hourly)
Nov 16 19:01:01 russell-server CROND[3889]: pam_unix(crond:session): session closed for user root
Nov 16 19:25:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 190 Airflow_Temperature_Cel changed from 64 to 63
Nov 16 19:25:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 194 Temperature_Celsius changed from 36 to 37
Nov 16 19:25:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Currently unreadable (pending) sectors
Nov 16 19:25:08 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Offline uncorrectable sectors
Nov 16 19:35:46 russell-server murmurd[255]: <W>2013-11-16 19:35:46.915 1 => <9:avoided(-1)> Connection closed: The remote host closed the connection [1]
Nov 16 19:35:49 russell-server murmurd[255]: <W>2013-11-16 19:35:49.064 1 => <13:(-1)> New connection: 105.236.150.144:54958
Nov 16 19:35:49 russell-server murmurd[255]: <W>2013-11-16 19:35:49.347 1 => <13:(-1)> Client version 1.2.4 (Win: 1.2.4)
Nov 16 19:35:49 russell-server murmurd[255]: <W>2013-11-16 19:35:49.399 1 => <13:avoided(-1)> Authenticated
Nov 16 19:53:33 russell-server murmurd[255]: <W>2013-11-16 19:53:33.536 1 => <12:baden(-1)> Connection closed: The remote host closed the connection [1]
Nov 16 19:54:00 russell-server murmurd[255]: <W>2013-11-16 19:54:00.049 1 => <11:Jason-iPad(-1)> Connection closed: The remote host closed the connection [1]
Nov 16 19:54:00 russell-server murmurd[255]: <W>2013-11-16 19:54:00.101 1 => CELT codec switch ffffffff8000000b ffffffff80000010 (prefer ffffffff80000010) (Opus 1)
Nov 16 19:55:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Prefailure Attribute: 1 Raw_Read_Error_Rate changed from 118 to 113
Nov 16 19:55:07 russell-server smartd[169]: Device: /dev/sda [SAT], SMART Usage Attribute: 195 Hardware_ECC_Recovered changed from 53 to 52
Nov 16 19:55:07 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Currently unreadable (pending) sectors
Nov 16 19:55:07 russell-server smartd[169]: Device: /dev/sdc [SAT], 1 Offline uncorrectable sectors
Nov 16 19:55:07 russell-server smartd[169]: Device: /dev/sdc [SAT], SMART Usage Attribute: 194 Temperature_Celsius changed from 109 to 107
Nov 16 19:56:31 russell-server murmurd[255]: <W>2013-11-16 19:56:31.625 1 => <8:Jason(-1)> Connection closed: The remote host closed the connection [1]
Nov 16 19:57:02 russell-server murmurd[255]: <W>2013-11-16 19:57:02.629 1 => <14:(-1)> New connection: 41.13.208.60:35279
Nov 16 19:57:02 russell-server murmurd[255]: <W>2013-11-16 19:57:02.929 1 => <14:(-1)> Connection closed: The remote host closed the connection [1]
Nov 16 19:57:03 russell-server murmurd[255]: <W>2013-11-16 19:57:03.087 1 => <15:(-1)> New connection: 41.13.208.60:35281
Nov 16 19:57:03 russell-server murmurd[255]: <W>2013-11-16 19:57:03.470 1 => <15:(-1)> Client version 1.2.4 (iPhone OS: Mumble for iOS 1.2.2)
Nov 16 19:57:03 russell-server murmurd[255]: <W>2013-11-16 19:57:03.522 1 => <15:Rob iPhone(-1)> Rejected connection: Invalid username
Nov 16 19:57:03 russell-server murmurd[255]: <W>2013-11-16 19:57:03.574 1 => <15:Rob iPhone(-1)> Connection closed:  [-1]
Nov 16 20:01:01 russell-server crond[4520]: pam_unix(crond:session): session opened for user root by (uid=0)
Nov 16 20:01:01 russell-server CROND[4521]: (root) CMD (run-parts /etc/cron.hourly)
Nov 16 20:01:01 russell-server CROND[4520]: pam_unix(crond:session): session closed for user root
-- Reboot --
Nov 16 20:22:20 russell-server systemd-journal[94]: Runtime journal is using 544.0K (max 150.7M, leaving 226.0M of free 1.4G, current limit 150.7M).
Nov 16 20:22:20 russell-server systemd-journal[94]: Runtime journal is using 548.0K (max 150.7M, leaving 226.0M of free 1.4G, current limit 150.7M).
Nov 16 20:22:20 russell-server kernel: Initializing cgroup subsys cpuset
Nov 16 20:22:20 russell-server kernel: Initializing cgroup subsys cpu
Nov 16 20:22:20 russell-server kernel: Initializing cgroup subsys cpuacct
Nov 16 20:22:20 russell-server kernel: Linux version 3.12.0-1-ARCH (tobias@T-POWA-LX) (gcc version 4.8.2 (GCC) ) #1 SMP PREEMPT Wed Nov 6 09:06:27 CET 2013
Nov 16 20:22:20 russell-server kernel: Command line: BOOT_IMAGE=../vmlinuz-linux root=/dev/sda2 rw initrd=../initramfs-linux.img
Nov 16 20:22:20 russell-server kernel: e820: BIOS-provided physical RAM map:
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x0000000000000000-0x000000000009f7ff] usable
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x000000000009f800-0x000000000009ffff] reserved
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x00000000000f0000-0x00000000000fffff] reserved
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x0000000000100000-0x00000000bffeffff] usable
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x00000000bfff0000-0x00000000bfff2fff] ACPI NVS
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x00000000bfff3000-0x00000000bfffffff] ACPI data
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x00000000d0000000-0x00000000dfffffff] reserved
Nov 16 20:22:20 russell-server kernel: BIOS-e820: [mem 0x00000000fec00000-0x00000000ffffffff] reserved

had to hold down power button, no keyboard lights respond, no signal from monitor. help sad


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

#8 2013-11-17 00:09:55

cfr
Member
From: Cymru
Registered: 2011-11-27
Posts: 7,178

Re: random crash of PC (ata2 errors in journal)

I wouldn't trust sda or sdc. sdb looks OK. [Note: Could be cable or could be disks. Probably the disk for sdc since it can't complete the tests which, as far as I know, run autonomously once started? Not sure about this.]

Last edited by cfr (2013-11-17 00:11:30)


CLI Paste | How To Ask Questions

Arch Linux | x86_64 | GPT | EFI boot | refind | stub loader | systemd | LVM2 on LUKS
Lenovo x270 | Intel(R) Core(TM) i5-7200U CPU @ 2.50GHz | Intel Wireless 8265/8275 | US keyboard w/ Euro | 512G NVMe INTEL SSDPEKKF512G7L

Offline

#9 2013-11-17 05:07:30

x33a
Forum Fellow
Registered: 2009-08-15
Posts: 4,587

Re: random crash of PC (ata2 errors in journal)

Wow, your sda is the oldest working drive I have seen. 32800 hours!!

While the scans passed for sda, the logs show some errors just a few hours earlier. So it was most likely this drive giving the errors. Also, it is super old, so you might want to replace it on time.

As cfr said, sdb looks fine.

As for sdc, while it is old, and the scans failed in its case, you might want to try a new cable for that drive, and run the scans again to see what happens. If the scans still fail, it might be on its way out too.

Offline

#10 2013-11-17 12:56:28

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

Re: random crash of PC (ata2 errors in journal)

haha, I guess Ill replace sda then smile  any chance ata2 is referring to the sata port used on the motherborad? or is it ata version 2 or something?

Also, considering that my whole PC locks up, could I assume that its sda or sda's cable (which is root) causing the problem? (ie would any other drive cause a complete lockup?) I could try mounting everything else with noauto perhaps?

And thanks a lot for the help smile


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

#11 2013-11-17 23:05:03

cfr
Member
From: Cymru
Registered: 2011-11-27
Posts: 7,178

Re: random crash of PC (ata2 errors in journal)

Well there's a problem with both sda and sdc (whether the drive itself or the cable). The tests smart runs should complete without issue. I'm not certain how the tests work but I thought they pretty much ran independently once triggered and so wouldn't rely on the cable connection. But I could very well be wrong about that.


CLI Paste | How To Ask Questions

Arch Linux | x86_64 | GPT | EFI boot | refind | stub loader | systemd | LVM2 on LUKS
Lenovo x270 | Intel(R) Core(TM) i5-7200U CPU @ 2.50GHz | Intel Wireless 8265/8275 | US keyboard w/ Euro | 512G NVMe INTEL SSDPEKKF512G7L

Offline

#12 2013-11-18 08:57:05

x33a
Forum Fellow
Registered: 2009-08-15
Posts: 4,587

Re: random crash of PC (ata2 errors in journal)

jrussell wrote:

haha, I guess Ill replace sda then smile  any chance ata2 is referring to the sata port used on the motherborad? or is it ata version 2 or something?

Yeah, ata2 refers to the sata port.

Also, considering that my whole PC locks up, could I assume that its sda or sda's cable (which is root) causing the problem? (ie would any other drive cause a complete lockup?) I could try mounting everything else with noauto perhaps?

Since root was installed on sda, the lockups were probably due to this.

Also, do try changing the cables for sdc (or try changing the port), and run the test again on that drive. If it still fails, then you have a problem.

Offline

#13 2013-11-26 15:58:16

jrussell
Member
From: Cape Town, South Africa
Registered: 2012-08-16
Posts: 510

Re: random crash of PC (ata2 errors in journal)

Well I replaced the cable for sdb which was plugged into sata2 on the motherboard and haven't had any problems since. I think it was the cable then.

Thanks for the help


bitcoin: 1G62YGRFkMDwhGr5T5YGovfsxLx44eZo7U

Offline

Board footer

Powered by FluxBB