You are not logged in.
Pages: 1
This recently started happening when trying to play GPU heavy games. The monitors go blank, some sounds still play but I think they are in a loop. I have to power cycle to get any display.
I fear this is a hardware problem, but would like another opinion. Anybody have suggestions?
Full journal: https://paste.c-net.org/BearingsSunday
Probably the relevant bit:
Jul 15 21:37:01 cb-d2022 kernel: pcieport 0000:40:03.1: DPC: containment event, status:0x1f01: unmasked uncorrectable error detected
Jul 15 21:37:01 cb-d2022 kernel: pcieport 0000:40:03.1: PCIe Bus Error: severity=Uncorrectable (Fatal), type=Transaction Layer, (Receiver ID)
Jul 15 21:37:01 cb-d2022 kernel: pcieport 0000:40:03.1: device [1022:1453] error status/mask=00040000/04400000
Jul 15 21:37:01 cb-d2022 kernel: pcieport 0000:40:03.1: [18] MalfTLP (First)
Jul 15 21:37:01 cb-d2022 kernel: pcieport 0000:40:03.1: AER: TLP Header: 0x1f000001 0x00000000 0x00000000 0x599a960d
Jul 15 21:37:01 cb-d2022 kernel: snd_hda_intel 0000:43:00.1: Unable to change power state from D3hot to D0, device inaccessible
Jul 15 21:37:01 cb-d2022 kernel: snd_hda_intel 0000:43:00.1: CORB reset timeout#2, CORBRP = 65535
Jul 15 21:37:01 cb-d2022 kernel: amdgpu 0000:43:00.0: PCI error: detected callback!!
Jul 15 21:37:01 cb-d2022 kernel: amdgpu 0000:43:00.0: pci_channel_io_frozen: state(2)!!
Jul 15 21:37:01 cb-d2022 kernel: amdgpu 0000:43:00.0: device lost from bus!
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State Completed
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: [drm] AMDGPU device coredump file has been created
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: ring comp_1.3.1 timeout, signaled seq=38711, emitted seq=38712
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Process GameThread pid 4636 thread dxvk-submit pid 4723
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Starting comp_1.3.1 ring reset
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Ring comp_1.3.1 reset failed
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU reset begin!. Source: 1
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: device lost from bus!
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU reset end with ret = -19
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU Recovery Failed: -19
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State Completed
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: ring gfx_0.0.0 timeout, signaled seq=663167, emitted seq=663169
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Process GameThread pid 4636 thread dxvk-submit pid 4723
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Starting gfx_0.0.0 ring reset
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Ring gfx_0.0.0 reset failed
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU reset begin!. Source: 1
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: device lost from bus!
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU reset end with ret = -19
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: GPU Recovery Failed: -19
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: Dumping IP State Completed
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: [drm] AMDGPU device coredump file has been created
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
Jul 15 21:37:03 cb-d2022 kernel: amdgpu 0000:43:00.0: ring sdma2 timeout, signaled seq=958, emitted seq=960For what it's worth "/sys/class/drm/card1/device/devcoredump" does not exist, at least after reboot.
"lspci -vvv": https://paste.c-net.org/FixedCaptive
"lspci -nnk": https://paste.c-net.org/FailuresMango
from "lspci -nnk":
40:03.1 PCI bridge [0604]: Advanced Micro Devices, Inc. [AMD] Family 17h (Models 00h-0fh) PCIe GPP Bridge [1022:1453]
Subsystem: Advanced Micro Devices, Inc. [AMD] Family 17h (Models 00h-0fh) PCIe GPP Bridge [1022:1453]
Kernel driver in use: pcieport
Kernel modules: shpchp"inxi --basic":
System:
Host: cb-d2022 Kernel: 7.1.3-arch1-3 arch: x86_64 bits: 64
Desktop: Hyprland v: 0.55.4 Distro: Arch Linux
Machine:
Type: Desktop Mobo: ASRock model: X399 Taichi serial: <superuser required>
Firmware: UEFI vendor: American Megatrends v: P3.92 date: 01/14/2021
CPU:
Info: 16-core AMD Ryzen Threadripper 1950X [MT MCP] speed (MHz): avg: 2200
min/max: 2200/4200
Graphics:
Device-1: Advanced Micro Devices [AMD/ATI] Navi 21 [Radeon RX 6800/6800 XT
/ 6900 XT] driver: amdgpu v: kernel
Display: wayland server: X.Org v: 24.1.13 with: Xwayland v: 24.1.13
compositor: Hyprland v: 0.55.4 driver: X: loaded: amdgpu dri: radeonsi
gpu: amdgpu resolution: 1: 2560x1080~120Hz 2: 1920x1080~60Hz
API: OpenGL v: 4.6 vendor: amd mesa v: 26.1.4-arch1.1 renderer: AMD
Radeon RX 6900 XT (radeonsi navi21 ACO DRM 3.64 7.1.3-arch1-3)
Info: Tools: api: eglinfo, glxinfo, vulkaninfo gpu: amdgpu_top, corectrl,
radeontop x11: xdpyinfo, xprop, xrandr
Network:
Device-1: Intel I211 Gigabit Network driver: igb
Device-2: Intel Dual Band Wireless-AC 3168NGW [Stone Peak] driver: iwlwifi
Device-3: Intel I211 Gigabit Network driver: igb
Drives:
Local Storage: total: 2.29 TiB used: 954.33 GiB (40.8%)
Info:
Memory: total: 64 GiB available: 62.67 GiB used: 9.74 GiB (15.5%)
Processes: 562 Uptime: 3d 11h 52m Shell: Zsh inxi: 3.3.41Please let me know if I should provide more information.
Offline
Does this happen w/ the LTS kernel?
Have you tried to re-seat the GPU?
Is this even the correct (a PEG) slot (the bus ID seems strange)?
Did you forget to attach the dedicated power supply or did it maybe come loose?
Online
My apologies I left out that information.
I did blow out the dust & re-seat the GPU.
Both 8 pin power cables are firmly seated in GPU & PSU
The GPU is seated in the closest PCI slot to the CPU
The crash does occur w/ the LTS kernel
Offline
Ewww… do you have a spare GPU (to check whether this is coming from the device or the bus)?
Online
I do have a GTX 970, though that PC does not post currently. I have no idea if it works or not. /fml
Offline
Pages: 1