You are not logged in.
Pages: 1
Hi, lately my computer has been randomly crashing, usually while using discord, but I can't associate it to anything specific inside discord.
Everything on my screen freezes, I don't see any blinking lights on the keyboard and I also can't manually toggle the lights by pressing for example the caps lock key. After a click on the power button, I get the ability to toggle the lights on the keyboard with the caps lock and num lock keys. I can't switch to the tty. After some time the system powers off.
I checked journalctl and here is what I get when it crashes (I've included the lines showing that it is detecting the power button and powering off the system after pressing it)
Sep 01 21:18:45 hp kernel: NVRM: GPU at PCI:0000:01:00: GPU-00bdbf8f-292a-508e-c19c-a8ef6903638d
Sep 01 21:18:45 hp kernel: NVRM: GPU Board Serial Number:
Sep 01 21:18:45 hp kernel: NVRM: Xid (PCI:0000:01:00): 79, pid=0, GPU has fallen off the bus.
Sep 01 21:18:45 hp kernel: NVRM: GPU 0000:01:00.0: GPU has fallen off the bus.
Sep 01 21:18:45 hp kernel: NVRM: GPU 0000:01:00.0: GPU is on Board .
Sep 01 21:18:45 hp kernel: NVRM: A GPU crash dump has been created. If possible, please run
NVRM: nvidia-bug-report.sh as root to collect this data before
NVRM: the NVIDIA kernel module is unloaded.
Sep 01 21:19:30 hp systemd-logind[402]: Power key pressed.
Sep 01 21:19:30 hp systemd-logind[402]: Powering Off...
Sep 01 21:19:30 hp systemd-logind[402]: System is powering down.As the subject says, I'm using the proprietary nvidia driver.
I got nvidia-drm.modeset=1 on my kernel parameters, someone told me to add it because of another problem that I'm still having (please help: https://bbs.archlinux.org/viewtopic.php?id=258037 ), still it crashes with or without it.
I ran nvidia-bug-report.sh. Should I include the generated reports here or should I just send them to nvidia?
I was hoping someone cold help me solving this problem
EDIT: it just happened while using instagram on the surf browser. This time i got some extra logs on journalctl:
Sep 01 22:21:08 hp kernel: NVRM: GPU at PCI:0000:01:00: GPU-00bdbf8f-292a-508e-c19c-a8ef6903638d
Sep 01 22:21:08 hp kernel: NVRM: GPU Board Serial Number:
Sep 01 22:21:08 hp kernel: NVRM: Xid (PCI:0000:01:00): 79, pid=441, GPU has fallen off the bus.
Sep 01 22:21:08 hp kernel: NVRM: GPU 0000:01:00.0: GPU has fallen off the bus.
Sep 01 22:21:08 hp kernel: NVRM: GPU 0000:01:00.0: GPU is on Board .
Sep 01 22:21:08 hp kernel: NVRM: A GPU crash dump has been created. If possible, please run
NVRM: nvidia-bug-report.sh as root to collect this data before
NVRM: the NVIDIA kernel module is unloaded.
Sep 01 22:21:19 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:21:24 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:21:43 hp systemd-logind[397]: Power key pressed.
Sep 01 22:21:43 hp systemd-logind[397]: Powering Off...
Sep 01 22:21:43 hp systemd-logind[397]: System is powering down.
. . .
Sep 01 22:21:43 hp kernel: audit: type=1131 audit(1598995303.107:77): pid=1 uid=0 auid=4294967295 ses=4294967295 msg='unit=systemd-random-seed comm="systemd" exe="/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success'
Sep 01 22:21:43 hp mkinitcpio[9471]: ==> Starting build: none
Sep 01 22:21:43 hp mkinitcpio[9471]: -> Running build hook: [sd-shutdown]
Sep 01 22:21:43 hp kernel: snd_hda_codec_hdmi hdaudioC1D0: out of range cmd 0:5:707:ffffffbf
Sep 01 22:21:43 hp systemd[457]: pulseaudio.service: Succeeded.
Sep 01 22:21:43 hp mkinitcpio[9471]: ==> Build complete.
Sep 01 22:21:43 hp systemd[1]: mkinitcpio-generate-shutdown-ramfs.service: Succeeded.
Sep 01 22:21:43 hp systemd[1]: Finished Generate shutdown-ramfs.
Sep 01 22:21:43 hp audit[1]: SERVICE_START pid=1 uid=0 auid=4294967295 ses=4294967295 msg='unit=mkinitcpio-generate-shutdown-ramfs comm="systemd" exe="/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success'
Sep 01 22:21:43 hp audit[1]: SERVICE_STOP pid=1 uid=0 auid=4294967295 ses=4294967295 msg='unit=mkinitcpio-generate-shutdown-ramfs comm="systemd" exe="/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success'
Sep 01 22:21:43 hp kernel: audit: type=1130 audit(1598995303.203:78): pid=1 uid=0 auid=4294967295 ses=4294967295 msg='unit=mkinitcpio-generate-shutdown-ramfs comm="systemd" exe="/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success'
Sep 01 22:21:43 hp kernel: audit: type=1131 audit(1598995303.203:79): pid=1 uid=0 auid=4294967295 ses=4294967295 msg='unit=mkinitcpio-generate-shutdown-ramfs comm="systemd" exe="/usr/lib/systemd/systemd" hostname=? addr=? terminal=? res=success'
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Stopping timed out. Killing.
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000987d:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:2:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:2:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:3:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:3:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000987d:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000917e:2:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:2:0:0x0000000f
. . .
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:2:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:3:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed detecting connected display devices
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed detecting connected display devices
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed detecting connected display devices
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:0:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:1:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:2:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000927c:3:0:0x0000000f
Sep 01 22:23:13 hp kernel: nvidia-modeset: ERROR: GPU:0: Failed to query display engine channel state: 0x0000987d:0:0:0x0000000f
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 469 (startx) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 492 (xinit) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 493 (Xorg) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 720 (brave) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 741 (Chrome_IOThread) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Killing process 9468 (n/a) with signal SIGKILL.
Sep 01 22:23:13 hp systemd[1]: session-1.scope: Failed with result 'timeout'.
Sep 01 22:23:13 hp systemd[1]: Stopped Session 1 of user jl.I can add a pastebin link to the full log it you want me to
Last edited by jlucas (2020-09-03 20:01:09)
Offline
Omit "nvidia-drm.modeset=1 ", wait for the screen to black out and dump the system journal (sudo journalct -b)
You can also try "rcutree.rcu_idle_gp_delay=1" if the GPU keeps fallen off the bus.
Offline
Omit "nvidia-drm.modeset=1 ", wait for the screen to black out and dump the system journal (sudo journalct -b)
You can also try "rcutree.rcu_idle_gp_delay=1" if the GPU keeps fallen off the bus.
Assuming that by "wait for the screen to black out" you mean "wait for the system to power off after it crashes", this is what I get:
Sep 01 23:55:39 hp kernel: ------------[ cut here ]------------
Sep 01 23:55:39 hp kernel: WARNING: CPU: 2 PID: 486 at /build/nvidia/src/nvidia/450.66/build/nvidia/nv.c:4630 nvidia_dev_put+0x92/0xa0 [nvidia]
Sep 01 23:55:39 hp kernel: Modules linked in: nvidia_drm(POE) nvidia_modeset(POE) snd_sof_pci snd_sof_intel_byt snd_sof_intel_ipc snd_sof_intel_hda_common snd_soc_hdac_hda snd_sof_xtensa_dsp intel_rapl_msr snd_sof_intel_hda intel_rapl_common snd_sof nvidia(POE) snd_soc_skl x86_pkg_temp_thermal intel_powerclamp coretemp snd_soc_sst_ipc snd_soc_sst_dsp iTCO_wdt snd_hda_ext_core intel_pmc_bxt iTCO_vendor_support ee1004 mei_hdcp snd_soc_acpi_intel_match 8250_dw kvm_intel snd_soc_acpi kvm snd_soc_core hp_wmi wmi_bmof sparse_keymap snd_hda_codec_realtek irqbypass snd_compress snd_hda_codec_generic crct10dif_pclmul snd_hda_codec_hdmi crc32_pclmul ac97_bus ledtrig_audio snd_pcm_dmaengine ghash_clmulni_intel snd_hda_intel aesni_intel crypto_simd ofpart cryptd snd_intel_dspcfg 8821ce(OE) cmdlinepart glue_helper snd_hda_codec rapl nls_iso8859_1 intel_spi_pci intel_spi intel_cstate nls_cp437 btusb snd_hda_core spi_nor intel_uncore r8169 vfat btrtl btbcm
fat pcspkr mtd realtek i2c_i801 i2c_smbus cfg80211 libphy
Sep 01 23:55:39 hp kernel: snd_hwdep btintel snd_pcm mei_me intel_lpss_pci intel_lpss bluetooth mei idma64 snd_timer joydev mousedev ecdh_generic snd input_leds rfkill ecc soundcore intel_pch_thermal ie31200_edac wmi tpm_crb tpm_tis tpm_tis_core tpm rng_core evdev mac_hid vboxnetflt(OE) vboxnetadp(OE) vboxdrv(OE) vboxvideo drm_vram_helper drm_ttm_helper ttm drm_kms_helper cec rc_core drm syscopyarea sysfillrect sysimgblt fb_sys_fops vboxsf vboxguest agpgart ip_tables x_tables ext4 crc32c_generic crc16 mbcache hid_generic jbd2 usbhid hid crc32c_intel sdhci_pci xhci_pci cqhci xhci_pci_renesas sdhci sr_mod xhci_hcd cdrom mmc_core
Sep 01 23:55:39 hp kernel: CPU: 2 PID: 486 Comm: Xorg Tainted: P OE 5.8.5-arch1-1 #1
Sep 01 23:55:39 hp kernel: Hardware name: HP HP Pavilion Gaming Desktop 690-00xx/843B, BIOS F.42 05/28/2020
Sep 01 23:55:39 hp kernel: RIP: 0010:nvidia_dev_put+0x92/0xa0 [nvidia]
Sep 01 23:55:39 hp kernel: Code: e7 e8 c2 7e 7c 00 85 c0 75 20 5b 4c 89 ef 5d 41 5c 41 5d e9 80 70 3d c8 5b 48 c7 c7 b0 e4 04 c2 5d 41 5c 41 5d e9 6e 70 3d c8 <0f> 0b eb dc 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 41 57 41
Sep 01 23:55:39 hp kernel: RSP: 0000:ffffaa180085bba8 EFLAGS: 00010202
Sep 01 23:55:39 hp kernel: RAX: 0000000000000026 RBX: 0000000000000100 RCX: 0000000000000000
Sep 01 23:55:39 hp kernel: RDX: 0000000000000087 RSI: 0000000000000246 RDI: 00000000ffffffff
Sep 01 23:55:39 hp kernel: RBP: ffff9446dd351800 R08: 0000000000000000 R09: ffff9446d2383000
Sep 01 23:55:39 hp kernel: R10: ffff9446d2648880 R11: ffff9446d2385e00 R12: ffff9446d2383000
Sep 01 23:55:39 hp kernel: R13: ffff9446dd351ce0 R14: ffff9446d8ef1c00 R15: ffff9446e4d75ba0
Sep 01 23:55:39 hp kernel: FS: 0000000000000000(0000) GS:ffff9446ebe80000(0000) knlGS:0000000000000000
Sep 01 23:55:39 hp kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
Sep 01 23:55:39 hp kernel: CR2: 00001d75f1678008 CR3: 0000000026e0a006 CR4: 00000000003606e0
Sep 01 23:55:39 hp kernel: Call Trace:
Sep 01 23:55:39 hp kernel: nvkms_close_gpu+0x49/0x80 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: _nv002731kms+0x11d/0x170 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: ? _nv002544kms+0x8c/0xd0 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: ? _nv000340kms+0x86/0xb0 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: ? nvKmsClose+0xab/0x170 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: ? nvkms_close+0x42/0xd0 [nvidia_modeset]
Sep 01 23:55:39 hp kernel: ? nvidia_frontend_close+0x2b/0x50 [nvidia]
Sep 01 23:55:39 hp kernel: ? __fput+0xca/0x230
Sep 01 23:55:39 hp kernel: ? task_work_run+0x5c/0x90
Sep 01 23:55:39 hp kernel: ? do_exit+0x369/0xab0
Sep 01 23:55:39 hp kernel: ? perf_event_task_tick+0x67/0x3c0
Sep 01 23:55:39 hp kernel: ? reweight_entity+0x2e/0x130
Sep 01 23:55:39 hp kernel: ? do_group_exit+0x33/0xa0
Sep 01 23:55:39 hp kernel: ? get_signal+0x148/0x900
Sep 01 23:55:39 hp kernel: ? timerqueue_add+0x96/0xb0
Sep 01 23:55:39 hp kernel: ? enqueue_hrtimer+0x39/0xa0
Sep 01 23:55:39 hp kernel: ? do_signal+0x3d/0x730
Sep 01 23:55:39 hp kernel: ? recalibrate_cpu_khz+0x10/0x10
Sep 01 23:55:39 hp kernel: ? ktime_get+0x38/0xa0
Sep 01 23:55:39 hp kernel: ? __prepare_exit_to_usermode+0x112/0x1c0
Sep 01 23:55:39 hp kernel: ? asm_sysvec_reschedule_ipi+0xa/0x20
Sep 01 23:55:39 hp kernel: ? prepare_exit_to_usermode+0x5/0x20
Sep 01 23:55:39 hp kernel: ? asm_sysvec_reschedule_ipi+0x12/0x20
Sep 01 23:55:39 hp kernel: ---[ end trace 60a4162e29991004 ]---The rest seems to be the same.
Offline
Omit "nvidia-drm.modeset=1 ", wait for the screen to black out and dump the system journal (sudo journalct -b)
You can also try "rcutree.rcu_idle_gp_delay=1" if the GPU keeps fallen off the bus.
It still crashes with "rcutree.rcu_idle_gp_delay=1", the logs seem to be the same, except
Sep 02 00:13:31 hp kernel: snd_hda_codec_hdmi hdaudioC1D0: out of range cmd 0:5:707:ffffffbfshows before I pressed the power button.
Offline
I meant when this happens (from your other thread):
Occasionally, my screen goes black for about 5 seconds, and most of the times I have no idea what's causing it.
Then post the entire system journal.
The backtrace smells like a stuck CPU (which apparently makes nvidia_modeset wet itself) which could have a different cause.
Offline
I meant when this happens (from your other thread):
Occasionally, my screen goes black for about 5 seconds, and most of the times I have no idea what's causing it.
Then post the entire system journal.
The backtrace smells like a stuck CPU (which apparently makes nvidia_modeset wet itself) which could have a different cause.
Oh ok, I just ran "sudo journalctl -f" and switched to the tty, then back to the X session, since that's the only way I can make that happen intentionally. It didn't show anything interesting on journalctl. A few seconds after that, it crashed.
I just got a black screen after opening my browser, so I will post "sudo journalctl -b".
I don't think we should be talking about this here, so I'm posting it on the other topic.
EDIT: Also I don't believe these two problems are related, it's been blacking out since I can remember and it only started crashing a few days ago.
Last edited by jlucas (2020-09-02 12:40:44)
Offline
Hey, I did a full system upgrade and it doesn't seem to be crashing
(I did upgrade the system before starting this topic)
/var/log/pacman.log:
[2020-09-02T13:23:19+0100] [PACMAN] starting full system upgrade
[2020-09-02T13:23:40+0100] [ALPM] transaction started
[2020-09-02T13:23:40+0100] [ALPM] upgraded linux-api-headers (5.7-1 -> 5.8-1)
[2020-09-02T13:23:40+0100] [ALPM] upgraded glibc (2.32-3 -> 2.32-4)
[2020-09-02T13:23:40+0100] [ALPM-SCRIPTLET] Generating locales...
[2020-09-02T13:23:42+0100] [ALPM-SCRIPTLET] en_US.UTF-8... done
[2020-09-02T13:23:43+0100] [ALPM-SCRIPTLET] pt_PT.UTF-8... done
[2020-09-02T13:23:43+0100] [ALPM-SCRIPTLET] Generation complete.
[2020-09-02T13:23:43+0100] [ALPM] upgraded gcc-libs (10.2.0-1 -> 10.2.0-2)
[2020-09-02T13:23:43+0100] [ALPM] upgraded binutils (2.35-1 -> 2.35-2)
[2020-09-02T13:23:43+0100] [ALPM] upgraded gcc (10.2.0-1 -> 10.2.0-2)
[2020-09-02T13:23:43+0100] [ALPM] upgraded lib32-glibc (2.32-3 -> 2.32-4)
[2020-09-02T13:23:44+0100] [ALPM] upgraded lib32-gcc-libs (10.2.0-1 -> 10.2.0-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded php (7.4.9-2 -> 7.4.10-1)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-common (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-devel (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-xephyr (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-xnest (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-xvfb (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] upgraded xorg-server-xwayland (1.20.9-1 -> 1.20.9-2)
[2020-09-02T13:23:44+0100] [ALPM] transaction completedI still have the other screen going black problem ![]()
I guess I'm marking this one as solved tomorrow if it doesn't crash
Offline
I'm thinking this could actually be hardware related. I accidentally kicked my desk and the screen froze. Then I remembered I was getting the same behavior while I was running windows. I checked journalctl and the logs were the same as the other crashes. It's the first crash since my last post, so if it's partially software related, I guess the software side is fixed.
Offline
Pages: 1