You are not logged in.
Pages: 1
Hello,
I have reinstalled my arch about a month ago . Since then, I am getting really serious crashes. Whole system just becomes unresponsive and the only possibility is to hard reboot it. From the logs I have found that the my initial problem is similar to this one.
I thought it might be problem with a kernel so I switched to LTS but the system still crashes from time to time. However, reason for this seems to be different (I will include two journalctl logs from before two different crashes). Now it seems like there is a problem with pulseaudio but I have not been able to google a solution.
My laptop is ThinkPad T480s with Intel GPU.
Two journalctl -b-1 logs that were too large for pastebin (sorry).
I am not sure if I did something wrong or it is a bug. I would be really grateful for any suggestions.
Offline
From log.txt
Dec 22 22:14:14 arch kernel: oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/user.slice/user-1000.slice/session-2.scope,task=Popcorn-Time,pid=4041,uid=1000
....
Dec 22 22:15:36 arch kernel: sysrq: This sysrq operation is disabled.System ran out of memory and sysrq did not work as disabled.
log-new.txt just ends.
Offline
From log.txt
Dec 22 22:14:14 arch kernel: oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/user.slice/user-1000.slice/session-2.scope,task=Popcorn-Time,pid=4041,uid=1000 .... Dec 22 22:15:36 arch kernel: sysrq: This sysrq operation is disabled.System ran out of memory and sysrq did not work as disabled.
log-new.txt just ends.
I do have 16 GB of RAM on my computer. I noticed the killing of processes inside logs so I tried to creating SWAP to see if it could help before the second crash. How can I prevent it otherwise? Or is it some memory leak inside the mentioned app?
Also, what do you mean by 'just ends'?
Thanks a lot for your answer btw!
Offline
The OOM killer does not necessarily pick the offending process, but apparently something™ eats up your memory, so monitor that resource and whether any process eats up large amounts of memory.
16GB fade very fast if eg. a videoplayer leaks decoded frames (142MB/s for FullHD @24Hz… - that's less than 2 minutes…)
Offline
The OOM killer does not necessarily pick the offending process, but apparently something™ eats up your memory, so monitor that resource and whether any process eats up large amounts of memory.
16GB fade very fast if eg. a videoplayer leaks decoded frames (142MB/s for FullHD @24Hz… - that's less than 2 minutes…)
Thanks, that sounds reasonable. I was running Android Emulator and Android Studio, I suppose they take a lot of RAM. I have system monitor opened and even during most resources hungry tasks, RAM usually stays below 10 GB. Is there some way to lets say kill the apps more aggressively? I would like to prevent my whole system to crash. Would it make difference to have 24 GB RAM instead of 16 GB? Or would it just let the app leak resources longer?
EDIT: I would like to note one thing though. I am using Arch for about a year now, for more or less similar use cases and I had no crashes for majority of this time.
Last edited by Recursed (2020-01-01 20:16:20)
Offline
If there's a leak, you'd still run OOM, just later.
"system monitor" is some GUI indicator displaying some™ aggregated number?
SWAP (file or partition) will allow you to stall the condition. The system might become very slow, but you stand a chance to investigate the issue.
You can run stuff in RAM quota'd cgroups, https://wiki.archlinux.org/index.php/Cgroups - but that requires you to know the offending process.
You can disable overcommitment, but I don't see how this would help us here, because it would just lead to random brokeness in whatever processes try to allocate memory after some hog has eaten it all.
I will say that the most memory consuming process has the biggest chance of being killed, so popcorn-time (bittorrent videoplayer, right? that should™ not allocate that vast amounts of RAM) is actually a good candidate of causing this. You could try and see whether this ever happens while it's not running.
Offline
If there's a leak, you'd still run OOM, just later.
"system monitor" is some GUI indicator displaying some™ aggregated number?SWAP (file or partition) will allow you to stall the condition. The system might become very slow, but you stand a chance to investigate the issue.
You can run stuff in RAM quota'd cgroups, https://wiki.archlinux.org/index.php/Cgroups - but that requires you to know the offending process.You can disable overcommitment, but I don't see how this would help us here, because it would just lead to random brokeness in whatever processes try to allocate memory after some hog has eaten it all.
I will say that the most memory consuming process has the biggest chance of being killed, so popcorn-time (bittorrent videoplayer, right? that should™ not allocate that vast amounts of RAM) is actually a good candidate of causing this. You could try and see whether this ever happens while it's not running.
Yes, by system monitor I mean the built-in one in KDE. I try to check htop too because it seems to be better for some tasks. Usually, the most resource heavy is android emulator or IDE. You might be right that it is caused by popcorntime, I will check it. If it does this leak though, it has to do it really slowly.
I did enable swap file recently to see if it works. In crash that happened today, It got partially responsive after short freeze which allowed me to open tty, but it was super slow and after a while froze completely.
Thanks for suggesting cgroups, I will try it right now.
Offline
If there's a leak, you'd still run OOM, just later.
"system monitor" is some GUI indicator displaying some™ aggregated number?SWAP (file or partition) will allow you to stall the condition. The system might become very slow, but you stand a chance to investigate the issue.
You can run stuff in RAM quota'd cgroups, https://wiki.archlinux.org/index.php/Cgroups - but that requires you to know the offending process.You can disable overcommitment, but I don't see how this would help us here, because it would just lead to random brokeness in whatever processes try to allocate memory after some hog has eaten it all.
I will say that the most memory consuming process has the biggest chance of being killed, so popcorn-time (bittorrent videoplayer, right? that should™ not allocate that vast amounts of RAM) is actually a good candidate of causing this. You could try and see whether this ever happens while it's not running.
My computer crashed again today when I plugged it into my docking station. Errors from logs seems slighly different this time. Link to logs. I do not see any oom-kill inside the logs, does that mean nothing was killed?
Could this be possible caused by hardware issue?
Thanks for answer!
Offline
Jan 02 22:20:45 arch kernel: [drm:intel_dp_start_link_train [i915]] *ERROR* failed to enable link training
Jan 02 22:20:45 arch kernel: [drm:intel_encoders_pre_enable.isra.0 [i915]] *ERROR* failed to allocate vcpi
Jan 02 22:20:45 arch kernel: [drm:intel_mst_enable_dp [i915]] *ERROR* Timed out waiting for ACT sent
Jan 02 22:20:45 arch kernel: usb 1-9: reset full-speed USB device number 7 using xhci_hcd
Jan 02 22:20:45 arch kernel: [drm:drm_atomic_helper_wait_for_flip_done [drm_kms_helper]] *ERROR* [CRTC:41:pipe A] flip_done timed out
Jan 02 22:20:45 arch kernel: [drm:pipe_config_err [i915]] *ERROR* mismatch in dp_m_n (expected tu 0 gmch 2813678/8388608 link 468946/1048576, or tu 0 gmch 0/0 link 0/0, found tu 64, gmch 2813678/8388608 link 468946/1048576)
Jan 02 22:20:45 arch kernel: ------------[ cut here ]------------
Jan 02 22:20:45 arch kernel: pipe state doesn't match!
Jan 02 22:20:45 arch kernel: WARNING: CPU: 0 PID: 246702 at drivers/gpu/drm/i915/intel_display.c:11909 intel_atomic_commit_tail+0xd15/0xd90 [i915]
Jan 02 22:20:45 arch kernel: Modules linked in: hid_logitech_hidpp snd_usb_audio snd_usbmidi_lib snd_rawmidi hid_logitech_dj snd_seq_device hid_generic usbhid hid ccm msr lz4 lz4_compress joydev mousedev elan_i2c iTCO_wdt iTCO_vendor_support mei_wdt arc4 wmi_bmof intel_rapl intel_wmi_thunderbolt iwlmvm snd_hda_codec_hdmi mac80211 snd_soc_skl snd_soc_skl_ipc snd_soc_sst_ipc snd_soc_sst_dsp snd_hda_ext_core snd_soc_acpi_intel_match snd_soc_acpi snd_soc_core nls_iso8859_1 snd_hda_codec_realtek nls_cp437 snd_compress vfat x86_pkg_temp_thermal ac97_bus fat snd_hda_codec_generic iwlwifi snd_pcm_dmaengine intel_powerclamp snd_hda_intel coretemp snd_hda_codec kvm_intel snd_hda_core intel_cstate intel_uncore snd_hwdep intel_rapl_perf snd_pcm pcspkr psmouse input_leds cfg80211 e1000e snd_timer i2c_i801 btusb btrtl btbcm thunderbolt
Jan 02 22:20:45 arch kernel: btintel uvcvideo bluetooth videobuf2_vmalloc videobuf2_memops videobuf2_v4l2 videobuf2_common cdc_mbim videodev cdc_wdm cdc_ncm mei_me usbnet pcc_cpufreq ecdh_generic cdc_acm media mii mei intel_pch_thermal processor_thermal_device intel_xhci_usb_role_switch intel_soc_dts_iosf roles ucsi_acpi typec_ucsi typec wmi thinkpad_acpi tpm_crb nvram rfkill snd soundcore ac int3403_thermal battery int340x_thermal_zone evdev int3400_thermal tpm_tis mac_hid acpi_thermal_rel tpm_tis_core tpm rng_core crypto_user ip_tables x_tables ext4 crc32c_generic crc16 mbcache jbd2 sd_mod uas usb_storage scsi_mod dm_crypt dm_mod crct10dif_pclmul crc32_pclmul crc32c_intel ghash_clmulni_intel pcbc serio_raw atkbd libps2 aesni_intel aes_x86_64 xhci_pci crypto_simd cryptd glue_helper xhci_hcd i8042 serio i915 kvmgt
Jan 02 22:20:45 arch kernel: vfio_mdev mdev vfio_iommu_type1 vfio kvm irqbypass i2c_algo_bit drm_kms_helper syscopyarea sysfillrect sysimgblt fb_sys_fops drm intel_agp intel_gtt agpgart
Jan 02 22:20:45 arch kernel: CPU: 0 PID: 246702 Comm: kworker/u16:72 Not tainted 4.19.91-1-lts #1
Jan 02 22:20:45 arch kernel: Hardware name: LENOVO 20L7S02Y00/20L7S02Y00, BIOS N22ET59W (1.36 ) 10/18/2019
Jan 02 22:20:45 arch kernel: Workqueue: events_unbound async_run_entry_fn
Jan 02 22:20:45 arch kernel: RIP: 0010:intel_atomic_commit_tail+0xd15/0xd90 [i915]
Jan 02 22:20:45 arch kernel: Code: 13 00 00 8d 71 41 48 c7 c7 68 30 4a c0 75 6d e8 61 d0 d6 ff e9 34 fb ff ff e8 36 f2 e9 dd 0f 0b e9 52 fd ff ff e8 2a f2 e9 dd <0f> 0b e9 fc f9 ff ff e8 1e f2 e9 dd 0f 0b 0f b6 44 24 18 e9 2a f9
Jan 02 22:20:45 arch kernel: RSP: 0018:ffffa364caf6fc68 EFLAGS: 00010286
Jan 02 22:20:45 arch kernel: RAX: 0000000000000000 RBX: ffff95f55b810000 RCX: 0000000000000006
Jan 02 22:20:45 arch kernel: RDX: 0000000000000007 RSI: 0000000000000096 RDI: ffff95f55e2165b0
Jan 02 22:20:45 arch kernel: RBP: ffff95f5261d7000 R08: 0000174c6095ea95 R09: ffffffff9fa7baf4
Jan 02 22:20:45 arch kernel: R10: 00000000000009d1 R11: 0000000000006f3c R12: ffff95f5261d3800
Jan 02 22:20:45 arch kernel: R13: ffff95f5261d7800 R14: ffff95f555300000 R15: ffff95f555300368
Jan 02 22:20:45 arch kernel: FS: 0000000000000000(0000) GS:ffff95f55e200000(0000) knlGS:0000000000000000
Jan 02 22:20:45 arch kernel: CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
Jan 02 22:20:45 arch kernel: CR2: 0000309837493000 CR3: 00000001d8a0a003 CR4: 00000000003626f0
Jan 02 22:20:45 arch kernel: Call Trace:
Jan 02 22:20:45 arch kernel: intel_atomic_commit+0x2b3/0x2f0 [i915]
Jan 02 22:20:45 arch kernel: drm_atomic_helper_commit_duplicated_state+0xc9/0xe0 [drm_kms_helper]
Jan 02 22:20:45 arch kernel: __intel_display_resume+0x84/0xd0 [i915]
Jan 02 22:20:45 arch kernel: intel_display_resume+0xd1/0x120 [i915]
Jan 02 22:20:45 arch kernel: ? drm_dp_mst_topology_mgr_resume+0xb7/0x110 [drm_kms_helper]
Jan 02 22:20:45 arch kernel: ? pci_pm_suspend_late+0x30/0x30
Jan 02 22:20:45 arch kernel: i915_drm_resume+0xc2/0x110 [i915]
Jan 02 22:20:45 arch kernel: dpm_run_callback+0x59/0x150
Jan 02 22:20:45 arch kernel: device_resume+0xac/0x200
Jan 02 22:20:45 arch kernel: async_resume+0x19/0x30
Jan 02 22:20:45 arch kernel: async_run_entry_fn+0x37/0x140
Jan 02 22:20:45 arch kernel: process_one_work+0x1da/0x3b0
Jan 02 22:20:45 arch kernel: worker_thread+0x4d/0x3f0
Jan 02 22:20:45 arch kernel: kthread+0xfb/0x130
Jan 02 22:20:45 arch kernel: ? process_one_work+0x3b0/0x3b0
Jan 02 22:20:45 arch kernel: ? kthread_park+0x80/0x80
Jan 02 22:20:45 arch kernel: ret_from_fork+0x35/0x40
…
…
…
Jan 02 22:21:16 arch kernel: [drm:drm_atomic_helper_wait_for_dependencies [drm_kms_helper]] *ERROR* [CRTC:41:pipe A] flip_done timed out
…
Jan 02 22:21:26 arch kernel: [drm:drm_atomic_helper_wait_for_flip_done [drm_kms_helper]] *ERROR* [CRTC:41:pipe A] flip_done timed out
…
Jan 02 22:21:36 arch kernel: [drm:drm_atomic_helper_wait_for_dependencies [drm_kms_helper]] *ERROR* [CRTC:41:pipe A] flip_done timed out
…
Jan 02 22:21:47 arch kernel: [drm:drm_atomic_helper_wait_for_dependencies [drm_kms_helper]] *ERROR* [PLANE:38:cursor A] flip_done timed out
Jan 02 22:21:54 arch systemd-logind[629]: Lid opened.
…
Jan 02 22:21:58 arch kernel: [drm:drm_atomic_helper_wait_for_dependencies [drm_kms_helper]] *ERROR* [PLANE:28:plane 1A] flip_done timed out
…
Jan 02 22:22:08 arch kernel: [drm:drm_atomic_helper_wait_for_flip_done [drm_kms_helper]] *ERROR* [CRTC:41:pipe A] flip_done timed out
Jan 02 22:22:09 arch systemd-logind[629]: Power key pressed.
Jan 02 22:22:09 arch systemd-logind[629]: Powering Off...Does "crash" mean "my display didn't work"? When the system "crashes", can you still ssh into it (oc. make sure sshd runs)?
I don't think this is the same or even just a related problem - can you reproduce the issue w/ the dock? Perhaps only after a S3 cycle?
Offline
Does "crash" mean "my display didn't work"? When the system "crashes", can you still ssh into it (oc. make sure sshd runs)?
I don't think this is the same or even just a related problem - can you reproduce the issue w/ the dock? Perhaps only after a S3 cycle?
Well the symptoms were almost the same. The system was not responsive and fans were maxed and loud (I suppose if only monitor was not working, it should be quiet?). I could not use Ctrl Alt F2 to do something. But yes, this time the display did not work.
I don't know how to really reproduce this. Sometimes it crashes after serveral days of uptime, sometime twice a day. Before I plugged the computer to my docking station I was doing some heavy lifting and everything was fine. I am just really confused about this, sorry for that.
From where should I ssh into it?
Offline
From where should I ssh into it?
Any available other machine in the subnet (a smartphone might do)
Offline
Pages: 1