You are not logged in.
Hello,
I have been experiencing a problem since I recently built a new PC. About once a day, seemingly at random, when alt-tabbing or changing window focus (in XFCE), the display will completely lock up and not recover. At first, I thought this might be an issue from full-screen applications like games, but it has also happened when I am not running any such applications, just normal desktop applications like Firefox, Discord, and Spotify. Normally, in Linux, it's possible to press ctrl+alt+f2 to go to a different login prompt, but I can't do that when the screen locks up. I had been using Arch on my old PC fine. I did a complete fresh install of Arch for my new PC. The relevant specs of the PC are:
DE: XFCE
Processor: AMD Ryzen 5 3600X
Video Card: Sapphire Pulse RX 5700 XT
Monitors:
-1440x900 running at 75 Hz
-1920x1080 running at 144 Hz
-1920x1080 running at 60 Hz
So far, I have tried switching to mesa-git packages, but the lockups were still present when using those. I have also tried unplugging both of the supplementary monitors and just running on my main monitor, the 1920x1080@144 monitor, but the lockups have occurred when only that monitor is plugged in as well.
The last time this happened, I took a look at the journalctl log, and this seems to be the relevant log entry:
Dec 08 02:50:30 simon-dt2019-arch kernel: RIP: 0010:CalculateVMAndRowBytes.constprop.0+0x211/0x910 [amdgpu]
Dec 08 02:50:30 simon-dt2019-arch kernel: Hardware name: To Be Filled By O.E.M. To Be Filled By O.E.M./X570 Phantom Gaming 4, BIOS P1.70 09/10/2019
Dec 08 02:50:30 simon-dt2019-arch kernel: CPU: 9 PID: 580 Comm: Xorg Tainted: G W 5.4.2-arch1-1 #1
Dec 08 02:50:30 simon-dt2019-arch kernel: divide error: 0000 [#1] PREEMPT SMP NOPTII would post the Xorg log, but it seems to have been overwritten by a more recent entry. If there are any other relevant logs or info I should be posting, please let me know.
I know support for the RX 5700 XT is still in its early stages, but I haven't seen any other posts recently like this from Duck Duck Go searches, so I am worried that it might be a hardware issue. I have a Windows dual boot that I could try out at some point. My thought is that if the lock ups also happen in Windows, it is probably a hardware issue. However, since it's an intermittent issue and I'm a bit strapped for time at the end of the school semester, I don't really have the days necessary to test this on Windows, at least not for another couple of weeks.
In the mean time, any info or help would be appreciated. I am first trying to find out if this is a hardware issue or a firmware/software issue. Second, if it is a software issue, I am looking for any fixes or workarounds to stop this from happening.
Thanks for your help.
Last edited by Doctor_Propain (2019-12-19 00:34:56)
Offline
Is there a backtrace for the dividie by zero? https://github.com/torvalds/linux/blob/ … _20.c#L858
Offline
There wasn't one in journalctl. Those were the only four log entries at that exact time.
Offline
As the system hangs when the issue triggers I added BUG_ON to trigger an oops to try and catch the issue.
diff --git a/drivers/gpu/drm/amd/display/dc/dml/dcn20/display_mode_vba_20.c b/drivers/gpu/drm/amd/display/dc/dml/dcn20/display_mode_vba_20.c
index 6c6c486b774a..da046862a786 100644
--- a/drivers/gpu/drm/amd/display/dc/dml/dcn20/display_mode_vba_20.c
+++ b/drivers/gpu/drm/amd/display/dc/dml/dcn20/display_mode_vba_20.c
@@ -922,6 +922,7 @@ static unsigned int CalculateVMAndRowBytes(
/ 256;
}
if (GPUVMEnable == true) {
+ BUG_ON(!VMMPageSize);
MetaPTEBytesFrame = (dml_ceil(
(double) (DCCMetaSurfaceBytes - VMMPageSize)
/ (8 * VMMPageSize),
@@ -954,6 +955,8 @@ static unsigned int CalculateVMAndRowBytes(
MacroTileSizeBytes = 262144;
MacroTileHeight = 32 * BlockHeight256Bytes;
}
+ BUG_ON(!BytePerPixel);
+ BUG_ON(!MacroTileHeight);
*MacroTileWidth = MacroTileSizeBytes / BytePerPixel / MacroTileHeight;
if (GPUVMEnable == true && mode_lib->vba.GPUVMMaxPageTableLevels > 1) {
@@ -1005,6 +1008,7 @@ static unsigned int CalculateVMAndRowBytes(
unsigned int EffectivePDEProcessingBufIn64KBReqs;
if (SurfaceTiling == dm_sw_linear) {
+ BUG_ON(!BytePerPixel);
PixelPTEReqHeight = 1;
PixelPTEReqWidth = 8.0 * VMMPageSize / BytePerPixel;
PTERequestSize = 64;
@@ -1035,6 +1039,9 @@ static unsigned int CalculateVMAndRowBytes(
EffectivePDEProcessingBufIn64KBReqs = PDEProcessingBufIn64KBReqs;
if (SurfaceTiling == dm_sw_linear) {
+ BUG_ON(!BytePerPixel);
+ BUG_ON(!Pitch);
+ BUG_ON(!PixelPTEReqWidth);
*dpte_row_height =
dml_min(
128,
@@ -1055,11 +1062,13 @@ static unsigned int CalculateVMAndRowBytes(
/ PixelPTEReqWidth,
1) + 1);
} else if (ScanDirection == dm_horz) {
+ BUG_ON(!PixelPTEReqWidth);
*dpte_row_height = PixelPTEReqHeight;
*PixelPTEBytesPerRow = PTERequestSize
* (dml_ceil(((double) SwathWidth - 1) / PixelPTEReqWidth, 1)
+ 1);
} else {
+ BUG_ON(!PixelPTEReqHeight);
*dpte_row_height = dml_min(PixelPTEReqWidth, *MacroTileWidth);
*PixelPTEBytesPerRow = PTERequestSize
* (dml_ceil(To apply
git clone git://git.archlinux.org/svntogit/packages.git --single-branch --branch "packages/linux"
cd packages/trunkApply the patch. As the PKGBUILD automatically applies any source entry ending .patch you can skip altering the prepare() function.
Then you can build and test the package.
Offline
Thanks a bunch, I was able to successfully apply the patch and install it with makepkg -si. I do have one question, though. After makepkg called pacman, I got a warning that linux was already up to date, reinstalling. To me, this warning made it a bit unclear to me if it was installing linux from pacman's cached version (from the official repos) or from the newly-built package. Is this a normal message, or perhaps I did something wrong?
Thanks again.
Last edited by Doctor_Propain (2019-12-09 23:49:20)
Offline
That is a normal message. You can check the installed version details with
pacman -Qi linuxIf the installed package is the locally built version it should show
Packager : Unknown Packager
Validated By : NoneOffline
Well, it's been about a week, and I haven't had any lock-ups since I applied that patch, so I am going to mark this as resolved for now. I'll update in case it happens in the future. If this turns out to be the correct fix, would this patch be a good candidate for the Linux Kernel?
Offline