That’s sadly the difficult truth. I don’t even know how far AMD gives testing devices to their Linux devs. If you have something very rare (for all I remember only one other person or so said to have the same or a similar bug) it becomes really difficult to figure these things out. One of the AMD guys even recommended to me to try to bisect the problem to figure out the offending commit. But when you can’t even tell how to reliably trigger the issue, and when you have not the slightest clue about bisecting (compiling the kernel with a given config as Debian package is the most I know, and that’s dead simple), that’s sadly also not helpful. For all I remember, AMD has added some stuff to the kernel that may or may not help (no idea if for 6.16 or 6.17), but when it gets just too intricate, no matter how much I’d like to help at least by providing logs and stuff, that’s just so far above my paygrade…
I used to feel like it was probably option 1, but I’m actually a little bit more over time leaning more towards option 2.
That’s also the direction I’m heading in. Sadly the request form for the phase change thermal pad is closed, the next time my current issue appears I’ll look at the thermals. Beyond that log message I posted, the only thing these occurrences have in common is that the device is quite warm (usually from charging). Maybe it’s actually a thermal issue, even though it never happens in the rare occasions I actually have my FW in performance mode and e.g. compile the kernel, which should produce a lot more heat than the situations where this occurs, but you never know.
I have had problems with a Mediatek network card, an AX210 network card, and an AMD GPU, to an extent that seems like it would have popped up and got dealt with, if it was a general expansion hardware / driver issue.
Wow, I have had a similar experience. Though, MediaTek is just an abomination, it’s guaranteed to cause issues no matter the device and the OS. I also switched to the AX210, the only issue I have is 6 GHz WiFi not working (but lacking a decent AP for testing, but I just ordered one, and also having the same issue across distros, including Windows and the MTK module), though I never owned the dGPU expansion bay module. Additionally I have a weird bug with the headphone jack module that at least appears with debian, Ubuntu and Debian with the upstream kernel compiled based on a Debian config, but not with Fedora…Also I fear that e.g. compared to the infamous Intel 13th and 14th gen processors that had these weird issues, the AMD Ryzen 7000 (and especially the laptop SKUs) just aren’t used by that many people to manifest in numbers large enough to raise suspicion. I’m not sure with which kernel version all this started, but even Ubuntu 25.04 has “only” 6.14, and the Linux user base will probably just not be that massive for rare errors to crop up that much. So there’s basically nothing to build theories on since it seems impossible to get hard evidence.
and probably if that does not work I will go way way back to kernel 6.6 and see if that produces some positive impact.
Good idea since that’s still LTS. Also, all the VRR/PSR/PR glitches would be gone since they weren’t supported back then. Maybe I’ll try that out soon too.