System crashes due to overheating

If I run the laptop too hard for too long it will overheat and cause the OS to freeze, requiring a hard reset. This is not a problem with the laptop heating up too easily - I do need to run it hard for a long time; but still I don’t think it’s supposed to crash when I do.

If thermal throttling is supposed to kick in I don’t think it does - when I go to reset it after a crash the power button is scalding hot, keeping it pressed for a few seconds is painful. If I weren’t using an external keyboard the laptop would become too hot to use long before it crashes.

I know nothing about how to troubleshoot thermal issues - any help is appreciated!

Which Linux distro are you using?

Ubuntu 22.04.2 LTS

Which kernel are you using?

7.0.0, but previous kernels had the same issue

Which BIOS version are you using?

4.05

Which Framework Laptop 16 model are you using?

Batch 15 of Framework Laptop 16 (AMD Ryzen™ 7040 Series)

You can replace the liquid metal to PTM7950/7958.

I’ll reiterate: I don’t have a problem with the laptop heating up too easily. It’s reasonable that it can’t run at max power for hours in the summer heat. What’s not right is that it keeps running at max power until it becomes unstable; besides the annoyance of it, I’m worried that some safety system is not working as it should and reducing the lifespan of the laptop.

The “safety system” is preventing your CPU from exceeding 100C, it doesn’t prevent other components from being heated by the CPU if the heat dissipation is poor. For an extreme example, the computer fails suspending in a backpack, causing battery swelling due to heat being transferred from the CPU to the battery. Although the CPU is capped to 100C, other components’ safety temperatures are much lower.

Earlier batches of FL16 use liquid metal that degraded faster than expected. If the thermal interface is not replaced, the CPU is going to continue transferring heat to other parts of the laptop including the keyboard.

Since not all components can withstand 100C this caused your computer to freeze or crash.

You don’t mention the fans. Are they running?
What surface have you put the laptop on?
If it is a soft surface like a sofa or bed, it will overheat.
Its best to put the laptop on some sort of hard surface, like a tray, so that the air can still get in from underneath.
Perhaps running “sensors” will tell you the temps and the fan speeds.

I don’t see any of the symptoms described in that thread though - there is no performance degradation, the fans are not running unless I put the CPU under a heavy load; after a crash and reset they typically take less than a minute to stop running. If I immediately run it hard again it takes an hour or more until the next crash. I’m no expert but this doesn’t look like the CPU is having trouble getting rid of the heat.

You say you are using Linux.
Install “amdgpu_top”.
You can get it here:

It will give you a better idea of what is going on.
How many watts it is using etc.

There’s a slight chance that the freezing is not caused by heat Framework 13 AMD 7840U freeze on Linux kernel 7.0.0 and 7.0.1 - #4 by dimitris

I booted a game and got a crash almost immediately (typically it takes much longer); this is journalctl:

Jul 22 15:52:48 raptor kernel: pcieport 0000:00:01.1: pciehp: Slot(0): Link Down
Jul 22 15:52:48 raptor kernel: pcieport 0000:00:01.1: pciehp: Slot(0): Card not present
Jul 22 15:52:48 raptor kernel: snd_hda_intel 0000:03:00.1: CORB reset timeout#2, CORBRP = 65535
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: device lost from bus!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: SMU: bus error for message: TransferTableSmu2Dram(18) response:0xFFFFFFFF 
Jul 22 15:52:49 raptor kernel: in params:00000005
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: Failed to export SMU metrics table!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: device lost from bus!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: SMU: bus error for message: TransferTableSmu2Dram(18) response:0xFFFFFFFF 
Jul 22 15:52:49 raptor kernel: in params:00000005
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: Failed to export SMU metrics table!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: device lost from bus!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: SMU: bus error for message: TransferTableSmu2Dram(18) response:0xFFFFFFFF 
Jul 22 15:52:49 raptor kernel: in params:00000005
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: Failed to export SMU metrics table!
Jul 22 15:52:49 raptor kernel: amdgpu 0000:03:00.0: device lost from bus!

and this is the output of sensors after the minute or so it took to reboot:

amdgpu-pci-0300
Adapter: PCI adapter
vddgfx:       18.00 mV 
fan1:           0 RPM  (min =    0 RPM, max = 4900 RPM)
edge:         +57.0°C  (crit = +100.0°C, hyst = -273.1°C)
                       (emerg = +105.0°C)
junction:     +58.0°C  (crit = +100.0°C, hyst = -273.1°C)
                       (emerg = +105.0°C)
mem:          +66.0°C  (crit = +105.0°C, hyst = -273.1°C)
                       (emerg = +110.0°C)
PPT:         1000.00 mW (cap = 100.00 W)

ucsi_source_psy_USBC000:004-isa-0000
Adapter: ISA adapter
in0:          20.00 V  (min =  +5.00 V, max = +20.00 V)
curr1:         3.00 A  (max =  +3.00 A)

cros_ec-isa-0000
Adapter: ISA adapter
fan1:               4748 RPM
fan2:               4593 RPM
ambient_f75303@4d:   +68.8°C  (high = +86.8°C, crit = +86.8°C)
                              (emerg = +104.8°C)
charger_f75303@4d:   +68.8°C  (high = +86.8°C, crit = +86.8°C)
                              (emerg = +104.8°C)
apu_f75303@4d:       +68.8°C  (high = +86.8°C, crit = +86.8°C)
                              (emerg = +104.8°C)
cpu@4c:              +73.8°C  (high = +104.8°C, crit = +104.8°C)
                              (emerg = +124.8°C)
gpu_amb_f75303@4d:   +61.9°C  (high = -273.1°C, crit = -273.1°C)
                              (emerg = -273.1°C)
gpu_vr_f75303@4d:    +65.8°C  (high = +70.8°C, crit = +96.8°C)
                              (emerg = +104.8°C)
gpu_vram_f75303@4d:  +62.9°C  (high = -273.1°C, crit = -273.1°C)
                              (emerg = -273.1°C)
gpu_temp@40:         +57.9°C  (high = +86.8°C, crit = +104.8°C)
                              (emerg = +107.8°C)

ucsi_source_psy_USBC000:002-isa-0000
Adapter: ISA adapter
in0:           5.00 V  (min =  +5.00 V, max =  +5.00 V)
curr1:         1.50 A  (max =  +1.50 A)

spd5118-i2c-2-50
Adapter: SMBus PIIX4 adapter port 0 at 0b00
temp1:        +69.8°C  (low  =  +0.0°C, high = +55.0°C)  ALARM (HIGH)
                       (crit low =  +0.0°C, crit = +85.0°C)

nvme-pci-0600
Adapter: PCI adapter
Composite:    +62.9°C  (low  =  -5.2°C, high = +89.8°C)
                       (crit = +93.8°C)

BAT1-acpi-0
Adapter: ACPI interface
in0:          17.31 V  
curr1:         0.00 A  

amdgpu-pci-c500
Adapter: PCI adapter
vddgfx:      969.00 mV 
vddnb:       760.00 mV 
edge:         +62.0°C  
PPT:          10.19 W  (avg =   8.15 W)

k10temp-pci-00c3
Adapter: PCI adapter
Tctl:         +72.2°C  

mt7921_phy0-pci-0500
Adapter: PCI adapter
temp1:        +66.0°C  

ucsi_source_psy_USBC000:003-isa-0000
Adapter: ISA adapter
in0:           0.00 V  (min =  +0.00 V, max =  +0.00 V)
curr1:         0.00 A  (max =  +0.00 A)

spd5118-i2c-2-51
Adapter: SMBus PIIX4 adapter port 0 at 0b00
temp1:        +68.0°C  (low  =  +0.0°C, high = +55.0°C)  ALARM (HIGH)
                       (crit low =  +0.0°C, crit = +85.0°C)

ucsi_source_psy_USBC000:001-isa-0000
Adapter: ISA adapter
in0:           5.00 V  (min =  +5.00 V, max = +20.00 V)
curr1:         5.00 A  (max =  +5.00 A)

nvme-pci-0400
Adapter: PCI adapter
Composite:    +60.9°C  (low  = -40.1°C, high = +83.8°C)
                       (crit = +87.8°C)
Sensor 1:     +70.8°C  (low  = -273.1°C, high = +65261.8°C)
Sensor 2:     +60.9°C  (low  = -273.1°C, high = +65261.8°C)

acpitz-acpi-0
Adapter: ACPI interface
temp1:        +68.8°C  
temp2:        +68.8°C  
temp3:        +68.8°C  
temp4:        +73.8°C  
temp5:        +61.8°C  
temp6:        +65.8°C  
temp7:        +62.8°C  
temp8:        +57.8°C  

so maybe it’s in fact not heat related?

Link down? Try reinstalling the interposer.

well, the interposer looked fine… reinstalled just in case… but then I had trouble reinstalling the touchpad and realized I have a swollen battery. So there’s definitely some problem with heat. On the bright side it doesn’t need a battery to boot, and maybe the extra space will help with convection!

Maybe too late now, but while you had it open, it might have been a good time to remove the expansion bay and clean the fan and vents. In case they are slightly blocked.
The FW16 has a problem in relation to overheating if the fan vents are blocked.

1 Like

Yesterday my dGPU was “device lost from bus” after heavy use while accidentaly not connecting the charger. Only when it crashed I noticed that the fans were not running at all and that the laptop had become really hot. The crash took my GNOME desktop down as well and the system was almost not reacting, so I shut it down with Alt+SysRQ combination.

After powering on again, fans were running at maximum speed for a short time, then termperatures normalised. I have never had this happen before, so I suspect the recent BIOS update (4.0.5) in combination with running on battery as the culprit - somehow the fans were not active although the dGPU was used. The “device lost from bus” seems to be the typical symptom when the dGPU overheats.

The battery is used to provide peak power. With the battery removed, the TGP is going to reduce significantly, leading to a lower temperature. It’s hard to differentiate whether the convection or the reduced TGP is the cause of a lower temperature. In addition to James3’s recommendation of clean the fan and vents, It’s better to repaste thermal interface of both CPU and GPU while waiting for the new battery, also check the condition of the thermal pads to make sure the VRMs have good cooling as well

What would prevent a second battery from swelling up again?

Even if the dies are not transferring heat properly and I manage to improve rather than ruin things with the thermal pads, wouldn’t that result in it running harder to reach the same temperature at whichever sensor it is using to monitor it, which is clearly not enough to keep the battery safe?

My impression is that long term the laptop can’t handle the way I use it and removing the battery may be the fix.

I’m not sure that sharing this will be terribly helpful for me to post, but I wanted to thank you for including the sensors output. If you’re concerned about overheating, I just wanted to chime in and say that I’ve made some (aggressive) modifications to my FW 16, and it appears to be able to sustain 80W draw on the CPU with burst peaks going just over 100W. My temps stay around 85c during this, and my fan speed averages 4000 rpm (4100 and 3900).

I don’t know what your CPU power draw/load is during the sensors grab you shared, but my fan speed is already 15% lower than that when at 100% load @ 80W+ draw. My unit is a batch 1 motherboard that shipped with the overheating issue, so I imagine if it’s possible on my board it’s possible on others as well.

Note: there’s not really much of a performance gain by unlocking the power target that far. It’s just fun.

If you’re interested in what those mods are, I’ve listed them here:
https://community.frame.work/t/rx-7700s-dgpu-vent-covers-v3-now-with-bottom-riser

All the modifications I’ve made to my FW 16 (so far) - Framework Laptop 16 - Framework Community

I agree, the fact that your laptop gets hot implies the PMT paste etc. is working well enough.
The problem appears to be the removal of that heat not working once it has left the CPU.

  1. Check the airway round the fans etc. is all clean so air flow works.
  2. I use a very simple approach. If my laptop feels hot, I just use “ectool fanduty 100” to put the fans on max until it cools down again. Note, only works in (1) is working.

My theory is that the FW 16 maybe does not have enough temp sensors, or that not all of them are used to determine fan speed. Thus some edge cases where it gets too hot.
I have found that Battery temp does not seem to trigger the fans to go on, or at least not as early as I would like.

I think, if the battery itself gets too hot, it starts swelling.

I removed the graphics module to clean it; I found some dust buildup which I’m sure was not helping but not enough to take all the blame. I am always careful with airflow around the laptop because I had this problem when it was brand new and tried to use it with the lid closed (with an external monitor). With the lid open it needs summer temperatures to act up.

I also found what I think is a surprising amount of warping in whatever this copper sheet is called:

I don’t have point of reference so I don’t know if this is normal or another symptom of overheating.

thanks to @John_Obscurant I realized that every crash I can remember happened when I was using the discrete graphics card - and while I never paid close attention to them, I think every time I’ve used discrete graphics the fans were on at full speed. So maybe I do have thermal transfer issues on the GPU; would that cause the whole laptop to overheat?

That warping is normal.

I’d recommend using a tool like HWInfo64 (if you’re on Windows) to monitor GPU temps to see if it’s overheating.

I don’t see mentioned what batch you were, or missed it.

My batch 1 dGPU had subpar thermal paste application. After I worked on the CPU, I also worked on the dGPUI and re-pasted it. That gave me ~50% performance gains on some benchmarks I ran, and also less fan ramp up.

I later on had the dGPU replaced on a warranty claim (the usb c partly failed) and that one was pasted better and didn’t have the degraded performance (I now had benchmarks from the previous re-pasting to compare).

Now I have a different dGPU (gen 2 fans) so it’s totally different.

Point is, might be worth to try and re-paste the dGPU module. I only remember it was not hard to do.