FW13 Pro: Suspend-Then-Hibernate system crash. possible firmware issue: s2idle resume — platform resets (Panther Lake, BIOS 3.02)

I have suspend-then-hibernate enabled on my laptop. Twice in the last 3 weeks I’ve come back to it hard reset with no explainable reason. I used claude to help me collect a diagnosis, but this is frankly beyond my expertise since there doesn’t appear to be any reason logged in the system.

Twice in the last three weeks my Framework 13 Pro has hard-reset itself out of
s2idle sleep. The firmware records a fatal SOC error in BERT each time, so
the reset is happening below the OS. Posting the decode in case it’s useful.

System

  • Framework Laptop 13 Pro (Intel Core Ultra Series 3), board FRANMJCP07
  • Core Ultra X7 358H, microcode 0x11a
  • BIOS 3.02 (2026-05-26), System Firmware 0.0.3.2 — fwupdmgr reports no update available
  • ME 21.0.6.1503
  • Fedora 44, kernel 7.2.5-200.fc44.x86_64
  • suspend-then-hibernate, s2idle — this platform offers S0 S4 S5 only

Symptom

The kernel log simply stops mid-sleep. There is no PM: suspend exit, no
shutdown sequence, no panic — the next line in the journal is a cold boot.

Sep 19 23:54:58 kernel: PM: suspend entry (s2idle)
-- boot ends here --
Sep 20 00:55:15 kernel: Linux version 7.2.5-200.fc44.x86_64 ...

Suspend was entered at 23:54:58. systemd’s suspend-then-hibernate wakes once
an hour to re-evaluate; that wake fired at 00:54:58 and the machine reset 17
seconds later. So the fault is on the resume path out of s2idle.

What the firmware recorded

/sys/fs/pstore is empty and efi_pstore is loaded, so the kernel never
panicked — it never regained control. But BERT has a record:

kernel: ACPI: BERT 0x000000006FFC7000 000030 (v01 INSYDE PTL 00000002 ACPI 00040000)
kernel: GHES: APEI firmware first mode is enabled by APEI bit.
kernel: BERT: [Hardware Error]: Skipped 1 error records
kernel: BERT: Total records found: 1

The kernel skipped printing it (records over 1 KB are truncated from dmesg), so
I pulled the region out of /dev/mem and decoded the CPER by hand:

block_status   0x00000000   (consumed by the kernel this boot)
data_length    27512 bytes
severity       1 = FATAL

section GUID   81212a96-09ed-4996-9471-8d729c8e69ed
section type   Firmware Error Record Reference
section sev    1 = FATAL     revision 0x0300
payload        19408 bytes

fw record type 2 = SOC Firmware Error Record Type2, rev 2
record GUID    8f87f311-c998-4d9e-a0c4-6065518c4f6d

The 19 KB payload is an Intel crashlog blob — Intel(R) Core(TM) Ultra X7 358H
with the rest is 32-bit register values. I’ve kept the raw dump and can attach
it.

Correlation

BERT records appear in exactly the boots following an unexplained reset, and
nowhere else:

Sep 14 03:24   reset out of s2idle    1 record
Sep 20 00:55   reset out of s2idle    1 record
8 other boots  clean shutdown         0 records

Rate is low: 989 suspends logged, 987 resumes. Two failures, both fatal.

Ruled out (grain of salt, Claude ruled out)

  • Not a kernel panic — pstore empty, efi_pstore loaded.
  • Not a failed hibernate — systemd-hibernate-resume reports PM: Image not found (code -22); no image was ever written. (which tracks, it was on AC, and I have it configured to not hibernate on AC)
  • Not thermal, not battery depletion — it was on AC, and the first failure happened 24 seconds into suspend.
  • No userspace involvement — the fault is recorded by the SoC, and the OS log ends before it.

Question

Is this a known Panther Lake s0ix exit issue, and is there anything past 3.02
in the pipeline that touches it? Happy to dig in more if you point me at what to dump.

Had another super strange behavior after hibernate, possibly related:

Follow-up: second failure mode, same subsystem — _WAK aborts on hibernate resume and takes the ACPI namespace with it

Four days after the s2idle reset above, the same machine failed again on a
sleep transition — this time out of hibernate (S4), and with a very
different and much more legible signature. Same BIOS 3.02, same kernel, no
config changes in between.

What happened

Hibernated at 23:15:43. On resume, ACPI’s wake method aborted:

23:15:43 kernel: PM: hibernation: hibernation entry
23:21:27 kernel: ACPI BIOS Error (bug): Could not resolve symbol [\_WAK.G4], AE_NOT_FOUND (20260408/psargs-365)
23:21:27 kernel: ACPI Error: Aborting method \_WAK due to previous error (AE_NOT_FOUND) (20260408/psparse-543)
23:21:27 kernel: PM: hibernation: hibernation exit

_WAK is the method the OS calls to bring the platform back up after sleep. It
referenced an undefined symbol G4 and bailed out partway. The machine
“resumed” but the ACPI namespace never finished coming back.

The damage

From that instant to the reboot 7 minutes later: 2,791 ACPI errors — 1,626
AE_NOT_FOUND and 1,822 AE_AML_PACKAGE_LIMIT. The methods that kept failing:

765  \_SB.PC00.LPCB.ACAD._PSR        AC adapter state unreadable
348  \ADBG
 99  \_SB.PC00.LPCB.EC0.EST3         EC thermal
 97  \_SB.IETM.SEN3._TMP             thermal sensor
 66  \_SB.PEPD._DSM                  power engine plugin
 56  \_SB.PC00.TDM0._S0W / TDM1._S0W
 25  \_SB.PC00.LPCB.EC0.LID0._LID    lid state

Observable effects:

  • BAT1 disappeared from /sys/class/power_supply entirely — BAT1._STA was
    failing on AE_NOT_FOUND for symbol RMS
  • /proc/acpi/button/lid/*/state returned unsupported
  • systemd-sleep began logging HibernateOnACPower=no was ignored because the system does not have a battery
  • Thermal sensors stopped reporting

The runaway

With the lid unreadable, logind kept deciding the lid was shut. It issued
29 suspend requests in seven minutes, each waking about two seconds later:

23:25:08 systemd-logind: Suspending, then hibernating...
23:25:10 kernel: PM: suspend exit
23:25:39 systemd-logind: Suspending, then hibernating...
23:25:41 kernel: PM: suspend exit

That is what made the issue obvious: the screen locking every 30 seconds,
because each suspend attempt fires the session’s before-sleep hook.

Recovery

systemctl poweroff, unplug USB-C, hold power for 30 s to reset the EC, boot.
Clean afterward: 0 ACPI errors, BAT1 back at 78%, lid reads open,
thermal sensors reporting, no BERT record. A warm reboot was not attempted —
given BAT1 had vanished I went straight for the EC reset.

Why I think these two reports are the same bug

Both failures are in the platform’s sleep-transition firmware, four days apart
on an untouched machine:

  • Sep 14 / Sep 20 — s2idle resume, fatal SOC error in BERT, platform reset
    itself with no OS involvement
  • Sep 24 — hibernate resume, _WAK aborts on an unresolved symbol, ACPI
    namespace left corrupt until reboot

Different exit paths, same subsystem. This one is the more useful of the two to
chase, because an unresolved symbol in _WAK is a concrete, addressable defect
in the DSDT rather than an opaque Intel crashlog.

Questions

  1. Is \_WAK.G4 a known unresolved reference in the 3.02 DSDT for this board?
    Happy to pull the tables with acpidump and attach them.
  2. Same for BAT1._STA.RMS — that one also resolved to nothing.
  3. Is there anything after 3.02 in the pipeline for Panther Lake sleep/resume?
    fwupdmgr still reports no update available here.

Raw BERT blob from the earlier crash is still on hand if anyone at Framework
wants it.

This may be better suited to be posted to the Framework GitHub Firmware issues area.

Without getting lost in the weeds of details, a few messages are not defined or missing at some point and a race condition results in failure.

No recovery mechanism exists for these conditions so resetting to a steady state is required for the system which is not expected.

It is difficult to regressively test for all cases on modern systems with so many unique aspects to all the players involved of subcomponents.