NetBSD-Bugs archive
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index][Old Index]
Re: port-amd64/60631: Intel N150 cpu hang leading to kernel panic in heartbeat
The following reply was made to PR port-amd64/60631; it has been noted by GNATS.
From: Taylor R Campbell <riastradh%NetBSD.org@localhost>
To: gnats-bugs%netbsd.org@localhost
Cc: port-amd64-maintainer%netbsd.org@localhost, gnats-admin%netbsd.org@localhost,
netbsd-bugs%netbsd.org@localhost, phil.netbsd%ashbysoft.com@localhost
Subject: Re: port-amd64/60631: Intel N150 cpu hang leading to kernel panic
in heartbeat
Date: Sat, 22 Aug 2026 15:20:05 +0000
Crash dump 2 makes it look like maybe there's an interrupt storm, or a
buggy interrupt handler going in a loop, preventing hardclock from
making progress:
crash> mach cpu 0
using CPU 0
crash> bt
__kernel_end() at 0
intr_kdtrace_wrapper() at intr_kdtrace_wrapper+0x32
Xhandle_ioapic_edge2() at Xhandle_ioapic_edge2+0x75
--- interrupt ---
Xspllower() at Xspllower+0xe
hardclock() at hardclock+0xb3
Xresume_lapic_ltimer() at Xresume_lapic_ltimer+0x1e
--- interrupt ---
Xspllower() at Xspllower+0xe
comopen() at comopen+0x2a1
cdev_open() at cdev_open+0x12a
spec_open() at spec_open+0x1e8
VOP_OPEN() at VOP_OPEN+0x3e
vn_open() at vn_open+0x32c
do_open() at do_open+0xb3
do_sys_openat() at do_sys_openat+0x72
sys_open() at sys_open+0x24
syscall() at syscall+0x9a
--- syscall (number 5) ---
syscall+0x9a:
Let's see how many interrupts have arrived so far:
crash> show event/i
evcnt type 1: TLB shootdown = 342
evcnt type 1: cpu0 timer = 301
evcnt type 1: msi1 vec 0 = 57
evcnt type 1: ioapic0 pin 16 = 5
evcnt type 1: msix2 vec 1 = 1084
evcnt type 1: msix2 vec 2 = 8587
evcnt type 1: msix2 vec 3 = 87
evcnt type 1: msix2 vec 4 = 3
Looks like these correspond to nvme queues:
nvme0: for admin queue interrupting at msix2 vec 0
nvme0: WD PC SN540 SDDPNPF-512G, firmware 33006000, serial 251051800551
nvme0: for io queue 1 interrupting at msix2 vec 1 affinity to cpu0
nvme0: for io queue 2 interrupting at msix2 vec 2 affinity to cpu1
nvme0: for io queue 3 interrupting at msix2 vec 3 affinity to cpu2
nvme0: for io queue 4 interrupting at msix2 vec 4 affinity to cpu3
Is 9-10k interrupts from nvme by this time a lot? Seems like it might
be a lot. But only 1k of them were on cpu0, and it doesn't seem like
that should be enough to hold up the timecounter for 15sec.
The x86_stihlt part is a red herring -- cpu3 is just idling (with the
HLT instruction as ACPI recommended), and then when the hardclock
timer fired on cpu3, it detected that the timecounter hasn't advanced
and panicked:
crash> bt
__kernel_end() at 0
vpanic() at vpanic+0x171
panic() at vprintf
heartbeat() at heartbeat+0x175
hardclock() at hardclock+0x9c
Xresume_lapic_ltimer() at Xresume_lapic_ltimer+0x1e
--- interrupt ---
x86_stihlt() at x86_stihlt+0x6
acpicpu_cstate_idle() at acpicpu_cstate_idle+0x128
idle_loop() at idle_loop+0x89
heartbeat+0x175 is not very helpful, but going back one instruction
gives:
(gdb) info line *(heartbeat+0x170)
Line 681 of "/usr/src/sys/kern/kern_heartbeat.c"
starts at address 0xffffffff80ce9f39 <heartbeat+359>
and ends at 0xffffffff80ce9f47 <heartbeat+373>.
676 stamp =
677 atomic_load_relaxed(&curcpu()->ci_heartbeat_uptime_stamp);
678 d = count - stamp;
679 if (__predict_false(d > period_ticks) &&
680 !heartbeat_timecounter_suspended()) {
=> 681 panic("%s: time has not advanced in %u heartbeats",
682 cpu_name(curcpu()), d);
683 }
https://nxr.NetBSD.org/xref/src/sys/kern/kern_heartbeat.c?r=1.14#681
Home |
Main Index |
Thread Index |
Old Index