Thank you for the suggestions.

I've applied `apport-collect 2163642` from srv05 while it was running
the affected kernel (6.8.0-137-generic) - diagnostic data has been
uploaded to this report.

I've also installed and tested `linux-generic-hwe-24.04` (kernel
7.0.0-29-generic) on srv05, the host with the highest VM density (20
guest VMs) and the one that had crashed most frequently under 6.8.0-137.
As of this update, srv05 has been running kernel 7.0.0-29-generic under
normal production load for **~15 hours without a kernel panic**,
compared to the 1-24 hour window in which panics previously occurred on
6.8.0-137 across all 5 affected hosts.

This is encouraging but not yet conclusive - I'll continue monitoring
and report back once we reach 24+ hours of stable uptime, and will
consider rolling the same kernel out to the remaining affected hosts if
stability holds.

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2163642

Title:
  Kernel panic "Attempted to kill the idle task" on AMD Opteron multi-
  node NUMA under KVM (6.8.0-137)

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163642/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to