Thank you for the suggestions. I've applied `apport-collect 2163642` from srv05 while it was running the affected kernel (6.8.0-137-generic) - diagnostic data has been uploaded to this report.
I've also installed and tested `linux-generic-hwe-24.04` (kernel 7.0.0-29-generic) on srv05, the host with the highest VM density (20 guest VMs) and the one that had crashed most frequently under 6.8.0-137. As of this update, srv05 has been running kernel 7.0.0-29-generic under normal production load for **~15 hours without a kernel panic**, compared to the 1-24 hour window in which panics previously occurred on 6.8.0-137 across all 5 affected hosts. This is encouraging but not yet conclusive - I'll continue monitoring and report back once we reach 24+ hours of stable uptime, and will consider rolling the same kernel out to the remaining affected hosts if stability holds. -- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2163642 Title: Kernel panic "Attempted to kill the idle task" on AMD Opteron multi- node NUMA under KVM (6.8.0-137) To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163642/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
