apport information ** Tags added: apport-collected cloud-image noble
** Description changed: **Title:** Kernel panic "Attempted to kill the idle task" on AMD Opteron 6300 series (multi-node NUMA) under KVM load - linux 6.8.0-137-generic (24.04) **Summary:** On 5 out of 5 identical HP ProLiant DL385p Gen8 servers (AMD Opteron 6386 SE, 4 NUMA nodes each, running as KVM/libvirt hypervisors with multiple guest VMs), kernel panics with "Attempted to kill the idle task!" occurred within roughly 1 to 24 hours after upgrading from Ubuntu 22.04 (kernel 5.15.0-187-generic, stable in production for years) to Ubuntu 24.04 (kernel 6.8.0-137-generic) via do-release-upgrade. Pinning the boot kernel back to 5.15.0-187-generic (via GRUB_DEFAULT) while keeping the 24.04 userspace fully resolved the instability on all affected hosts, with no further crashes since. **Hardware:** - HP ProLiant DL385p Gen8, BIOS A28 (03/14/2018) - AMD Opteron 6386 SE, microcode 0x6000852 - 4 NUMA nodes per host - KVM/libvirt hypervisor, 9-20 guest VMs per host under production load (mixed Kubernetes control-plane/worker nodes, haproxy, monitoring, storage services) **Steps to reproduce:** 1. Fresh `do-release-upgrade` from 22.04 (5.15.0-187-generic) to 24.04 (6.8.0-137-generic) on the described hardware 2. Boot into 6.8.0-137-generic with multiple KVM guests running under moderate-to-high, mixed CPU/memory load 3. Panic occurs non-deterministically, anywhere from ~1 hour to ~24 hours after boot; not tied to any single reproducible user action **Observed panic signature (consistent across all 5 affected hosts):** ``` Kernel panic - not syncing: Attempted to kill the idle task! Call Trace: <TASK> ? multi_cpu_stop+0xa0/0x120 ? __pfx_multi_cpu_stop+0x10/0x10 cpu_stopper_thread+0x9c/0x170 ? __pfx_smpboot_thread_fn+0x10/0x10 smpboot_thread_fn+0xe3/0x1e0 kthread+0xf2/0x120 ? __pfx_kthread+0x10/0x10 ret_from_fork+0x47/0x70 ? __pfx_kthread+0x10/0x10 ret_from_fork_asm+0x1b/0x30 </TASK> ``` One instance on a related host (same hardware) additionally showed `membarrier_private_expedited+0x1cf/0x2d0` in RIP immediately preceding the panic. **Possible related upstream discussion:** A 2018 LKML thread by Matt Fleming (SUSE), "cpu stopper threads and load balancing leads to deadlock" (https://lkml.iu.edu/hypermail/linux/kernel/1804.2/01642.html), describes a closely matching mechanism: a CPU stopper thread entering idle/active load balancing while, concurrently, NUMA balancing on another CPU attempts to migrate a task via `stop_two_cpus()`, resulting in a deadlock over the stopper lock. Given that our affected hosts are multi-node NUMA systems (4 nodes) and the panic call trace centers on `cpu_stopper_thread`/`multi_cpu_stop`, this looks like a strong candidate mechanism, though we cannot confirm whether it is the same unresolved issue, a regression of a past fix, or a new but related bug in 6.8. **Workaround applied:** GRUB kernel pinning to boot `5.15.0-187-generic` while keeping the 24.04 userspace intact (`GRUB_DEFAULT` set to the jammy kernel's advanced boot entry ID, `update-grub`, reboot). All 5 hosts stable since applying this workaround. **Additional diagnostics performed:** - Hardware (BIOS version, CPU model, microcode) verified identical across all 5 affected hosts - Memtester (4GB sample, full test cycle) passed clean on the two hosts tested - HP Integrated Management Log (IML) reviewed on two hosts: one shows a documented Uncorrectable CPU Machine Check Exception from 2020 (Bank 0x04, both physical processors) - possibly a contributing factor for that specific host, but does not explain the same panic signature occurring on hosts without such MCE history - Network/bonding layers (LACP, active-backup) verified healthy and unrelated to the panics - `dmesg`/`journalctl` EDAC output shows no correctable/uncorrectable memory errors logged around panic times on the hosts checked - **Please let me know what additional information (full dmesg, kernel - config, /proc/cpuinfo, NUMA topology via `numactl --hardware`, etc.) - would help diagnose this further.** + **Please let me know what additional information (full dmesg, kernel config, /proc/cpuinfo, NUMA topology via `numactl --hardware`, etc.) would help diagnose this further.** + --- + ProblemType: Bug + AlsaDevices: + total 0 + crw-rw---- 1 root audio 116, 1 Aug 18 15:57 seq + crw-rw---- 1 root audio 116, 33 Aug 18 15:57 timer + AplayDevices: Error: [Errno 2] No such file or directory: 'aplay' + ApportVersion: 2.28.3-0ubuntu0.1 + Architecture: amd64 + ArecordDevices: Error: [Errno 2] No such file or directory: 'arecord' + AudioDevicesInUse: Error: command ['fuser', '-v', '/dev/snd/seq', '/dev/snd/timer'] failed with exit code 1: + CRDA: N/A + CasperMD5CheckResult: unknown + CloudBuildName: server + CloudSerial: 20240524 + DistroRelease: Ubuntu 24.04 + IwConfig: Error: [Errno 2] No such file or directory: 'iwconfig' + MachineType: HP ProLiant DL385p Gen8 + Package: linux (not installed) + PciMultimedia: + + ProcEnviron: + LANG=C.UTF-8 + PATH=(custom, no user) + SHELL=/bin/bash + TERM=xterm-256color + ProcFB: + + ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-6.8.0-137-generic root=UUID=9d93cee9-3b2d-4aff-b422-59c0ba9a1198 ro console=ttyS0,115200n8 console=tty0 vga=normal nomodeset nointremap + ProcVersionSignature: Ubuntu 6.8.0-137.137-generic 6.8.12 + RelatedPackageVersions: + linux-restricted-modules-6.8.0-137-generic N/A + linux-backports-modules-6.8.0-137-generic N/A + linux-firmware 20240318.git3b128b60-0ubuntu2.29 + RfKill: Error: [Errno 2] No such file or directory: 'rfkill' + Tags: cloud-image noble + Uname: Linux 6.8.0-137-generic x86_64 + UpgradeStatus: Upgraded to noble on 2026-08-13 (5 days ago) + UserGroups: N/A + _MarkForUpload: True + dmi.bios.date: 03/14/2018 + dmi.bios.vendor: HP + dmi.bios.version: A28 + dmi.chassis.type: 23 + dmi.chassis.vendor: HP + dmi.ec.firmware.release: 2.82 + dmi.modalias: dmi:bvnHP:bvrA28:bd03/14/2018:efr2.82:svnHP:pnProLiantDL385pGen8:pvr:cvnHP:ct23:cvr:sku653203-B21: + dmi.product.family: ProLiant + dmi.product.name: ProLiant DL385p Gen8 + dmi.product.sku: 653203-B21 + dmi.sys.vendor: HP ** Attachment added: "CurrentDmesg.txt" https://bugs.launchpad.net/bugs/2163642/+attachment/5993119/+files/CurrentDmesg.txt -- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2163642 Title: Kernel panic "Attempted to kill the idle task" on AMD Opteron multi- node NUMA under KVM (6.8.0-137) To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163642/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
