Hi everyone, I managed to make a slightly better reproducer, this can put systemd into D state for between 2 to 3 to 5 minutes, depending on the run. Its not deterministic, eventually each core will get an opportunity to reschedule, which lets the kworker run to drain the LRU cache, and eventually unblocks each core, but it does the job to prove there is a problem.
Launch a c8a.metal-48xl with ubuntu resolute AMI. $ sudo apt update $ sudo apt upgrade $ sudo apt install build-essential Setup isolcpus, nohz_full and rcu_nocb by setting the following grub command line in /etc/default/grub.d/99-repro.cfg: $ cat << 'EOF' | sudo tee /etc/default/grub.d/99-repro.cfg GRUB_CMDLINE_LINUX_DEFAULT="$GRUB_CMDLINE_LINUX_DEFAULT isolcpus=nohz,18-95,108-178 nohz_full=18-95,108-178 rcu_nocbs=18-95,108-178 rcu_nocb_poll skew_tick=1 nosmt=force idle=poll numa_balancing=disable transparent_hugepage=never audit=0 nmi_watchdog=0 nowatchdog mce=ignore_ce tsc=reliable clocksource=tsc" EOF Run $ sudo update-grub $ sudo reboot Once the system is up again, check the cmdline: $ cat /proc/cmdline Extract the attached tarball, $ tar -xf reproducer.tar.xz $ cd reproducer $ chmod +x run_repro.sh $ sudo -s # ./run_repro.sh There are three files in the tarball: repro_worker.c: This program pins a SCHED_FIFO/80 thread to an isolated core, then continuously dirties and discards a 24 MB private buffer to generate un-drained local per-CPU LRU folio batches. anchor_alloc.c: This program allocates and holds 50 GB of memory bound to NUMA node 1 inside the target cgroup slice. run_repro.sh: spawns worker load on isolated CPUs 18-95 and 108-178 inside a systemd slice, runs anchor_alloc to allocate the 50GB memory, and starts a NUMA page migration to NUMA node 0. I'll add a comment with some analysis next. Thanks, Matthew ** Attachment added: "reproducer.tar.xz" https://bugs.launchpad.net/ubuntu/+source/linux-aws/+bug/2165410/+attachment/5998157/+files/reproducer.tar.xz ** Also affects: linux-aws (Ubuntu Resolute) Importance: Undecided Status: New ** Also affects: linux-aws (Ubuntu Stonking) Importance: Undecided Status: Confirmed ** Changed in: linux-aws (Ubuntu Stonking) Assignee: (unassigned) => Matthew Ruffell (mruffell) ** Changed in: linux-aws (Ubuntu Resolute) Assignee: (unassigned) => Matthew Ruffell (mruffell) ** Changed in: linux-aws (Ubuntu Stonking) Status: Confirmed => In Progress ** Changed in: linux-aws (Ubuntu Resolute) Status: New => In Progress ** Changed in: linux-aws (Ubuntu Resolute) Importance: Undecided => Medium ** Changed in: linux-aws (Ubuntu Stonking) Importance: Undecided => Medium -- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2165410 Title: Ubuntu LTS 26.04 linux-aws: systemd enters D state and blocks SSH during cpuset migration on nohz_full CPUs To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux-aws/+bug/2165410/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
