Hi everyone,

I managed to make a slightly better reproducer, this can put systemd into D
state for between 2 to 3 to 5 minutes, depending on the run. Its not 
deterministic, eventually each core will get an opportunity to reschedule, which
lets the kworker run to drain the LRU cache, and eventually unblocks each core,
but it does the job to prove there is a problem.

Launch a c8a.metal-48xl with ubuntu resolute AMI.

$ sudo apt update
$ sudo apt upgrade
$ sudo apt install build-essential

Setup isolcpus, nohz_full and rcu_nocb by setting the following grub
command line in /etc/default/grub.d/99-repro.cfg:

$ cat << 'EOF' | sudo tee /etc/default/grub.d/99-repro.cfg
GRUB_CMDLINE_LINUX_DEFAULT="$GRUB_CMDLINE_LINUX_DEFAULT 
isolcpus=nohz,18-95,108-178 nohz_full=18-95,108-178 rcu_nocbs=18-95,108-178 
rcu_nocb_poll skew_tick=1 nosmt=force idle=poll numa_balancing=disable 
transparent_hugepage=never audit=0 nmi_watchdog=0 nowatchdog mce=ignore_ce 
tsc=reliable clocksource=tsc"
EOF

Run

$ sudo update-grub
$ sudo reboot

Once the system is up again, check the cmdline:

$ cat /proc/cmdline

Extract the attached tarball,

$ tar -xf reproducer.tar.xz
$ cd reproducer
$ chmod +x run_repro.sh
$ sudo -s
# ./run_repro.sh

There are three files in the tarball:

repro_worker.c: This program pins a SCHED_FIFO/80 thread to an isolated core,
then continuously dirties and discards a 24 MB private buffer to generate
un-drained local per-CPU LRU folio batches.

anchor_alloc.c: This program allocates and holds 50 GB of memory bound to NUMA 
node 1 inside the target cgroup slice.

run_repro.sh: spawns worker load on isolated CPUs 18-95 and 108-178 inside
a systemd slice, runs anchor_alloc to allocate the 50GB memory, and starts a
NUMA page migration to NUMA node 0.

I'll add a comment with some analysis next.

Thanks,
Matthew

** Attachment added: "reproducer.tar.xz"
   
https://bugs.launchpad.net/ubuntu/+source/linux-aws/+bug/2165410/+attachment/5998157/+files/reproducer.tar.xz

** Also affects: linux-aws (Ubuntu Resolute)
   Importance: Undecided
       Status: New

** Also affects: linux-aws (Ubuntu Stonking)
   Importance: Undecided
       Status: Confirmed

** Changed in: linux-aws (Ubuntu Stonking)
     Assignee: (unassigned) => Matthew Ruffell (mruffell)

** Changed in: linux-aws (Ubuntu Resolute)
     Assignee: (unassigned) => Matthew Ruffell (mruffell)

** Changed in: linux-aws (Ubuntu Stonking)
       Status: Confirmed => In Progress

** Changed in: linux-aws (Ubuntu Resolute)
       Status: New => In Progress

** Changed in: linux-aws (Ubuntu Resolute)
   Importance: Undecided => Medium

** Changed in: linux-aws (Ubuntu Stonking)
   Importance: Undecided => Medium

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2165410

Title:
  Ubuntu LTS 26.04 linux-aws: systemd enters D state and blocks SSH
  during cpuset migration on nohz_full CPUs

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-aws/+bug/2165410/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to