per-CPU LRU batches
->drain work queued to every core
->PID 1 stuck in flush_work because the kworker can never preempt the
FIFO spinner on the isolated core.
The trigger is cpuset_migrate_mm (a cpuset.mems change or a task moving
between cgroups with different mems). SSH is just the first visible
victim.
Our understanding so far
Every CPU keeps a small batch of memory pages it hasn't yet added to the
LRU list. Those are the per-CPU LRU batches. Normally the batches get
drained constantly and nobody notices.
Some operations need every batch on every CPU drained before they can
carry on. Changing a cgroup's memory settings is one of them. That's
cpuset_migrate_mm, and it looks like that's what systemd was doing here.
To drain the batches, the kernel asks a helper thread on each CPU to do
it, then waits in flush_work for all of them to finish. That's the drain
work queued to every core.
One CPU has been set aside and a single program is running on it flat
out. That program appears to have been given top scheduling priority,
SCHED_FIFO, so nothing else on that CPU gets a turn. The helper thread
on that CPU is stuck behind it and never runs, so its batch never gets
drained.
If that's right, systemd, PID 1, is waiting for a helper that will never
answer, and anything else that needs the batches drained queues up
behind it. SSH broke first but was probably just the first in line.
On the interrupt approach
Forcing the isolated CPU to drain its batch from an interrupt does seem
to break the wait, but we have some doubts about it. The drain code
doesn't look like it was written to run in interrupt context, so it
might collide with the program on that CPU if that program is in the
kernel touching the same batch at the time.
Some things that might be worth looking at
These are suggestions only. You're closer to this than we are.
1. Is SCHED_FIFO intended for that program? If it was set by a start script or
a systemd unit and isn't actually needed, dropping it back to normal priority
might make the problem go away on its own, since the
helper thread could then share the CPU.
2. Could the isolated CPU avoid batching at all? One idea is a small
kernel change so that isolated CPUs add each page to the LRU list
immediately rather than batching. The batch would always be empty, so
nothing would be queued to that CPU and nothing would wait on it. The
kernel seems to do something similar already for the buffer_head cache
on isolated CPUs, so there may be a pattern to follow. There'd be a
small per-page cost on the isolated CPU, which might be acceptable given
how little memory work it should be doing.
Happy to be corrected on any of this. :-)
--
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2165410
Title:
Ubuntu LTS 26.04 linux-aws: systemd enters D state and blocks SSH
during cpuset migration on nohz_full CPUs
To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux-aws/+bug/2165410/+subscriptions
--
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs