The bcachefs ktest allocation-leak check writes rcutree.do_rcu_barrier
before reading /proc/allocinfo. While testing bcachefs performance
changes, small objects released with kfree_rcu() remained visible after
repeated writes to the hook and 20 seconds of waiting, causing otherwise
clean tests to fail their leak check.

The test assumes a stronger contract than the hook currently documents:
rcu_barrier() waits for ordinary callbacks, but does not flush objects
still held in kfree_rcu() batching or per-CPU SLUB sheaves. The retained
population eventually fell as a sheaf filled; there is no evidence here
of unbounded growth or OOM.

Changing the hook to drain kvfree_rcu() work let the same unmodified
bcachefs workload pass its allocation check. All eight checkpoints in
one VM, after 50 through 400 option changes, reported zero retained
reconcile_scan objects. This motivated the separate private-cache test
used to isolate the incomplete drain from bcachefs.

Calling kvfree_rcu_barrier() from rcu_barrier_throttled() was proposed
when kvfree_rcu_barrier() was added in 2024, to restore a clean baseline
between userspace benchmark runs. The discussion concluded that keeping
the existing hook name, adding the second operation and documenting both
was the safest compatibility choice, but the follow-up was not added.

Add that drain and document the stronger test interface. Always retain
the existing start-rate limit and perform the kvfree_rcu() drain: an
unrelated ordinary barrier does not establish that this work completed.

Retain the entry ordinary-barrier sequence snapshot. After draining,
skip the final ordinary barrier only if that snapshot is complete,
preserving the memory barrier on the completion path. Otherwise, invoke
rcu_barrier() explicitly. This keeps the ordinary-callback guarantee
independent of whether kvfree_rcu_barrier() embeds an ordinary barrier.

Clarify that the documented completion guarantee covers work queued
before the request, without preventing new work from being queued.

Earlier validation of the unconditional-drain version used four fresh
VM pairs with a private-cache fixture: controls retained the queued
object (60 to 60 active objects), and treatments drained it (60 to 59).
An ordinary-callback test passed on both kernels. Those runs predated
the guarded skip and do not validate that change. No elapsed-time
improvement is claimed.

Link: https://lore.kernel.org/all/[email protected]/
Signed-off-by: Matthias Goergens <[email protected]>
---
 .../admin-guide/kernel-parameters.txt         |  9 ++++--
 kernel/rcu/tree.c                             | 30 ++++++++++++-------
 2 files changed, 26 insertions(+), 13 deletions(-)

diff --git a/Documentation/admin-guide/kernel-parameters.txt 
b/Documentation/admin-guide/kernel-parameters.txt
index 68647ff4bdd2..914b65ae9413 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -5699,9 +5699,12 @@ Kernel parameters
                        there is an ongoing too-long CSD-lock wait.
 
        rcutree.do_rcu_barrier= [KNL]
-                       Request a call to rcu_barrier().  This is
-                       throttled so that userspace tests can safely
-                       hammer on the sysfs variable if they so choose.
+                       Wait for deferred kfree_rcu() frees and ordinary
+                       call_rcu() callbacks queued before this request to
+                       complete.  This does not prevent new work from being
+                       queued concurrently.  Requests are throttled so that
+                       userspace tests can safely hammer on the sysfs
+                       variable if they so choose.
                        If triggered before the RCU grace-period machinery
                        is fully active, this will error out with EAGAIN.
 
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 96848fc1f02b..93b71682306c 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -3989,12 +3989,12 @@ EXPORT_SYMBOL_GPL(rcu_barrier);
 static unsigned long rcu_barrier_last_throttle;
 
 /**
- * rcu_barrier_throttled - Do rcu_barrier(), but limit to one per second
+ * rcu_barrier_throttled - Drain deferred RCU frees, but rate-limit starts
  *
- * This can be thought of as guard rails around rcu_barrier() that
- * permits unrestricted userspace use, at least assuming the hardware's
- * try_cmpxchg() is robust.  There will be at most one call per second to
- * rcu_barrier() system-wide from use of this function, which means that
+ * This can be thought of as guard rails around the deferred-free barriers
+ * that permit unrestricted userspace use, at least assuming the hardware's
+ * try_cmpxchg() is robust.  There will be at most one drain operation started
+ * per sixteenth of a second from use of this function, which means that
  * callers might needlessly wait a second or three.
  *
  * This is intended for use by test suites to avoid OOM by flushing RCU
@@ -4016,14 +4016,24 @@ static void rcu_barrier_throttled(void)
        while (time_in_range(j, old, old + HZ / 16) ||
               !try_cmpxchg(&rcu_barrier_last_throttle, &old, j)) {
                schedule_timeout_idle(HZ / 16);
-               if (rcu_seq_done(&rcu_state.barrier_sequence, s)) {
-                       smp_mb(); /* caller's subsequent code after above 
check. */
-                       return;
-               }
                j = jiffies;
                old = READ_ONCE(rcu_barrier_last_throttle);
        }
-       rcu_barrier();
+       /*
+        * kfree_rcu() can retain objects outside the ordinary callback lists in
+        * per-CPU SLUB sheaves and kvfree_rcu batches.  Always drain those 
queues:
+        * an ordinary barrier does not establish that this work was drained.
+        */
+       kvfree_rcu_barrier();
+       /*
+        * A completed barrier can still cover ordinary callbacks queued before
+        * our entry snapshot.  Otherwise, retain an explicit ordinary barrier
+        * without depending on the implementation of kvfree_rcu_barrier().
+        */
+       if (rcu_seq_done(&rcu_state.barrier_sequence, s))
+               smp_mb(); /* caller's subsequent code after above check. */
+       else
+               rcu_barrier();
 }
 
 /*
-- 
2.55.0


Reply via email to