On Fri Aug 28, 2026 at 9:53 PM CEST, Shakeel Butt wrote:
On Thu, Aug 27, 2026 at 06:36:29PM +0800, Hui Zhu wrote:
From: Hui Zhu <[email protected]>
Add bpf_proactive_reclaim(), a sleepable kfunc which performs one
proactive reclaim pass on a given memory cgroup, similar to a write
to memory.reclaim but without retrying until the target is reached.
The kfunc refuses to reclaim if the calling task is already in a
reclaim context, as a nested reclaim would corrupt the outer reclaim
state.
Signed-off-by: Hui Zhu <[email protected]>
---
mm/bpf_memcontrol.c | 46 +++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 46 insertions(+)
diff --git a/mm/bpf_memcontrol.c b/mm/bpf_memcontrol.c
index 716df49d7647..297ff7f05042 100644
--- a/mm/bpf_memcontrol.c
+++ b/mm/bpf_memcontrol.c
@@ -6,6 +6,7 @@
*/
#include <linux/memcontrol.h>
+#include <linux/swap.h>
#include <linux/bpf.h>
__bpf_kfunc_start_defs();
@@ -159,6 +160,49 @@ __bpf_kfunc void bpf_mem_cgroup_flush_stats(struct
mem_cgroup *memcg)
mem_cgroup_flush_stats(memcg);
}
+/*
+ * Reclaim must not recurse: try_to_free_mem_cgroup_pages() overwrites
+ * current->reclaim_state, so a nested call would corrupt the outer
+ * reclaim state. Reclaim windows are marked with PF_MEMALLOC;
+ * reclaim_state is also checked because it is installed slightly
+ * before PF_MEMALLOC.
+ */
+static bool bpf_in_reclaim_context(void)
+{
+ return (current->flags & PF_MEMALLOC) || current->reclaim_state;
+}
+
+/**
+ * bpf_proactive_reclaim - proactively reclaim memory from a memory
+ * cgroup
+ * @memcg: the target memory cgroup to reclaim from
+ * @size: the amount of memory to reclaim, in bytes
+ *
+ * Trigger one proactive reclaim pass on @memcg, similar to a write to
+ * memory.reclaim, but without retrying until @size is reached.
+ * Must not be called with a filesystem lock held: the reclaim path
+ * may deadlock on it via filesystem shrinkers.
+ *
+ * Return: The amount of memory reclaimed, in bytes, or 0 if @size is
+ * smaller than a page or the task is already in a reclaim context.
+ */
+__bpf_kfunc unsigned long bpf_proactive_reclaim(struct mem_cgroup *memcg,
+ unsigned long size)
+{
+ unsigned long nr_reclaimed;
+
+ if (size < PAGE_SIZE || unlikely(bpf_in_reclaim_context()))
+ return 0;
I have been thinking about this more and more and after looking at the reasoning
behind your check current->reclaim_state and also Sashiko's comment on NOIO/NOFS
contexts, I am more convinced that this kfunc can not be a simple sleepable
function. We need more than that. We need clean process context as well similar
to the userspace poking memory.reclaim. Something like kthread or workqueue.
Kumar & Andrii, is there a way to restrict a kfunc to only be called from
special BPF threads/workqueues? Is there some similar concept in BPF world?
Yeah, I think the concern is valid. E.g. inode_rmdir() is sleepable but will be
problematic here, I think. My first instinct was if bpf_in_reclaim_context() +
nofs/noio save-restore might provide enough protection to let it be callable
from generic sleepable contexts, but I guess that will disable invocation of
filesystem shrinkers unconditionally.
So my suggestion would be to fix the context to BPF_PROG_TYPE_SYSCALL. There, we
should be able to init and schedule timers which poll specific state and arms wq
execution etc. It then remains invocable from sleepable async contexts (wq,
task_work) or the syscall program, all of which should be ok. Once BPF kthread
lands we can let it be callable from those threads as well, but that is for
later.
Hi Kumar and Shakeel,
The patch has been modified according to your comments.
Please help me review it.
Best,
Hui
[...]