On Mon, Sep 14, 2026 at 11:59 PM Hao Ge <[email protected]> wrote: > > With profiling toggled off, the overflow check in > reserve_module_tags() did not run, a module could load with more > tags than the page flags can address, and re-enabling profiling > then silently corrupted /proc/allocinfo. On overflow the fix shuts > profiling down, releases the reservation and returns -EAGAIN, and > the codetag section lands as regular module data in the same load, > so the module loads without profiling. > > Review of the earlier series by Sashiko turned up two more problems. > > One is a race. layout_sections() and move_module() both asked > codetag_needs_module_section() where a codetag section goes, and > mem_profiling_support can change between the two calls, for instance > when another module load overflows the tag index and shuts profiling > down. move_module() then copied the codetag section to offset 0 of > its regular destination and clobbered the first section placed in > that region. > > v7 reworks where codetag sections are allocated, on a prototype by > Petr Pavlu [1]. The allocation runs before layout_sections() and the > placement is decided in one step, so nothing re-asks the question > and the race is gone. The retry is gone too, on -EAGAIN the section > is laid out as regular module data right in the same load. > > [1] https://lore.kernel.org/all/[email protected]/ > > Following review feedback from Petr and Suren the series is now > split into six patches. The two alloc_tag fixes from [2] are folded > in as patches 5 and 6 and replace the versions currently in the mm > tree. Sashiko keeps flagging the percpu counter leak [3], patch 5 > fixes it, and now the whole set goes through review again. > > [2] https://lore.kernel.org/all/[email protected]/ > [3] https://lore.kernel.org/all/[email protected]/ > > Patch 1 moves release_module_tags() above reserve_module_tags(), > since the failure paths now have to call it. > > Patch 2 cleans up the populate failure path: the reservation is > released, module_tags.size rolled back and the PTEs a failed > vmap_pages_range() left behind unmapped, so a later populate of the > same range is safe. It carries both Fixes tags so it backports > wherever patch 4 goes, which uses its prev_size. > > Patch 3 introduces SH_ENTSIZE_STANDALONE to mark sections with a > separate allocation. The percpu section was previously excluded > from the layout by clearing its SHF_ALLOC, which per the ELF spec > says the section occupies memory during execution, and percpu does, > only outside the regular module layout. The mark lives in sh_entsize > now, find_sec(".data..percpu") gives stable results again and > apply_relocations() goes back to testing only SHF_ALLOC. The > section would show up under /sys/module/*/sections/, but it has one > instance per CPU and no single address, and the entry never > existed, so add_sect_attrs() and add_notes_attrs() skip it. > > Patch 4 moves the codetag allocation out of move_module() in front > of layout_sections(), so the placement is decided in one step and > the race is gone. On overflow profiling is shut down, the > reservation released, module_tags.size rolled back and -EAGAIN > returned, the section is laid out as regular module data and the > module loads without profiling instead of failing. Any other error > fails the load. The release and the fallback belong together, > without the release rmmod hits the stale entry and panics. > > Patch 5 skips the percpu counter allocation when profiling is off. > After the shutdown modules load their codetag section as regular > data, load_module() still allocated counters for every tag and > release_module_tags() cannot find them on unload, so they leaked > (Suggested by Suren). > > Patch 6 defers the /proc/allocinfo removal to a workqueue. > shutdown_mem_profiling() runs under mod_lock, and the synchronous > remove_proc_entry() deadlocks with a reader taking mod_lock for > read in allocinfo_start(). The file is also created at the end of > alloc_tag_init(), a leftover file after a failed init would panic > its readers (Found by Sashiko). > > Tested on an x86_64 virtual machine: > > Booted without sysctl.vm.mem_profiling=1,compressed: > # cat /proc/allocinfo is fine > > Booted with sysctl.vm.mem_profiling=1,compressed: > # cat /proc/allocinfo is fine > # insmod overflow_tag.ko > # dmesg > With module overflow_tag there are too many tags to fit in 13 page > flag bits. Memory allocation profiling is disabled!
Do you have your overflow_tag module posted anywhere in public (github perhaps?) > # rmmod overflow_tag > The module loads without profiling and unloads cleanly. > > Also ran continuous LTP stress for a few days, nothing abnormal > so far. > > Changes in v10: > - fold in the two alloc_tag fixes from [2] as patches 5 and 6, they > replace the versions in the mm tree and fix the percpu leak > Sashiko keeps flagging [3] > - create /proc/allocinfo at the end of alloc_tag_init(), a > leftover file after a failed init would panic its readers > (Found by Sashiko) > - count note sections with sect_visible() in add_notes_attrs() too, > the count has to match the fill loop > > Changes in v9: > - do not export .data..percpu under /sys/module/*/sections/ (Petr > Pavlu). > - move the populate failure cleanup in front of the rework, v8 patch > 4 is patch 2 now. It declares prev_size itself, which the overflow > path of the rework also uses, so it carries both Fixes tags and the > two patches backport together > > Changes in v8: > - roll back module_tags.size when the reservation is released, so a > concurrent load which already passed needs_section_mem() does not > skip populate for the freed gap (Sashiko) > - unmap the PTEs a failed vmap_pages_range() installed, a retry to > populate the same range would BUG on them > - zero the separately allocated codetag memory for an SHT_NOBITS > section, the tag area pages are not zeroed on allocation (Sashiko) > - .data..percpu is exported under /sys/module/*/sections/ with the > boot CPU instance of the per-cpu area, and the interface is > documented in the ABI docs (Petr Pavlu) > - keep a comment in apply_relocations() on how .data..percpu is > relocated (Petr Pavlu) > - replace the "Based-on-a-patch-by:" tag with an in-body > attribution and a numbered Link: (Andrew Morton) > > Changes in v7: > - split the rework following review feedback (Petr Pavlu, Suren > Baghdasaryan) > - new patch 2 marks separately allocated sections with > SH_ENTSIZE_STANDALONE instead of clearing SHF_ALLOC > (suggested by Petr Pavlu) > - split the populate failure release into its own patch (suggested > by Suren Baghdasaryan) > > Changes in v6: > - rework on Petr's prototype and allocate codetag sections before > layout_sections(), the retry and its state resets are gone > - fix the layout_sections()/move_module() race (Found by Sashiko) > - release the reservation on populate failure as well (Found by > Sashiko) > - only -EAGAIN keeps the fallback, other errors fail the load > > Changes in v5: > - add Fixes: and Cc: stable to patch 1/2 as well, since 2/2 does not > compile without it (Andrew Morton) > - restore frob-adjusted mem[type].size on retry instead of zeroing, > as s390 and parisc add GOT/PLT space there in > module_frob_arch_sections() (Reported by Sashiko) > - drop the load_module() mem_profiling_support check; the percpu > counter leak is pre-existing and orthogonal to this fix > > Changes in v4: > - add a new patch (1/2) to move release_module_tags() above > reserve_module_tags(); the overflow fix is 2/2 > - release the reservation on the -EAGAIN path > - return -EAGAIN instead of -ENOMEM so the module can still load > without profiling (Suren) > - reset sh_addr, mem[type].size and sym/str SHF_ALLOC before retry > - skip percpu counters in load_module() when profiling is off > > Changes in v3: > - use pr_warn_once() instead of pr_warn() > - return -ENOMEM instead of -ENOSPC (Suren) > - expand the commit message to describe the /proc/allocinfo impact > (Andrew) > > Changes in v2: > - return an error after shutdown_mem_profiling() to skip > vm_module_tags_populate() > > v1: https://lore.kernel.org/all/[email protected]/ > v2: https://lore.kernel.org/all/[email protected]/ > v3: https://lore.kernel.org/all/[email protected]/ > v4: https://lore.kernel.org/all/[email protected]/ > v5: https://lore.kernel.org/all/[email protected]/ > v6: https://lore.kernel.org/all/[email protected]/ > v7: https://lore.kernel.org/all/[email protected]/ > v8: https://lore.kernel.org/all/[email protected]/ > v9: https://lore.kernel.org/all/[email protected]/ > > Hao Ge (6): > alloc_tag: move release_module_tags() above reserve_module_tags() > alloc_tag: clean up the populate failure path > module: introduce SH_ENTSIZE_STANDALONE for separately allocated > sections > module: allocate codetag sections before the regular module layout > alloc_tag: skip percpu counter allocation when profiling is disabled > alloc_tag: Defer /proc/allocinfo removal to a workqueue > > include/linux/module.h | 2 + > kernel/module/internal.h | 8 +++ > kernel/module/kallsyms.c | 13 +--- > kernel/module/main.c | 133 +++++++++++++++++++----------------- > kernel/module/sysfs.c | 17 +++-- > lib/codetag.c | 10 ++- > mm/alloc_tag.c | 141 +++++++++++++++++++++++---------------- > 7 files changed, 188 insertions(+), 136 deletions(-) > > -- > 2.25.1 >

