There are intermittent failures in collapse_max_ptes_swap() and
collapse_max_ptes_shared() when using the khugepaged_context:
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0
# Run test: collapse_max_ptes_swap (khugepaged:anon)
# Swapout 257 of 2048 pages... OK
# Maybe collapse with max_ptes_swap exceeded.... OK
# Swapout 256 of 2048 pages... OK
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 17)
# Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0
This happens because khugepaged may collapse the pages before wait_for_scan()
is called, causing a sanity check that expects uncollapsed pages to fail.
For example, in collapse_max_ptes_swap(), after faulting the pages back in
and paging out up to max_ptes_swap pages, khugepaged may collapse them again
before c->collapse() is called.
To prevent this, change the khugepaged setting from ALWAYS to MADVICE for
the affected tests, and mark the VMA with MADV_NOHUGEPAGE after it has been
collapsed by wait_for_scan(). This prevents khugepaged from collapsing it
again before c->collapse() is called.
Also, fix false-positive results when a child process fails in tests
such as collapse_fork*() or collapse_max_ptes_shared():
# -------------------------
# running ./khugepaged -s 2
# -------------------------
#
# Run test: collapse_max_ptes_shared (khugepaged:anon)
# Allocate huge page... OK
# Share huge page over fork()... OK
# Trigger CoW on page 1023 of 2048... OK
# Maybe collapse with max_ptes_shared exceeded.... OK
# Trigger CoW on page 1024 of 2048... Fail
Bail out! Unexpected huge page
# Planned tests != run tests (26 != 23)
# Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0 // child failed.
# Check if parent still has huge page... OK // parent hpage
success
ok 24 collapse_max_ptes_shared // considered as
success
...
# Totals: pass:26 fail:0 xfail:0 xpass:0 skip:0 error:0
This failure was observed on NVIDIA Spark with 16KB page and this patch
is based on mm/mm-unstable
---
Yeoreum Yun (2):
kselftest: mm: return fail when child test result is fail in khugepaged
kselftest: mm: fix intermittent failure khugepaged test
tools/testing/selftests/mm/khugepaged.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
---
base-commit: 6b41451631cabf9ea3b384c2a099088e1598f963
change-id: 20260915-fix_khugepagd_fail-9d8932689200
Best regards,
--
Sincerely,
Yeoreum Yun