There are intermittent failures in collapse_max_ptes_swap() and collapse_max_ptes_shared() when using the khugepaged_context:
// while running ./khugepaged -s 2 # Run test: collapse_max_ptes_shared (khugepaged:anon) # Allocate huge page... OK # Share huge page over fork()... OK # Trigger CoW on page 1023 of 2048... OK # Maybe collapse with max_ptes_shared exceeded.... OK # Trigger CoW on page 1024 of 2048... Fail Bail out! Unexpected huge page # Planned tests != run tests (26 != 23) # Totals: pass:23 fail:0 xfail:0 xpass:0 skip:0 error:0 # Run test: collapse_max_ptes_swap (khugepaged:anon) # Swapout 257 of 2048 pages... OK # Maybe collapse with max_ptes_swap exceeded.... OK # Swapout 256 of 2048 pages... OK Bail out! Unexpected huge page # Planned tests != run tests (26 != 17) # Totals: pass:17 fail:0 xfail:0 xpass:0 skip:0 error:0 This happens because khugepaged may collapse the pages before wait_for_scan() is called, causing a sanity check that expects uncollapsed pages to fail. For example, in collapse_max_ptes_swap(), after faulting the pages back in and paging out up to max_ptes_swap pages, khugepaged may collapse them again before c->collapse() is called. To prevent this, mark the VMA with MADV_NOHUGEPAGE after it has been collapsed by wait_for_scan() for anon. This prevents khugepaged from collapsing it again before c->collapse() is called. This failure was observed on NVIDIA Spark with 16KB page. Signed-off-by: Yeoreum Yun <[email protected]> --- tools/testing/selftests/mm/khugepaged.c | 3 +++ 1 file changed, 3 insertions(+) diff --git a/tools/testing/selftests/mm/khugepaged.c b/tools/testing/selftests/mm/khugepaged.c index c32244b565658..1aad4bb427ece 100644 --- a/tools/testing/selftests/mm/khugepaged.c +++ b/tools/testing/selftests/mm/khugepaged.c @@ -578,6 +578,9 @@ static bool wait_for_scan(const char *msg, char *p, size_t len, usleep(TICK); } + if (!strncmp(ops->name, "anon", 4)) + madvise(p, len, MADV_NOHUGEPAGE); + return timeout == -1; } -- 2.43.0

