From: Jack Wang <[email protected]> check_sb_changes() activates a spare another node added, reached from md_reload_sb() -> process_metadata_update() with reconfig_mutex held. Stacking the device's limits there waits for q->limits_lock under that mutex, which deadlocks: the lock's holder waits for the queue to drain, and that I/O can be waiting for a superblock update needing reconfig_mutex.
The device is already a member, so its limits are stacked. Add it with MDDEV_STACK_SKIP and leave them alone. Assisted-by: LLM Signed-off-by: Jack Wang <[email protected]> --- drivers/md/md.c | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/drivers/md/md.c b/drivers/md/md.c index 28fc903ffeea..e60dc2c7eb90 100644 --- a/drivers/md/md.c +++ b/drivers/md/md.c @@ -10885,7 +10885,14 @@ static void check_sb_changes(struct mddev *mddev, struct md_rdev *rdev) rdev2->saved_raid_disk = -1; else rdev2->saved_raid_disk = role; - ret = remove_and_add_spares(mddev, rdev2, NULL); + /* + * reconfig_mutex is held, so q->limits_lock + * cannot be taken here. The device is + * already a member, its limits are stacked, + * so add it without touching them. + */ + ret = remove_and_add_spares(mddev, rdev2, + MDDEV_STACK_SKIP); pr_info("Activated spare: %pg\n", rdev2->bdev); /* wakeup mddev->thread here, so array could -- 2.43.0

