sunnysabor opened a new issue, #7399: URL: https://github.com/apache/shenyu/issues/7399
## Description `FailureRegistryTask` removes a pending registration unconditionally when its retry budget is exhausted. A newer failed registration for the same key can arrive after the last retry attempt has restored its holder, but before the task's final exhaustion callback runs. `addToFail` replaces the holder and assumes the existing task will handle it; the next invocation sees the retry limit, skips `doRetry`, and `onRetryExhausted(key)` removes that newer holder. This is separate from the in-flight retry race fixed by #7259: the new failure arrives after the final failed retry has returned, while the old timer task is still scheduled for its exhaustion invocation. ## Reproduction On current `master` (`09c6a5287330c8d6ada64cd29f2a08a570b0b2ab`), the following deterministic sequence reproduces the loss: 1. Queue a failed registration and retain its `FailureRegistryTask`. 2. Run the task through all 18 failed retry attempts; each retry restores the pending holder. 3. Before the next (exhaustion) invocation, queue another failed registration with the same key. The holder is replaced, but no new timer is scheduled because the map already contains that key. 4. Run the old task once more. Its retry-exhaustion callback removes the key without retrying the newer holder. I verified this with a temporary test in `FailbackRegistryRepositoryTest`: the module's Checkstyle passed, and the final assertion failed with `expected: <1> but was: <0>`. The probe was removed afterward. ## Impact The newest failed URI, metadata, API-doc, or MCP-tools registration can be discarded without another retry. Admin and gateway registration state can remain stale until a later registration event or restart. ## Suggested fix Make retry exhaustion ownership-aware. For example, associate the scheduled task with the holder/version it owns and remove only that holder, or atomically arrange a fresh retry task when a newer holder replaces the exhausted task's payload. Add a deterministic regression test for a same-key failure queued after the final failed attempt and before exhaustion cleanup. ## Related issues and changes - #6558 covers a different race between a successful retry and its cleanup. - #7259 fixes failures arriving during an in-flight retry, but this post-retry exhaustion window remains. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
