Aias00 opened a new issue, #6502:
URL: https://github.com/apache/shenyu/issues/6502

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/shenyu/issues) and found no similar issues.
   
   ### Apache ShenYu Component
   
   shenyu-admin
   
   ### What happened
   
   `UpstreamCheckService` keeps all submitted async checks in the shared 
`futures` list and clears it only after `CompletableFuture.allOf(...).join()` 
succeeds:
   
   ```java
   private void waitFinish() {
       CompletableFuture.allOf(futures.toArray(new 
CompletableFuture[0])).join();
       futures.clear();
   }
   ```
   
   The scheduled task catches exceptions around `doCheck()`/`waitFinish()`:
   
   ```java
   private void scheduled() {
       try {
           doCheck();
           waitFinish();
       } catch (Exception e) {
           LOG.error("upstream scheduled check error", e);
       }
   }
   ```
   
   However, `checkZombie(...)` submits `CompletableFuture.runAsync(() -> 
checkZombie0(...), invokeExecutor)` without an `exceptionally` handler, and the 
aggregation future inside `check(...)` can also fail if `updateHandler(...)` 
throws. For example, database writes, selector conversion, or event publishing 
inside `updateHandler(...)` can throw a runtime exception.
   
   When any future completes exceptionally, `allOf(...).join()` throws 
`CompletionException`, so `futures.clear()` is skipped. The failed future 
remains in the shared list. Every later scheduled run includes the same 
already-failed future, so `join()` keeps throwing before the list can be 
cleared. This leaves the upstream health check loop stuck logging errors and 
retaining old futures instead of processing future health check results 
normally.
   
   ### Expected behavior
   
   One failed async health-check/update operation should be logged and isolated 
to that cycle. The service should always clear completed futures for the cycle, 
for example by using `try/finally`, attaching exception handlers to every 
submitted future, or snapshotting and clearing the list before joining.
   
   ### How to reproduce
   
   1. Enable HTTP register upstream health checking.
   2. Have at least one upstream/zombie entry in `UPSTREAM_MAP` or `ZOMBIE_SET`.
   3. Trigger a runtime exception from an async check path, for example make 
`updateSelectorHandler(...)` fail while processing a health status change.
   4. `waitFinish()` throws from `CompletableFuture.allOf(...).join()` and 
skips `futures.clear()`.
   5. Later scheduled runs keep joining the same failed future and cannot 
recover without restarting admin.
   
   ### Debug logs
   
   _No response_
   
   ### Environment
   
   Current `master` branch.
   
   ### Are you willing to submit a PR?
   
   - [ ] Yes I am willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to