Aias00 opened a new issue, #6844:
URL: https://github.com/apache/shenyu/issues/6844

   ## Description
   `ShenyuWebsocketClient.healthCheck()` calls `this.reconnectBlocking()` — a 
synchronous TCP-connect + WS-handshake call — when `!isOpen()`. `healthCheck` 
runs on `WheelTimerFactory.getSharedTimer()`, whose `taskExecutor` is a 
single-thread `ThreadPoolExecutor(1,1)`. While one client's 
`reconnectBlocking()` blocks, ALL other timer tasks stall — including other 
websocket clients' health checks and the `WebsocketSyncDataService.masterCheck` 
round task. No backoff, no jitter, no async reconnect.
   
   ## Location
   - 
`shenyu-sync-data-center/shenyu-sync-data-websocket/src/main/java/org/apache/shenyu/plugin/sync/data/websocket/client/ShenyuWebsocketClient.java:238-251`
 (healthCheck → reconnectBlocking), `:110,148` (runs on shared wheel timer); 
`HierarchicalWheelTimer` taskExecutor = `ThreadPoolExecutor(1,1)`
   
   ## Impact
   With multiple websocket URLs (HA admin), if one admin is unreachable, 
`reconnectBlocking()` blocks the shared single-thread executor for the TCP 
timeout duration (potentially minutes). All other clients' health checks and 
the masterCheck task are delayed, causing missed heartbeats, delayed failover, 
and lost config updates across all sync connections.
   
   ## Suggested fix
   Run `reconnectBlocking` on a dedicated executor (not the shared wheel 
timer), or use the async `reconnect()` variant with a dedicated 
connection-accepted callback; add exponential backoff + jitter.
   
   ## Related existing
   Distinct from #6569 (long-polling single-thread scheduler) — this is the 
websocket client reconnect path.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to