Aias00 opened a new issue, #7316:
URL: https://github.com/apache/shenyu/issues/7316

   ## Current Behavior
   
   `ShenyuWebsocketClient#onMessage` catches any `RuntimeException` raised 
while parsing or applying a configuration message and logs:
   
   ```text
   Failed to handle websocket message ..., the message will be ignored
   ```
   
   The connection remains open and the client does not request `MYSELF`, 
reconnect, mark itself unready, or notify Admin. The ignored message can 
represent a plugin, selector, rule, auth, metadata, proxy-selector, discovery, 
or API-key update.
   
   Because the WebSocket protocol has no event sequence, ACK, NACK, or replay 
cursor, one handler exception permanently diverges gateway cache from Admin 
state until another update happens to overwrite it or the gateway restarts.
   
   This behavior was introduced as a reader-thread safety improvement after 
#6843, but reader-thread survival alone does not restore consistency.
   
   ## Reproduction
   
   - Bind a subscriber whose update handler throws once.
   - Deliver a valid WebSocket UPDATE message.
   - Verify the exception is caught and ignored while the socket remains open.
   - Verify no full sync is requested and the gateway cache remains stale.
   - Restart/reconnect and verify `MYSELF` restores the current snapshot.
   
   ## Expected Behavior
   
   Failure to apply a configuration message must move the client into an 
explicit desynchronized state and initiate bounded recovery. It must not 
silently continue as healthy.
   
   ## Scope
   
   - Record message group/event/namespace and a safe event identifier or digest 
on application failure.
   - Mark synchronization state unhealthy/not-ready for the failed connection.
   - Define bounded recovery: NACK/replay if supported, otherwise 
close/reconnect and request a full snapshot.
   - Avoid infinite reconnect loops for deterministic poison data; expose a 
clear terminal error and metrics after bounded attempts.
   - Ensure the connection is not declared synchronized until the replacement 
snapshot is fully applied.
   - Coordinate with initial-sync readiness work in #7283 and 
delivery/backpressure work in #6587.
   
   ## Acceptance Criteria
   
   - [ ] A subscriber/application exception cannot be silently ignored while 
readiness remains healthy.
   - [ ] The client either replays the failed event or completes a full 
resynchronization.
   - [ ] Recovery is bounded and observable.
   - [ ] Poison data does not cause an unbounded reconnect storm.
   - [ ] Logs/metrics identify the failed group, event type, recovery attempt, 
and final result without leaking secrets.
   - [ ] Tests cover transient handler failure, permanent failure, successful 
full-sync recovery, and subsequent incremental updates.
   - [ ] End-to-end cache state converges to Admin state after a recoverable 
failure.
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to