Aias00 opened a new issue, #7316: URL: https://github.com/apache/shenyu/issues/7316
## Current Behavior `ShenyuWebsocketClient#onMessage` catches any `RuntimeException` raised while parsing or applying a configuration message and logs: ```text Failed to handle websocket message ..., the message will be ignored ``` The connection remains open and the client does not request `MYSELF`, reconnect, mark itself unready, or notify Admin. The ignored message can represent a plugin, selector, rule, auth, metadata, proxy-selector, discovery, or API-key update. Because the WebSocket protocol has no event sequence, ACK, NACK, or replay cursor, one handler exception permanently diverges gateway cache from Admin state until another update happens to overwrite it or the gateway restarts. This behavior was introduced as a reader-thread safety improvement after #6843, but reader-thread survival alone does not restore consistency. ## Reproduction - Bind a subscriber whose update handler throws once. - Deliver a valid WebSocket UPDATE message. - Verify the exception is caught and ignored while the socket remains open. - Verify no full sync is requested and the gateway cache remains stale. - Restart/reconnect and verify `MYSELF` restores the current snapshot. ## Expected Behavior Failure to apply a configuration message must move the client into an explicit desynchronized state and initiate bounded recovery. It must not silently continue as healthy. ## Scope - Record message group/event/namespace and a safe event identifier or digest on application failure. - Mark synchronization state unhealthy/not-ready for the failed connection. - Define bounded recovery: NACK/replay if supported, otherwise close/reconnect and request a full snapshot. - Avoid infinite reconnect loops for deterministic poison data; expose a clear terminal error and metrics after bounded attempts. - Ensure the connection is not declared synchronized until the replacement snapshot is fully applied. - Coordinate with initial-sync readiness work in #7283 and delivery/backpressure work in #6587. ## Acceptance Criteria - [ ] A subscriber/application exception cannot be silently ignored while readiness remains healthy. - [ ] The client either replays the failed event or completes a full resynchronization. - [ ] Recovery is bounded and observable. - [ ] Poison data does not cause an unbounded reconnect storm. - [ ] Logs/metrics identify the failed group, event type, recovery attempt, and final result without leaking secrets. - [ ] Tests cover transient handler failure, permanent failure, successful full-sync recovery, and subsequent incremental updates. - [ ] End-to-end cache state converges to Admin state after a recoverable failure. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
