Aias00 opened a new issue, #7356:
URL: https://github.com/apache/shenyu/issues/7356
### Is there an existing issue for this?
- [x] I have searched the existing issues.
Related issues such as #7312 cover Admin cluster mode. This issue is
specifically about multiple Admin nodes running with cluster mode disabled
while sharing one database.
### Current Behavior
Consider this topology:
```text
Admin 1 -- WebSocket -- Bootstrap 1
\
shared database
/
Admin 2 -- WebSocket -- Bootstrap 2
```
Both Admin nodes run with `shenyu.cluster.enabled=false`. Each Bootstrap
maintains a healthy WebSocket connection to its local Admin node.
When configuration is changed through Admin 1:
- the database is updated;
- Bootstrap 1 receives the change immediately;
- Admin 2 can read the new database state, but no local data-change event is
generated;
- Bootstrap 2 remains stale until an operator clicks synchronize on Admin 2
or restarts/reconnects Bootstrap 2.
### Expected Behavior
All Bootstrap instances connected to standalone Admin nodes that share the
same database should eventually converge to the latest configuration without
manual synchronization or restart.
The mechanism must preserve namespace isolation and avoid repeatedly pushing
unchanged full configuration.
### Steps To Reproduce
1. Start two Admin nodes with `shenyu.cluster.enabled=false` and the same
database.
2. Start two Bootstrap nodes using WebSocket data sync.
3. Connect Bootstrap 1 only to Admin 1 and Bootstrap 2 only to Admin 2.
4. Update a plugin, selector, rule, metadata, app-auth, discovery,
proxy-selector, or AI proxy key through Admin 1.
5. Verify that Bootstrap 1 applies the update.
6. Verify that Bootstrap 2 does not apply it while its WebSocket connection
remains healthy.
7. Click synchronize on Admin 2 or restart Bootstrap 2 and verify that it
then receives the latest database state.
### Root Cause
`DataChangedEvent` is a local Spring application event. A database write on
Admin 1 is dispatched only inside the Admin 1 JVM.
`WebsocketDataChangedListener` then sends the event through
`WebsocketCollector`, whose session collections are also local JVM state.
Sharing the database does not propagate the event to Admin 2 or to the
Bootstrap sessions connected to Admin 2.
Bootstrap requests a full `MYSELF` synchronization when a WebSocket
connection is established, but the normal health-check task only sends ping and
instance information. There is no periodic configuration reconciliation while
the connection stays healthy.
HTTP long polling already refreshes its database-backed MD5 cache
periodically, but that refresh is private to the HTTP sync strategy and does
not produce WebSocket notifications.
### Proposed Solution
Add a configurable Admin-side reconciliation mechanism for WebSocket sync.
1. Enable it only when WebSocket sync is enabled; it must support the
`cluster=false` topology described above.
2. Periodically compare a lightweight configuration version or deterministic
digest per `namespace + config group`.
3. When a version changes, reload only the affected namespace/group and
publish a local `REFRESH` to the Bootstrap sessions connected to that Admin.
4. Do not push data when the version/digest is unchanged.
5. Use fixed-delay execution, prevent overlapping runs, and add initial
jitter so multiple Admin nodes do not query the shared database simultaneously.
6. Do not advance the local reconciliation cursor after a failed load or
failed notification; the next cycle must retry.
7. Document the consistency delay and database-load trade-off of the
configured interval.
A database-backed monotonic version table keyed by `namespace + config
group` is preferred for production use. A deterministic full-data digest can be
used as an initial implementation if schema changes are intentionally avoided.
A raw periodic call to `syncAllByNamespaceId` from every Bootstrap is not
preferred because database reads and full pushes scale with the number of
Bootstrap instances.
### Required Correctness Fix
Full reconciliation must support clearing a configuration group.
Currently:
- `WebsocketDataChangedListener` does not send an empty data list;
- Bootstrap `AbstractDataHandler` ignores an empty data list before
processing `REFRESH` or `MYSELF`.
As a result, if the final rule, selector, metadata entry, or other group
item is deleted while an event is missed, a later full synchronization cannot
clear the stale Bootstrap cache.
For `REFRESH` and `MYSELF`, an empty list must clear the corresponding
namespace/group. Existing incremental `CREATE`, `UPDATE`, and `DELETE`
semantics must remain unchanged.
### Acceptance Criteria
- [ ] Two standalone Admin nodes sharing one database are supported with
WebSocket sync.
- [ ] A change made through Admin 1 reaches Bootstrap 2 within the
configured reconciliation interval.
- [ ] No manual synchronize action or Bootstrap restart is required.
- [ ] Only changed namespaces and configuration groups are refreshed.
- [ ] Unchanged polling cycles do not push configuration data.
- [ ] Deleting the last item in a group clears the corresponding Bootstrap
cache.
- [ ] Multiple namespaces remain isolated.
- [ ] Repeated detection is idempotent and does not create a refresh loop.
- [ ] Reconciliation failures are logged and retried without advancing the
successful version/cursor.
- [ ] Cluster mode and existing WebSocket reconnect/full-sync behavior do
not regress.
- [ ] Configuration properties and operational trade-offs are documented.
### Test Plan
Unit tests:
- unchanged version/digest produces no notification;
- one changed group produces one namespace-scoped refresh;
- multiple changed groups are handled deterministically;
- an empty `REFRESH` clears subscriber state;
- duplicate polling is idempotent;
- a failed refresh is retried;
- `REFRESH`/`MYSELF` does not increment the shared change version and cause
a loop.
Integration/E2E test:
- start two standalone Admin nodes against one database;
- connect one Bootstrap to each Admin;
- create, update, and delete configuration through only one Admin;
- assert that both Bootstrap nodes converge, including deletion of the final
item in a group.
### Environment
ShenYu version: current `master` (`2.7.2-SNAPSHOT`)
Data sync: WebSocket
Admin mode: multiple standalone Admin nodes, shared database,
`shenyu.cluster.enabled=false`
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]