allthingssecurity opened a new pull request, #26944: URL: https://github.com/apache/camel/pull/26944
# Description [CAMEL-25061](https://issues.apache.org/jira/browse/CAMEL-25061) Since CAMEL-22784 (#20578, backported to 4.14.x in #20780), `FileLockClusterService` reads and writes the cluster data file on a separate executor, created lazily by `getClusterDataTaskExecutor()` when the field is null. `doStop()` shuts that executor down but does not reset the field, unlike the scheduled executor handled just above it. So when the service is stopped and started again in the same JVM, for example with the stop and start JMX operations of the cluster service, the views start again but every cluster data task is submitted to the executor that was shut down and is rejected (`RejectedExecutionException` from `CompletableFuture.supplyAsync` in `FileLockClusterTaskExecutor`). The leadership check logs it at DEBUG level and retries at every interval, so the node never becomes the leader again, and its `master:` routes and clustered routes never run on it until the JVM is restarted. Affected: 4.14.5 and later, and 4.17.0 and later. This change: `doStop()` sets `clusterDataTaskExecutor` to null after shutting it down, as it already does for the scheduled executor. A new one is created on the next start. No upgrade guide entry: the only visible change is that a restarted service works again. Tests: new `FileLockClusterServiceRestartTest` in camel-core, next to `FileLockClusteredRoutePolicyTest`. One CamelContext with a `FileLockClusterService` on a temporary root: the node becomes the leader and writes its heartbeat, the service is stopped (the node is no longer the leader and the lock file is free), then started again, and the node must become the leader and write its heartbeat again. This is repeated once more. Without the main-code change the test fails after the first restart: ``` testLeaderAgainAfterServiceRestart:82->awaitLeader:92 » ConditionTimeout Condition with Lambda expression in FileLockClusterServiceRestartTest was not fulfilled within 10 seconds. ``` With the change, the test passed 3 times in a row (`-Dsurefire.rerunFailingTestsCount=0`, under 1.1 s each). `FileLock*` and `*Cluster*` in camel-core (36 tests), the whole camel-file suite (22 tests) and the whole camel-master suite (30 tests, including `FileLockClusterServiceBasicFailoverTest` and `FileLockClusterServiceAdvancedFailoverTest`) pass. They also pass with #26932 (CAMEL-25052, `FileLockClusterView`) merged on top. # Target - [x] I checked that the commit is targeting the correct branch (Camel 4 uses the `main` branch) # Tracking - [x] If this is a large change, bug fix, or code improvement, I checked there is a [JIRA issue](https://issues.apache.org/jira/browse/CAMEL) filed for the change (usually before you start working on it). # Apache Camel coding standards and style - [x] I checked that each commit in the pull request has a meaningful subject line and body. - [ ] I have run `mvn clean install -DskipTests` locally from root folder and I have committed all auto-generated changes. (I built and tested the affected modules, including the formatter and import-sort plugins. I did not run the full root build.) # AI-assisted contributions - [x] If this PR includes AI-generated code, commits have proper co-authorship attribution (e.g., `Co-authored-by` trailers) and the PR description identifies the AI tool used. This PR was prepared with Claude Code (Claude Opus 5.5). The commit carries a `Co-Authored-By` trailer. _Claude Code on behalf of allthingssecurity_ 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
