CH-Abhinav commented on issue #16729:
URL: https://github.com/apache/lucene/issues/16729#issuecomment-5954137856
Thanks for sharing your ideas @abernardi597!
Just to clarify on the approach in this PR: this isn't static work
assignment — work remains fully dynamic. Instead of pre-assigning static slices
to specific workers, it converts the batches themselves into discrete tasks
submitted to `TaskExecutor#invokeAll`, backed by a worker pool
(`BlockingQueue<ConcurrentMergeWorker>`). When a helper thread becomes
available, it leases any idle worker from the pool, runs the next available
batch, and returns the worker. Because each batch is small (~2048 vectors), an
inline execution on the caller thread only takes ~50ms rather than serializing
the entire merge.
That said, I see the appeal of your **re-entrant worker detection / dynamic
recruitment** idea:
1. It avoids creating $O(N/\text{batchSize})$ task objects and queue leasing
per batch.
2. Workers stay dedicated threads draining the atomic counter as originally
designed.
One question regarding implementing the caller-thread detection: since
`TaskExecutor#invokeAll()` is a blocking call that waits for all tasks to
complete, how would you structure the recruitment retry loop? Would the caller
thread invoke single tasks in a loop while doing batches, or are you
envisioning a mechanism inside `TaskExecutor` itself?
And `trySubmit` API requires changes in underlying Lucene's architecture.
Are you sure doing that is better than simple
`BlockingQueue<ConcurrentMergeWorker>` ?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]