andygrove commented on PR #6261: URL: https://github.com/apache/datafusion-comet/pull/6261#issuecomment-5861059774
I also checked this at the JVM level. With `COMET_WORKER_THREADS=1`, `local[4]` and 96m or 128m of off-heap memory, a native scan feeding `sortWithinPartitions` hung on `main` in 4 runs out of 4. The only worker was parked in `ExecutionMemoryPool.acquireMemory` waiting for 1/2N of the pool, with all four task threads in `executePlan`. With this PR merged onto `main` it passed 11 runs out of 11, in about 5 s. One of those runs hit the same wait and got through it. A side effect worth knowing about: because each acquire can hand the worker's core to another thread, a one-worker runtime now runs several spawned plans at once. The same query without memory pressure took 13.0 s on `main` with one worker and 7.7 to 8.6 s with this PR, against 5.5 s with four workers either way. This PR doesn't change one related thing, and I've filed it separately as #6294. A producer dropped by the runtime shutting down still looks like end of stream to `executePlan`, so the task finishes with whatever it had. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
