andygrove commented on PR #6261:
URL: 
https://github.com/apache/datafusion-comet/pull/6261#issuecomment-5861059774

   I also checked this at the JVM level. With `COMET_WORKER_THREADS=1`, 
`local[4]` and 96m or 128m of off-heap memory, a native scan feeding 
`sortWithinPartitions` hung on `main` in 4 runs out of 4. The only worker was 
parked in `ExecutionMemoryPool.acquireMemory` waiting for 1/2N of the pool, 
with all four task threads in `executePlan`. With this PR merged onto `main` it 
passed 11 runs out of 11, in about 5 s. One of those runs hit the same wait and 
got through it.
   
   A side effect worth knowing about: because each acquire can hand the 
worker's core to another thread, a one-worker runtime now runs several spawned 
plans at once. The same query without memory pressure took 13.0 s on `main` 
with one worker and 7.7 to 8.6 s with this PR, against 5.5 s with four workers 
either way.
   
   This PR doesn't change one related thing, and I've filed it separately as 
#6294. A producer dropped by the runtime shutting down still looks like end of 
stream to `executePlan`, so the task finishes with whatever it had.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to