Samyak2 commented on issue #23194: URL: https://github.com/apache/datafusion/issues/23194#issuecomment-5858808212
> * **Post-shuffle partition coalescing** — collapse many tiny partitions (over-estimated repartition) into a smaller number of right-sized ones. Needs per-partition row counts. > * **Adaptive aggregation strategy** — pick between hash and sort aggregation, or between `Partial` and `Single`, based on observed cardinality. To add to these points, one "adaptive execution" method we had seen the need for was this: https://github.com/apache/datafusion/issues/20847 Being able to dynamically decide whether we need to partitioned aggregate/partitioned join is something that DataFusion doesn't naturally support. The output partitions (and their nature, like unknown/hash) are fixed at plan time. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
