goutamadwant commented on issue #25437: URL: https://github.com/apache/datafusion/issues/25437#issuecomment-5842763421
sure @stuhood looked through the current scan path and think this could be split into smaller PRs: - First, handle simple range-key filters where we can prove a partition has no matching rows. If we can't prove it, keep the partition. - For `ListingTable`, leave skipped file groups empty but keep their partition numbers. This avoids reading those files without disrupting joins that rely on matching partition numbers. I'd test split-point boundaries, NULL/DESC cases, and join results. Let me know if you have any suggestion here. - Follow up with pruning for other Range-producing plans and reducing the number of scheduled partitions where it's safe to do so. so the first step would save file reads, but empty file groups alone won't eliminate every downstream task. Does this look good or you have any suggestions i can add here? thanks! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
