mohitgurav20 commented on issue #25797: URL: https://github.com/apache/datafusion/issues/25797#issuecomment-5857628268
Hey! I've been looking into this and I'd love to take it on if nobody else is already working on it. My plan is to fold the bloom filter evaluation directly into the load_bloom_filters async loop. This way, we can check predicate.prune() incrementally right after each column's filter is loaded, and simply short-circuit out of the I/O loop if it rules out the row group. Since we'd be evaluating eagerly, we won't need to retain the filters in memory for a secondary pass anymore. I'm planning to completely drop the BloomFiltersLoadedParquetOpen and PruneWithBloomFilters states, making the state machine cleaner by transitioning directly to BuildStream. I'll also add a quick mutable accessor to RowGroupAccessPlanFilter so we can skip the row groups right from the loader. Could you assign this to me? I can have a PR up for this shortly! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
