haohuaijin opened a new issue, #25797:
URL: https://github.com/apache/datafusion/issues/25797

   ### Is your feature request related to a problem or challenge?
   
   Parquet Bloom filter pruning currently loads filters for the relevant 
predicate columns before evaluating whether a row group can be excluded.
   
   For queries with multiple filtering conditions, some of these reads may be 
unnecessary. For example:
   
   ```sql
   SELECT COUNT(*)
   FROM traces
   WHERE trace_id = 'abc123'
     AND service_name = 'checkout'
     AND host = 'host-42';
   ```
   
   If a row group survives statistics pruning, but its `trace_id` Bloom filter 
proves that `'abc123'` is absent, the entire row group can already be excluded. 
Reading the `service_name` and `host` Bloom filters cannot change that decision.
   
   This is particularly relevant to observability workloads that filter across 
several columns, especially when Parquet files reside in remote object storage.
   
   ### Describe the solution you'd like
   
   Evaluate Bloom filters incrementally and stop reading additional filters for 
a row group as soon as a necessary condition proves that the group cannot match.
   
   The desired behavior is:
   
   1. Read a relevant Bloom filter.
   2. Check whether it rules out the row group.
   3. If it does, skip the remaining Bloom filter reads for that group.
   4. Otherwise, continue with the remaining relevant columns.
   
   Existing literal guarantees could provide the necessary conditions. For 
example, an `IN` condition can exclude a group only when all candidate values 
are definitely absent. A possible Bloom hit must remain inconclusive.
   
   This would also allow filters to be released after evaluation instead of 
retaining all filters until a separate pruning pass. The implementation should 
preserve conservative handling of compound predicates, NULLs, missing filters, 
and read failures.
   
   ### Describe alternatives you've considered
   
   Optimizing evaluation after loading all filters could reduce CPU overhead, 
but would not avoid unnecessary Bloom filter reads.
   
   ### Additional context
   
   The relevant loading and pruning flow is in 
`datafusion/datasource-parquet/src/opener/mod.rs`, particularly 
`load_bloom_filters()` and `prune_bloom_filters()`.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to