mbutrovich opened a new pull request, #3302:
URL: https://github.com/apache/iceberg-rust/pull/3302

   ## Which issue does this PR close?
   
   - Closes #3298.
   
   ## What changes are included in this PR?
   
   - Add `residual_for_missing_fields` in `arrow/reader/predicate_visitor.rs`. 
It replaces each predicate leaf on a top-level field that is missing from the 
data file with `AlwaysTrue` or `AlwaysFalse`, evaluated against the value 
projection returns for that field: the identity partition value, otherwise 
`initial-default`. Leaves on missing fields with neither keep the existing null 
handling. This is the same idea as Java's `ResidualEvaluator`, extended to 
`initial-default`.
   - Apply the residual in the reader pipeline right after the field ID map is 
built, so the row filter, row group pruning, bloom filter pruning, and page 
index pruning all see the same predicate. This includes the predicate built 
from equality deletes.
   - Leave leaves on nested fields unchanged, because a nested field also reads 
as null in any row where an ancestor struct is null. This keeps filtering 
consistent with projection, which doesn't apply nested defaults yet (#3261).
   - Reuse `ExpressionEvaluatorVisitor` for leaf evaluation and `constants_map` 
for identity partition values. Both, and `LogicalExpression::new`, go from 
private to `pub(crate)`. There are no public API changes.
   
   ## Are these changes tested?
   
   Yes, with unit tests and a `MemoryCatalog` scan test:
   
   - The three reproduction tests from #3298 in `row_filter.rs` cover the row 
filter with `initial-default`, page index pruning with `initial-default`, and 
the row filter with an identity partition value. The page index test now 
applies the residual before calling `get_row_selection_for_filter_predicate`, 
as the pipeline does.
   - Residual unit tests in `predicate_visitor.rs` follow Java's 
[`TestMetricsRowGroupFilter`](https://github.com/apache/iceberg/blob/48330b8dacab6662242d252b39c8444190979bb2/data/src/test/java/org/apache/iceberg/data/TestMetricsRowGroupFilter.java#L475-L641)
 cases for string, date, and double defaults. Other cases check that the 
partition value takes precedence over `initial-default`, that bucket partition 
values are ignored, and that a present column, a column with no value, and a 
nested field are left unchanged. They also check that `NOT`, `AND`, and `OR` 
keep their structure.
   - `test_scan_filter_on_initial_default_column_absent_from_file` in 
`scan/mod.rs` follows Java's 
[`testFilterPushdownOnInitialDefaultColumnAbsentFromFile`](https://github.com/apache/iceberg/blob/48330b8dacab6662242d252b39c8444190979bb2/spark/v4.1/spark/src/test/java/org/apache/iceberg/spark/sql/TestFilterPushDown.java#L651-L691).
 It appends a file, adds a column with a default through `update_schema`, 
appends a second file, and scans with filters, with row selection off and on. 
It fails without the fix.
   
   ## AI Disclosure
   
   I wrote this PR with help from Claude but understand and support the core 
approach, implementation, and test coverage.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to