Ayoubhm07 opened a new issue, #4003:
URL: https://github.com/apache/iceberg-python/issues/4003

   ### Apache Iceberg version
   
   main (development)
   
   ### Please describe the bug 🐞
   
   While profiling the delete path from #3129, I noticed that plan_files() 
stops using file bounds as soon as an IN predicate contains more than 200 
values.
   
   In _InclusiveMetricsEvaluationVisitor.visit_in, we return ROWS_MIGHT_MATCH 
before checking the file bounds:
   
   keys=200   planned_files=1  /20 rewritten=1
   keys=201   planned_files=20 /20 rewritten=1
   keys=1000  planned_files=20 /20 rewritten=1
   
   This was on an unpartitioned table with 20 files of 10k rows each, disjoint 
key ranges, with all deleted keys belonging to the first file.
   
   So the 201st value does not make the predicate significantly more expensive 
— it disables pruning and makes every file a candidate.
   
   I also see the same effect when scaling the table: with the same layout, the 
200 → 201 transition added about 0.15s on 5 files vs 3.60s on 80 files (median 
of 7 runs).
   
   The limit comes from #1588 / #1672. The original concern was that evaluating 
large IN predicates could cost more than the pruning saves. However, the bounds 
check can potentially be reduced to a single min() / max() computation per 
predicate instead of scanning all literals for every file.
   
   I would keep the current precise evaluation below the 200-value limit and 
only restore bounds-based pruning above it, where we currently don't prune at 
all.
   
   One more thing: the same evaluator is used by conflict detection in 
table/update/validate.py, so this may also affect false-positive conflicts for 
large IN predicates.
   
   Questions
   Is disabling all pruning above 200 values intentional?
   Would a min/max bounds check above the limit be acceptable?
   Should the Python change be mirrored in Java?
   
   I haven't implemented the fix yet; I'd rather confirm the intended behavior 
first.
   
   Benchmark: local SQLite catalog, Python 3.10, pyarrow 25.0.1, pyiceberg 
0d58407, median of 7 runs.
   
   I used an AI assistant to help run the benchmarks and inspect the code path; 
the reproduction and measurements are mine.
   
   ### Willingness to contribute
   
   - [x] I can contribute a fix for this bug independently
   - [ ] I would be willing to contribute a fix for this bug with guidance from 
the Iceberg community
   - [ ] I cannot contribute a fix for this bug at this time


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to