andygrove commented on PR #6263:
URL: 
https://github.com/apache/datafusion-comet/pull/6263#issuecomment-5876125842

   This is a light fully automated review since there are so many PRs open.
   
   The TopK guidance at `docs/source/user-guide/latest/metrics.md:124` and 
`docs/source/user-guide/latest/tuning/operators.md:145` points readers at 
`row_groups_pruned_dynamic_filter` for row-group savings. That counter only 
sees part of the pruning when a task reads more than one file, which is common 
once Spark packs small files into one partition. DataFusion 55.1 snapshots the 
live threshold when it opens each file and prunes in the open-time statistics 
pass, so groups skipped in files opened after the heap fills are counted in 
`row_groups_pruned_statistics` instead. The runtime pruner is also only built 
when more than one row group remains. With one row group per file, 
`row_groups_pruned_dynamic_filter` stays at 0 even when TopK skips nearly every 
file. The `separate_files` case in 
`native/core/src/execution/operators/dynamic_filter/topk/tests/statistics.rs` 
already exercises that open-time path. Could both docs say that TopK pruning of 
later files shows up in `row_groups_pruned_
 statistics`, the way the hash join section does at `metrics.md:107`?
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to