andygrove commented on PR #6263: URL: https://github.com/apache/datafusion-comet/pull/6263#issuecomment-5876125842
This is a light fully automated review since there are so many PRs open. The TopK guidance at `docs/source/user-guide/latest/metrics.md:124` and `docs/source/user-guide/latest/tuning/operators.md:145` points readers at `row_groups_pruned_dynamic_filter` for row-group savings. That counter only sees part of the pruning when a task reads more than one file, which is common once Spark packs small files into one partition. DataFusion 55.1 snapshots the live threshold when it opens each file and prunes in the open-time statistics pass, so groups skipped in files opened after the heap fills are counted in `row_groups_pruned_statistics` instead. The runtime pruner is also only built when more than one row group remains. With one row group per file, `row_groups_pruned_dynamic_filter` stays at 0 even when TopK skips nearly every file. The `separate_files` case in `native/core/src/execution/operators/dynamic_filter/topk/tests/statistics.rs` already exercises that open-time path. Could both docs say that TopK pruning of later files shows up in `row_groups_pruned_ statistics`, the way the hash join section does at `metrics.md:107`? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
