dwsmith1983 opened a new pull request, #6280: URL: https://github.com/apache/datafusion-comet/pull/6280
## Which issue does this PR close? No issue. This only changes tests. ## Rationale for this change Two Iceberg DPP tests use the native scan's `numPartitions == 1` as evidence that DPP pruned the fact table: - "runtime filtering - join with dynamic partition pruning" - "AQE DPP - CometSubqueryBroadcastExec replaces SubqueryAdaptiveBroadcastExec" Iceberg packs the tests' small files into one Spark partition whether or not DPP prunes, so both checks pass with DPP off. Running each query on main with `spark.sql.optimizer.dynamicPartitionPruning.enabled=false` still gives one partition, but the scan plans 3 file tasks instead of 1. ## What changes are included in this PR? Both tests now count the Iceberg file tasks the scan planned (the sum of `getFileScanTasksCount` over `perPartitionData`) and expect 1 of the 3 files. The AQE test also expects the `num_splits` metric to be 1. The join test does not check `num_splits`, because its `ORDER BY` runs the scan once to sample range bounds and again for the shuffle, so each split is read twice. The other two Iceberg DPP tests that check `numPartitions` are left alone. Their unpruned scans have 8 and 2 partitions, so the check already fails without pruning. ## How are these changes tested? The two tests pass on Spark 3.4, 3.5, 4.0 and 4.1. With DPP turned off for the query, the old `numPartitions` checks still pass and the new checks fail, planning 3 tasks where 1 is expected. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
