Neuw84 commented on PR #6268: URL: https://github.com/apache/datafusion-comet/pull/6268#issuecomment-5913466701
We tested this PR together with #6270 on TPC-DS at SF1000: Parquet on S3, a Spark 4.1 cluster on EKS, 8 executors x 13 cores on m5.4xlarge, one availability zone, AQE on with 300 shuffle partitions. The build was main `b58b2f3a` with #6268 (`588c029f`) and #6270 (`b3f4f058`) merged on top (HEAD `0f22d064`), with the native library built for x86-64-v3. Comet ran with native scan, exec and shuffle enabled. - **q5 now completes.** 23.1 s, 100 rows, the same checksum as vanilla Spark, with 77/77 operators on Comet (11x `CometNativeScanExec`). Before this fix, q5 failed at this scale with native scans. - **The whole run passes.** All 103 queries complete. Row counts match vanilla Spark on every query, and checksums match on all but q65, which has ties in its ORDER BY and differs the same way for every engine we compare. Thanks for the fix. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
