Neuw84 commented on PR #6268:
URL: 
https://github.com/apache/datafusion-comet/pull/6268#issuecomment-5913466701

   We tested this PR together with #6270 on TPC-DS at SF1000: Parquet on S3, a 
Spark 4.1 cluster on EKS, 8 executors x 13 cores on m5.4xlarge, one 
availability zone, AQE on with 300 shuffle partitions.
   
   The build was main `b58b2f3a` with #6268 (`588c029f`) and #6270 (`b3f4f058`) 
merged on top (HEAD `0f22d064`), with the native library built for x86-64-v3. 
Comet ran with native scan, exec and shuffle enabled.
   
   - **q5 now completes.** 23.1 s, 100 rows, the same checksum as vanilla 
Spark, with 77/77 operators on Comet (11x `CometNativeScanExec`). Before this 
fix, q5 failed at this scale with native scans.
   - **The whole run passes.** All 103 queries complete. Row counts match 
vanilla Spark on every query, and checksums match on all but q65, which has 
ties in its ORDER BY and differs the same way for every engine we compare.
   
   Thanks for the fix.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to