Neuw84 commented on PR #18258:
URL: https://github.com/apache/iceberg/pull/18258#issuecomment-5845424027

   A second, independent measurement on today's code, for both engines. It's 
the same CDC `MERGE INTO` as in the description: a v2 table with 432 M rows, a 
change batch of 42.0 M updates, 11.0 M deletes and 4.6 M inserts, and 
425,665,996 rows after. It runs on 8 executors × 14 cores on m5.4xlarge, 
reading from S3. Stock Apache Iceberg 1.11.0 is compared with the same runtime 
jar carrying this change (plus a backport of #17864, which the description 
found independent). The executor cache limits are 2 / 4 GiB on every run.
   
   | Engine | Iceberg | MERGE | Target scan (the stage that applies the 
deletes): wall / task time / CPU / GC |
   |---|---|---|---|
   | OSS Spark 4.1.3 | 1.11.0 | 189.3 s, 164.1 s | 138 s, 118 s / 11,655 s, 
11,069 s / 4,080 s, 4,051 s / 4,850 s, 4,655 s |
   | OSS Spark 4.1.3 | **with this PR** | **85.9 s, 86.5 s** | **41 s, 41 s / 
3,837 s, 3,952 s / 1,914 s, 1,911 s / 224 s, 246 s** |
   | Spark with a columnar engine plugin (spark-vector) | 1.11.0 | 157.7 s, 
156.6 s | 112 s, 115 s / 10,118 s, 10,434 s / 4,142 s, 4,458 s / 3,258 s, 3,294 
s |
   | Spark with a columnar engine plugin (spark-vector) | **with this PR** | 
**80.6 s, 84.5 s** | **40 s, 40 s / 4,136 s, 4,098 s / 2,113 s, 2,099 s / 267 
s, 280 s** |
   
   The two runs of each combination were interleaved.
   
   - **The MERGE is 1.9-2.2x faster in both engines.** The table's contents 
after the merge are the same in every run.
   - **Where it comes from:** the target scan does about half the CPU and a 
twentieth of the GC. The merged `StructLikeSet` is built once per executor 
rather than once per task, so executors stop allocating it again and again.
   - **The same holds for a different execution engine:** the plugin replaces 
Spark's operators above the scan, but the delete application is Iceberg's 
reader in both, so the gain doesn't depend on Spark's own operators.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to