Neuw84 opened a new pull request, #18258: URL: https://github.com/apache/iceberg/pull/18258
Closes #18257. ## What `BaseDeleteLoader.loadEqualityDeletes` caches each equality delete file's rows, but it builds the merged `StructLikeSet` again in every scan task. On a table with many equality deletes, those inserts dominate the task, and every task on an executor repeats them for the same delete files. This change caches the merged set through the loader's existing `canCache` / `getOrLoad` hooks: - **The key:** the projection plus the sorted delete file locations. Tasks with exactly the same delete files share one set. A task whose file list differs in any file (sequence numbers can exclude deletes for newer data files) gets its own entry. - **Reading the files:** when the merged set is cached, its files are read directly, not through their per-file entries. A cache load must not start other cache loads, since the delete worker threads would wait on the cache this load holds, and the merged set supersedes those entries anyway. - **When the set isn't cacheable:** nothing changes; it takes the existing path. - **Sizing:** the size handed to the cache is the sum of the per-file estimates (the convention `estimateEqDeletesSize` already uses), so `max-entry-size` / `max-total-size` decide as they do today. - **Thread safety:** the set isn't mutated after it's built, and `StructLikeSet.contains` is safe for concurrent readers. ## Results This is a Spark 4.1.3 cold `MERGE INTO` (CDC: 42.0 M updates, 11.0 M deletes, 4.6 M inserts) into a v2 table with 432 M rows, on 8 × m5.4xlarge: - **Timing:** 162.1 s on 1.11.0, **74.6 s** with this change (2.17×). - **Where it comes from:** the two scan stages that apply the equality deletes drop from 118 s and 96 s to 10 s and 32 s. - **Correctness:** both runs leave the same 425,665,996 rows. - **Cache limits:** the executor cache limits were raised to 2 / 4 GiB on both runs so the merged set fits (see the open question on the issue). ## Related On the same workload, the next-largest cost in the scan was building position-delete indexes: `Deletes.toPositionIndexes` hashed and looked up the data-file path for every row. That was 20.5 % of the target-scan stage in JFR. #17864 has since fixed it on `main` (not in a release yet). We backported it to 1.11.0 for our measurements, and the two changes are independent: this one is equality deletes only, #17864 is v2 position delete files only. ## Tests `TestBaseDeleteLoaderEqualityDeleteCache`: - the merged set is built once per group of files; - different files, or different equality fields, get their own set; - without caching every call builds its own set; - a merged set too large for the cache falls back to the per-file entries. `TestGenericReaderDeletes` passes unchanged. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
