Neuw84 opened a new pull request, #18258:
URL: https://github.com/apache/iceberg/pull/18258

   Closes #18257.
   
   ## What
   
   `BaseDeleteLoader.loadEqualityDeletes` caches each equality delete file's 
rows, but it builds the merged `StructLikeSet` again in every scan task. On a 
table with many equality deletes, those inserts dominate the task, and every 
task on an executor repeats them for the same delete files.
   
   This change caches the merged set through the loader's existing `canCache` / 
`getOrLoad` hooks:
   
   - **The key:** the projection plus the sorted delete file locations. Tasks 
with exactly the same delete files share one set. A task whose file list 
differs in any file (sequence numbers can exclude deletes for newer data files) 
gets its own entry.
   - **Reading the files:** when the merged set is cached, its files are read 
directly, not through their per-file entries. A cache load must not start other 
cache loads, since the delete worker threads would wait on the cache this load 
holds, and the merged set supersedes those entries anyway.
   - **When the set isn't cacheable:** nothing changes; it takes the existing 
path.
   - **Sizing:** the size handed to the cache is the sum of the per-file 
estimates (the convention `estimateEqDeletesSize` already uses), so 
`max-entry-size` / `max-total-size` decide as they do today.
   - **Thread safety:** the set isn't mutated after it's built, and 
`StructLikeSet.contains` is safe for concurrent readers.
   
   ## Results
   
   This is a Spark 4.1.3 cold `MERGE INTO` (CDC: 42.0 M updates, 11.0 M 
deletes, 4.6 M inserts) into a v2 table with 432 M rows, on 8 × m5.4xlarge:
   
   - **Timing:** 162.1 s on 1.11.0, **74.6 s** with this change (2.17×).
   - **Where it comes from:** the two scan stages that apply the equality 
deletes drop from 118 s and 96 s to 10 s and 32 s.
   - **Correctness:** both runs leave the same 425,665,996 rows.
   - **Cache limits:** the executor cache limits were raised to 2 / 4 GiB on 
both runs so the merged set fits (see the open question on the issue).
   
   ## Related
   
   On the same workload, the next-largest cost in the scan was building 
position-delete indexes: `Deletes.toPositionIndexes` hashed and looked up the 
data-file path for every row. That was 20.5 % of the target-scan stage in JFR. 
#17864 has since fixed it on `main` (not in a release yet). We backported it to 
1.11.0 for our measurements, and the two changes are independent: this one is 
equality deletes only, #17864 is v2 position delete files only.
   
   ## Tests
   
   `TestBaseDeleteLoaderEqualityDeleteCache`:
   
   - the merged set is built once per group of files;
   - different files, or different equality fields, get their own set;
   - without caching every call builds its own set;
   - a merged set too large for the cache falls back to the per-file entries.
   
   `TestGenericReaderDeletes` passes unchanged.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to