anoopj opened a new pull request, #4085:
URL: https://github.com/apache/iceberg-python/pull/4085

   # Rationale for this change
   When a scan applies deletes, we  loads the deletion vector that applies to 
each data file. For Puffin deletion vectors it read the entire file into memory 
and parsed the footer to locate and deserialize every blob, then returned the 
one for the referenced data file.
   
   Instead, read only the referenced blob with a single ranged read using 
`content_offset` and `content_size_in_bytes` from the manifest, and take the 
referenced data file from the manifest as well, matching the Java and Rust 
readers. Validate the blob's DV_MAGIC and CRC-32 while stripping the framing. 
   
   # Details
   
   - Performance: a deletion vector read is now a single ranged read of one 
blob rather than loading the whole Puffin file and deserializing every blob it 
contains. That cuts I/O (notably against object storage, where only the blob's 
byte range is fetched), CPU, and memory, and scales with the referenced vector 
rather than the size of the shared container.
   - Compatibility: deletion vectors that are not fully-formed Puffin files for 
example Delta-compatible vectors that omit the footer become readable, since 
the footer is never consulted.
   
   Note: Reading one blob per manifest entry surfaces gaps that reading every 
blob previously masked, so match and preserve each deletion vector by its 
target:
   
   - Route deletion vectors by referenced_data_file in DeleteFileIndex. 
Deletion vectors need not carry path bounds, so without this they fall into a 
shared partition bucket and collapse by file_path, giving every data file in 
the partition the same vector.
   - Deduplicate delete files on (file_path, content_offset) in 
_read_all_delete_files. DataFile equality keys only on file_path, so multiple 
deletion vectors packed into one Puffin file would otherwise collapse into a 
single read.
   - Fill referenced_data_file from the scan task's data file when converting 
REST position deletes. The field is optional in the REST schema, but the offset 
read requires it.
   
   ## Are these changes tested?
   
   Added unit tests
   
   ## Are there any user-facing changes?
   
   No
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to