anuragmantri opened a new pull request, #18393:
URL: https://github.com/apache/iceberg/pull/18393

   Depends on #18324
   
   This branch includes the commits of #18324. Only the last commit 
a4a473b1080ce6f8c8536a0bbfe3d147294ac6dd is new. 
   
   #18324 adds the `Stitcher` API and row-based reads of data files with column 
files. This PR adds a `VectorizedStitchingIterable` lines up the batches of the 
data file and its column files by row position, slices each to the rows they 
share, and places their column vectors by reference, so no values are copied. 
`VectorizedParquetReader` now tracks batch positions. Its `advanceTo` keeps the 
batch holding the target because vectorized readers can't stop mid-batch.
   
   Test plan:
   - Batches that cut each other. Batch sizes 3 and 6 run against row groups of 
4, 7 and 5 rows, so most output batches start or end inside some part's batch. 
Aligned files would hide off-by-one errors in the windows.
   - One part ahead of the others. A split of the data file, and a filter that 
prunes the first row group of a column file. These are the only paths through 
`VectorizedParquetReader.advanceTo`.
   - Deletes on stitched batches. A deletion vector, an equality delete on an 
updated column, and `_deleted`. `BatchDeleteFilter` reads `_pos`, column-file 
values and the `_deleted` vector through slices.
   - Metadata and identity-partition columns, including the `_partition` 
struct, pass through slice views.
   - Stitched columns are the parts' own vectors, and only partial batches are 
sliced.
   
   ---
   **AI Disclosure**
   - Model: Claude Opus 5.5
   - Platform/Tool: Claude Code 
   - Human Oversight: I reviewed the changes manually.
   - Prompt Summary: Add vectorized column-file reader  onto the `Stitcher` API 
from #18324 for Spark 4.2. 
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to