Neuw84 opened a new issue, #18259:
URL: https://github.com/apache/iceberg/issues/18259

   `ParquetValueReaders.StringReader` decodes every value with 
`column.nextBinary().toStringUsingUTF8()`. When a page is dictionary-encoded, 
Parquet's dictionary reader hands back the same `Binary` instance for every 
occurrence of a dictionary entry. The reader still decodes each of them into a 
new `String`.
   
   This shows up when position deletes are loaded. 
`BaseDeleteLoader.readPosDeletes` reads a whole position delete file through 
the generic reader. `file_path` is dictionary-encoded and, since the spec sorts 
position deletes by `file_path`, it repeats for thousands of rows in a row. We 
profiled a Spark 4.1 `MERGE INTO` over a v2 table with position and equality 
deletes (8 executors, JFR on all of them). Loading the position deletes was 
10.5 % of the target-scan stage, even with #17864 applied. Inside that load, 
`StringReader.read` for `file_path` was 22.9 % of the samples.
   
   **Proposal**
   
   In `StringReader`, keep the last `Binary` read and its `String`, and return 
that `String` while the next value is the same instance. Values are unchanged. 
A PR with a unit test and a JMH benchmark (loading a whole position delete file 
through `BaseDeleteLoader`: 143.3 ms on `main`, 113.1 ms with the change) 
follows.
   
   Related: #17864, #16440, #16052.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to