Neuw84 opened a new issue, #18259: URL: https://github.com/apache/iceberg/issues/18259
`ParquetValueReaders.StringReader` decodes every value with `column.nextBinary().toStringUsingUTF8()`. When a page is dictionary-encoded, Parquet's dictionary reader hands back the same `Binary` instance for every occurrence of a dictionary entry. The reader still decodes each of them into a new `String`. This shows up when position deletes are loaded. `BaseDeleteLoader.readPosDeletes` reads a whole position delete file through the generic reader. `file_path` is dictionary-encoded and, since the spec sorts position deletes by `file_path`, it repeats for thousands of rows in a row. We profiled a Spark 4.1 `MERGE INTO` over a v2 table with position and equality deletes (8 executors, JFR on all of them). Loading the position deletes was 10.5 % of the target-scan stage, even with #17864 applied. Inside that load, `StringReader.read` for `file_path` was 22.9 % of the samples. **Proposal** In `StringReader`, keep the last `Binary` read and its `String`, and return that `String` while the next value is the same instance. Values are unchanged. A PR with a unit test and a JMH benchmark (loading a whole position delete file through `BaseDeleteLoader`: 143.3 ms on `main`, 113.1 ms with the change) follows. Related: #17864, #16440, #16052. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
