zhf999 opened a new issue, #50703:
URL: https://github.com/apache/arrow/issues/50703

   ### Describe the enhancement requested
   
   **Component(s):** C++, Parquet
   
   ## Describe the enhancement requested
   
   When full-scanning a large Parquet file (~6.5 GB, ZSTD-compressed) whose 
columns are mostly
   BYTE_ARRAY (strings) with RLE_DICTIONARY encoding (falling back to PLAIN for 
some pages),
   single-threaded reads through 
`parquet::arrow::FileReader::GetRecordBatchReader()` are about
   20% slower than Velox and StarRocks on the same file.
   
   Profiling with VTune shows two related problems on the hot path (screenshot 
below):
   
   <img width="1390" height="431" alt="Image" 
src="https://github.com/user-attachments/assets/2ddda405-3c7a-48dc-a592-dfdd371b7729";
 />
   
   ### 1. Value-data buffer growth causes repeated full-buffer copies 
(`__memmove` is the #1 hotspot, ~18% of CPU time)
   
   - For dictionary-encoded pages, 
`DictByteArrayDecoderImpl::DecodeArrowDense()` passes
     `estimated_data_length = {}` to `ArrowBinaryHelper`, so no space is 
reserved in the
     accumulator's value-data buffer at all. The decoded size cannot be derived 
from page
     metadata (the page's `uncompressed_size` only reflects the RLE-encoded 
index stream),
     so every value goes through the checked `BinaryBuilder::Append()` path and 
the buffer
     grows incrementally, copying previously accumulated data on each resize.
   - For PLAIN pages, `PlainByteArrayDecoder::DecodeArrowDense()` does reserve, 
but
     `BufferBuilder::Reserve()` resizes to exactly `size + additional` 
(`grow_by_factor`
     defaults to `false`). Since the accumulator persists across multiple 
decode calls
     within one record batch, each call's exact reservation triggers a 
reallocation that
     copies the entire accumulated buffer again.
   
   ### 2. Per-value append bookkeeping (~25% of CPU time in aggregate)
   
   For each decoded value, the `AppendValue` → `BinaryBuilder::Append()` chain 
performs a
   capacity check (`Reserve(1)`), `ValidateOverflow`, a single-value offset 
append, and sets
   the validity bitmap one bit at a time (`SetBitTo` alone accounts for ~3% of 
total CPU).
   None of this is amortized across a batch.
   
   ## Proposed improvements
   
   1. **Dictionary path: reserve exactly, once per decode call.** The RLE 
indices have to be
      decoded anyway; decoding them up front (into a reusable scratch buffer) 
allows computing
      the exact decoded data size by summing dictionary entry lengths, followed 
by a single
      `ReserveData()`. As a bonus, `IndexInBounds` validation can also be 
hoisted out of the
      append loop, and the append loop can then use the `UnsafeAppend` fast 
path safely, since
      the exact size is a guaranteed upper bound.
   2. **Reserve with slack for cross-call growth.** When the helper reserves on 
a builder that
      already contains data from previous decode calls, grow by at least the 
current capacity
      (`max(requested, capacity)`) instead of resizing to the exact requested 
size, so repeated
      reservations are amortized.
   3. **Append validity bits per run instead of per value.** `VisitBitRuns` 
already provides
      run boundaries; the bitmap can be appended in bulk with
      `UnsafeAppendToBitmap(num_bits, value)` once per run (and null runs 
likewise), removing
      the per-value bit manipulation.
   
   ## Component(s)
   
   C++, Parquet
   
   
   ### Component(s)
   
   C++


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to