pingzh opened a new pull request, #6281:
URL: https://github.com/apache/datafusion-comet/pull/6281

   ## Which issue does this PR close?
   
   Closes #5910.
   
   ## Rationale for this change
   
   An off-heap Spark `UTF8String` currently passes through `getByteBuffer()`, 
which allocates a JVM byte array and copies the payload before Arrow copies it 
into its own buffer. Copying directly into reserved Arrow storage removes that 
intermediate allocation while retaining independent ownership of the output 
bytes.
   
   ## What changes are included in this PR?
   
   - In ordinary `StringWriter`, reserve the value range with 
`setValueLengthSafe`, obtain the current buffer address after any reallocation, 
copy with `writeToMemory`, and mark the value defined. Check offset arithmetic 
and destination bounds before the raw copy. Heap-backed inputs retain the 
existing setter path.
   - Add byte-level regression coverage for source reuse, 
null/empty/multibyte/invalid UTF-8 values, heap slices, separate data and 
metadata growth, retained C Data exports, allocation-failure cleanup, and 
offset overflow. Register the suite in Linux and macOS CI.
   - Add `CometStringWriterBenchmark`, comparing the old setter against the 
actual writer with explicit off-heap inputs and heap-backed controls, 
preallocated and growing destinations, allocation counters, and output 
validation.
   
   ## How are these changes tested?
   
   - `make core` and full Maven compilation succeeded. Spotless, Scalastyle, 
and `dev/ci/check-suites.py` passed.
   - All 35 tests in `CometStringWriterSuite` and `CometArrowStreamSuite` 
passed in each configuration below. Spark profiles were built with JDK 17; 
Spark 4.x test classes were additionally executed with JDK 21.
   
   | Spark | JDK | Tests |
   |---|---|---:|
   | 3.4.3 | 17 | 35/35 |
   | 3.5.9 | 17 | 35/35 |
   | 4.0.4 | 17, 21 | 35/35 each |
   | 4.1.3 | 17, 21 | 35/35 each |
   
   Focused Maven command (select the corresponding Spark profile):
   
   ```sh
   ./mvnw test -Pspark-4.1 -Dtest=none 
-Dsuites=org.apache.spark.sql.comet.execution.arrow.CometStringWriterSuite,org.apache.spark.sql.comet.execution.arrow.CometArrowStreamSuite
   ```
   
   The full benchmark ran on Spark 4.1.3 / Temurin JDK 17.0.20.1, Linux x86-64, 
AMD EPYC 9V74. Each case uses a 1-second warmup and at least 1 second of 
measured writing. Source construction and initial destination allocation are 
outside the measured region; growth during writing is included.
   
   Preallocated off-heap destinations:
   
   | Value size | Baseline JVM bytes/row | Direct copy JVM bytes/row | 
Conversion throughput vs baseline |
   |---|---:|---:|---:|
   | 16 B | 32 | 0 | 1.09x |
   | 1 KiB | 1,096 | 0 | 1.82x |
   | 64 KiB | 65,608 | 0 | 3.16x |
   | 1 MiB | 1,048,648 | 0 | 1.58x |
   | Varied (0 B–1 MiB) | 209,163.5 | 0 | 3.47x |
   
   Growing off-heap destinations retain 0.8–81 JVM bytes/row of Arrow growth 
bookkeeping, while still removing the payload-sized array. Heap-array and 
heap-slice controls have identical allocation counts and throughput within 
approximately 2% of baseline across the measured cases. These are conversion 
microbenchmark results, not end-to-end query speedups, and the JVM allocation 
counter excludes Arrow's native storage.
   
   A separate JDK 21 JFR allocation run (`--smoke`) attributes sampled baseline 
byte arrays to `UTF8String.getBytes → getByteBuffer → 
ByteBufferStringWriter.setValue`; the direct writer has no sampled 
payload-array allocation stack. The thread-allocation counters above provide 
the per-row measurement.
   
   Reproduce with:
   
   ```sh
   make 
benchmark-org.apache.spark.sql.comet.execution.arrow.CometStringWriterBenchmark
   ```
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to