WZhuo opened a new pull request, #961:
URL: https://github.com/apache/iceberg-cpp/pull/961

   HashBytes passed a size_t length straight to the vendored 
MurmurHash3_x86_32, whose length parameter is a signed 32-bit int. For inputs 
of 2^31 bytes or more the length wrapped negative, making the hash loop read 
far out of bounds (and nblocks * 4 signed-overflow UB); for inputs above 2^32 
bytes a silently wrong prefix was hashed, breaking the bucket contract against 
other implementations.
   
   Every literal hash path (kDecimal/kString/kUuid/kBinary/kFixed) funnels 
through this one function, so guard it there: reject lengths above INT32_MAX 
with ICEBERG_CHECK_OR_DIE before the narrowing cast. The bound matches the 
limit implied by iceberg-java's byte arrays, so no legal bucket input is 
rejected. The vendored murmurhash3_internal.cc is left untouched (tagged 
third-party code) and the now-checked narrowing is made explicit with 
static_cast<int>.
   
   Add BucketUtilsTest.HashBytesRejectsOversizedInput, which fabricates an 
oversized span over a 1-byte buffer (never dereferenced, so no 2 GiB allocation 
is needed in CI). Without the guard the test segfaults.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to