WZhuo opened a new pull request, #961: URL: https://github.com/apache/iceberg-cpp/pull/961
HashBytes passed a size_t length straight to the vendored MurmurHash3_x86_32, whose length parameter is a signed 32-bit int. For inputs of 2^31 bytes or more the length wrapped negative, making the hash loop read far out of bounds (and nblocks * 4 signed-overflow UB); for inputs above 2^32 bytes a silently wrong prefix was hashed, breaking the bucket contract against other implementations. Every literal hash path (kDecimal/kString/kUuid/kBinary/kFixed) funnels through this one function, so guard it there: reject lengths above INT32_MAX with ICEBERG_CHECK_OR_DIE before the narrowing cast. The bound matches the limit implied by iceberg-java's byte arrays, so no legal bucket input is rejected. The vendored murmurhash3_internal.cc is left untouched (tagged third-party code) and the now-checked narrowing is made explicit with static_cast<int>. Add BucketUtilsTest.HashBytesRejectsOversizedInput, which fabricates an oversized span over a 1-byte buffer (never dereferenced, so no 2 GiB allocation is needed in CI). Without the guard the test segfaults. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
