tohuya6 opened a new pull request, #25813: URL: https://github.com/apache/datafusion/pull/25813
## Which issue does this PR close? - Closes #25799. ## Rationale for this change Spark `xxhash64` returns different results for the same data depending on whether a value inside a struct or list is dictionary-encoded. The dictionary fast path restarted from the default seed, so earlier arguments and earlier list elements had no effect. ## What changes are included in this PR? Ports the fix from apache/datafusion-comet#5757: - The dictionary fast path is used only when every row has the same running hash. - The fast path starts from that running hash instead of 42. ## What is the testing strategy for this PR? Two unit tests ported from Comet, both fail on main: - `test_dictionary_element_in_list_matches_decoded`: a dictionary inside a list hashes the same as its decoded values. - `test_dictionary_with_nonuniform_seeds_matches_decoded`: rows with different running hashes hash the same as the decoded values. The existing Spark hash unit tests and sqllogictests pass. Both queries from the issue now return matching results for dictionary and plain values. ## Are there any user-facing changes? `xxhash64` now matches Spark for dictionary-encoded values inside structs and lists. No API changes. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
