tohuya6 opened a new pull request, #25813:
URL: https://github.com/apache/datafusion/pull/25813

   ## Which issue does this PR close?
   
   - Closes #25799.
   
   ## Rationale for this change
   
   Spark `xxhash64` returns different results for the same data depending on 
whether a value inside a struct or list is dictionary-encoded. The dictionary 
fast path restarted from the default seed, so earlier arguments and earlier 
list elements had no effect.
   
   ## What changes are included in this PR?
   
   Ports the fix from apache/datafusion-comet#5757:
   
   - The dictionary fast path is used only when every row has the same running 
hash.
   - The fast path starts from that running hash instead of 42.
   
   ## What is the testing strategy for this PR?
   
   Two unit tests ported from Comet, both fail on main:
   
   - `test_dictionary_element_in_list_matches_decoded`: a dictionary inside a 
list hashes the same as its decoded values.
   - `test_dictionary_with_nonuniform_seeds_matches_decoded`: rows with 
different running hashes hash the same as the decoded values.
   
   The existing Spark hash unit tests and sqllogictests pass. Both queries from 
the issue now return matching results for dictionary and plain values.
   
   ## Are there any user-facing changes?
   
   `xxhash64` now matches Spark for dictionary-encoded values inside structs 
and lists. No API changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to