jayzhan211 commented on PR #25434:
URL: https://github.com/apache/datafusion/pull/25434#issuecomment-5797220880

   My concern now is that 
   
   Sampling adds a lot of machinery: four mutable flags (tried_preallocation, 
sampled_preallocation, compact_sampled_growth, compaction_headroom) plus magic 
thresholds (128 * SAMPLES, /32, > 8, the triples/pairs test). It only helps one 
case: a single non-null Utf8/Binary Column key with 262k or more rows.
   
   Sampling is what delivers the mid-cardinality win (mid_20k: 59.1 MB → 26.6 
MB). But the low-cardinality win (duplicates 47.7 → 16.4 MB) and the headline 2 
MiB-pool fix are already in 57791ea2a.
   
   Is it possible to have a much simpler design for low-cardinality builds? 🤔


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to