jayzhan211 commented on PR #25434: URL: https://github.com/apache/datafusion/pull/25434#issuecomment-5797220880
My concern now is that Sampling adds a lot of machinery: four mutable flags (tried_preallocation, sampled_preallocation, compact_sampled_growth, compaction_headroom) plus magic thresholds (128 * SAMPLES, /32, > 8, the triples/pairs test). It only helps one case: a single non-null Utf8/Binary Column key with 262k or more rows. Sampling is what delivers the mid-cardinality win (mid_20k: 59.1 MB → 26.6 MB). But the low-cardinality win (duplicates 47.7 → 16.4 MB) and the headline 2 MiB-pool fix are already in 57791ea2a. Is it possible to have a much simpler design for low-cardinality builds? 🤔 -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
