cetra3 commented on PR #23565: URL: https://github.com/apache/datafusion/pull/23565#issuecomment-5863566906
## Benchmark bot results after the heuristics commit (head `7effeab5c`, base `ed43e69fd`) I rebased onto `main` and changed the compaction to address the regressions from the last round: - Deduplication samples the first 256 non-inline values. If they contain almost no repeats, it falls back to plain `gc()`, so all-distinct data no longer pays for hashing. - Deduplication now uses a hash table over the views directly, instead of rebuilding the array through `GenericByteViewBuilder`. A/B = main vs this PR, A/A = main vs main. Ratios are branch / base, so <1.00 is faster with this PR. ### spill_views "Default limits" are the suite's own limits (40M for q01–q03, 96M for q04–q05). | Query | A/B run 1 (default limits) | A/B run 2 (default limits) | A/A (default limits) | A/B (q05 at 128M) | A/A (q05 at 128M) | | --- | --- | --- | --- | --- | --- | | q01 sort, 1 distinct value | 108.0 → 84.3 ms (**0.78x**) | 107.8 → 85.4 ms (**0.79x**) | 0.99x | 105.0 → 83.5 ms (**0.80x**) | 1.01x | | q02 sort, 1000 distinct values | 102.2 → 94.4 ms (**0.92x**) | 103.3 → 95.5 ms (**0.92x**) | 0.99x | 101.6 → 93.0 ms (**0.92x**) | 1.01x | | q03 sort, BinaryView, 1000 distinct | 100.4 → 94.2 ms (**0.94x**) | 101.8 → 95.2 ms (**0.94x**) | 0.98x | 99.2 → 93.3 ms (**0.94x**) | 1.01x | | q04 sort, all distinct | 1.00x | 1.02x | 0.99x | 1.00x | 1.01x | | q05 GROUP BY, all distinct | 0.99x | 1.01x | 0.98x | 1.01x | 1.00x | q05 now passes at the default 96M limit (2 of 2 runs; it previously failed on the branch side in 5 of 5). ### sort_tpch, 512M limit, 4 partitions Q1–Q7 and Q10 show no change; Q3 fails on both sides, as before. | Query | A/B run 1 | A/B run 2 | A/A | | --- | --- | --- | --- | | Q8 (`l_comment`) | 1.01x | 1.00x | 0.99x | | Q9 (`l_comment`) | 1.00x | 1.00x | 1.00x | | Q11 (`l_comment`) | 1.01x | 1.00x | 0.99x | | Total | 1.00x | 1.00x | 0.99x | ### Compared with the previous round | | Before | Now | | --- | --- | --- | | q01 | 0.84–0.86x | **0.78–0.80x** | | q02 / q03 | ~1.00x | **0.92x / 0.94x** | | q04 | 1.38–1.45x | 1.00–1.02x | | q05 at 96M | OOM on the branch side, 5 of 5 runs | passes | | sort_tpch Q11 | 1.08–1.11x | 1.00–1.01x | The repeated-value queries are faster, and the all-distinct and `sort_tpch` cases are now within A/A noise. The reduction in spilled bytes is covered by `test_gc_copies_repeated_values_once` (the bot's peak-spill figure is sampled once per second and too coarse to show it). @adriangb @kumarUjjawal could you take another look? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
