cetra3 commented on PR #23565:
URL: https://github.com/apache/datafusion/pull/23565#issuecomment-5863566906

   ## Benchmark bot results after the heuristics commit (head `7effeab5c`, base 
`ed43e69fd`)
   
   I rebased onto `main` and changed the compaction to address the regressions 
from the last round:
   
   - Deduplication samples the first 256 non-inline values. If they contain 
almost no repeats, it falls back to plain `gc()`, so all-distinct data no 
longer pays for hashing.
   - Deduplication now uses a hash table over the views directly, instead of 
rebuilding the array through `GenericByteViewBuilder`.
   
   A/B = main vs this PR, A/A = main vs main. Ratios are branch / base, so 
<1.00 is faster with this PR.
   
   ### spill_views
   
   "Default limits" are the suite's own limits (40M for q01–q03, 96M for 
q04–q05).
   
   | Query | A/B run 1 (default limits) | A/B run 2 (default limits) | A/A 
(default limits) | A/B (q05 at 128M) | A/A (q05 at 128M) |
   | --- | --- | --- | --- | --- | --- |
   | q01 sort, 1 distinct value | 108.0 → 84.3 ms (**0.78x**) | 107.8 → 85.4 ms 
(**0.79x**) | 0.99x | 105.0 → 83.5 ms (**0.80x**) | 1.01x |
   | q02 sort, 1000 distinct values | 102.2 → 94.4 ms (**0.92x**) | 103.3 → 
95.5 ms (**0.92x**) | 0.99x | 101.6 → 93.0 ms (**0.92x**) | 1.01x |
   | q03 sort, BinaryView, 1000 distinct | 100.4 → 94.2 ms (**0.94x**) | 101.8 
→ 95.2 ms (**0.94x**) | 0.98x | 99.2 → 93.3 ms (**0.94x**) | 1.01x |
   | q04 sort, all distinct | 1.00x | 1.02x | 0.99x | 1.00x | 1.01x |
   | q05 GROUP BY, all distinct | 0.99x | 1.01x | 0.98x | 1.01x | 1.00x |
   
   q05 now passes at the default 96M limit (2 of 2 runs; it previously failed 
on the branch side in 5 of 5).
   
   ### sort_tpch, 512M limit, 4 partitions
   
   Q1–Q7 and Q10 show no change; Q3 fails on both sides, as before.
   
   | Query | A/B run 1 | A/B run 2 | A/A |
   | --- | --- | --- | --- |
   | Q8 (`l_comment`) | 1.01x | 1.00x | 0.99x |
   | Q9 (`l_comment`) | 1.00x | 1.00x | 1.00x |
   | Q11 (`l_comment`) | 1.01x | 1.00x | 0.99x |
   | Total | 1.00x | 1.00x | 0.99x |
   
   ### Compared with the previous round
   
   | | Before | Now |
   | --- | --- | --- |
   | q01 | 0.84–0.86x | **0.78–0.80x** |
   | q02 / q03 | ~1.00x | **0.92x / 0.94x** |
   | q04 | 1.38–1.45x | 1.00–1.02x |
   | q05 at 96M | OOM on the branch side, 5 of 5 runs | passes |
   | sort_tpch Q11 | 1.08–1.11x | 1.00–1.01x |
   
   The repeated-value queries are faster, and the all-distinct and `sort_tpch` 
cases are now within A/A noise. The reduction in spilled bytes is covered by 
`test_gc_copies_repeated_values_once` (the bot's peak-spill figure is sampled 
once per second and too coarse to show it).
   
   @adriangb @kumarUjjawal could you take another look?
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to