Rich-T-kid commented on PR #24227:
URL: https://github.com/apache/datafusion/pull/24227#issuecomment-5873009385

   > Is the goal of this PR to improve performance for some usecases? If so, do 
we have any benchmark results that show the performance improves?
   
   @alamb yes & no. The end goal is faster queries & providing users with more 
control over how their data is read. This initial PR solves the second problem. 
As for the query speedup, I would have expected a larger speedup for certain 
queries. As we suspected, and as @yinli-systems pointed out 
[here](https://github.com/apache/datafusion/issues/24117#issuecomment-5671192732),
 there is a cutoff for the usefulness of dictionaries. I made a follow-up issue 
to address this and bridge the performance gaps.
   
   On a slightly unrelated note, it will be much easier to track performance 
changes and gains with this PR, as we can just run the regular benchmark suite 
with this flag enabled. This means a faster iteration cycle.
   
   I plan on going through the list of optimizations pointed out in the issue 
and then running the benchmarks with the flag enabled vs disabled. for any 
other discrepancies with performance ill do in in depth profile to see where we 
may be missing optimizations


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to