Rich-T-kid commented on PR #24227: URL: https://github.com/apache/datafusion/pull/24227#issuecomment-5873009385
> Is the goal of this PR to improve performance for some usecases? If so, do we have any benchmark results that show the performance improves? @alamb yes & no. The end goal is faster queries & providing users with more control over how their data is read. This initial PR solves the second problem. As for the query speedup, I would have expected a larger speedup for certain queries. As we suspected, and as @yinli-systems pointed out [here](https://github.com/apache/datafusion/issues/24117#issuecomment-5671192732), there is a cutoff for the usefulness of dictionaries. I made a follow-up issue to address this and bridge the performance gaps. On a slightly unrelated note, it will be much easier to track performance changes and gains with this PR, as we can just run the regular benchmark suite with this flag enabled. This means a faster iteration cycle. I plan on going through the list of optimizations pointed out in the issue and then running the benchmarks with the flag enabled vs disabled. for any other discrepancies with performance ill do in in depth profile to see where we may be missing optimizations -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
