ion-elgreco opened a new issue, #25775: URL: https://github.com/apache/datafusion/issues/25775
### Is your feature request related to a problem or challenge? In delta-rs, we replay the Delta log at planning time to find the files to read. We then put these files in FileScanConfig.file_groups, so the file list is fixed before execution starts. It would be useful if DataSourceExec could have a child node that supplies its files while the query runs. The child would be a metadata scan, for example the Delta (or iceberg) log replay. Filters that are known only at runtime, such as join dynamic filters could then be pushed into that child and skip whole files. Our use case is our copy-on-write MERGE. It rewrites whole files, so it must **read every row of each target file** that it keeps. But we want to skip target files at runtime with statistics coming from the source, which are known only after the source is read, this might come from a source which can be exclusively streamed once. PR https://github.com/delta-io/delta-rs/pull/4790 adds this to delta-rs. When the target scan starts, it runs file skipping again (albeit with an already materialized snapshot) and builds a copy of the planned DataSourceExec with only the kept files. @alamb do you have any thoughts on this? Or know otherwise who best to discuss this with 😄 ### Describe the solution you'd like _No response_ ### Describe alternatives you've considered The PR listed above is the alternative. We rewrite the parquet nodes at execution time. ### Additional context _No response_ -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
