ErikBPF commented on issue #5964: URL: https://github.com/apache/datafusion-comet/issues/5964#issuecomment-5691442726
Update: the investigation now includes SQL-level regression coverage, beyond the earlier decoder probe. The parquet-java fix is published as apache/parquet-java#3796, with reproduction tracked in apache/parquet-java#3795. Its current-master port passes 112 targeted tests plus Spotless and package checks. The isolated Spark candidate selects the first physical root descriptor while retaining requested field IDs and suppresses ambiguous predicate pushdown without dropping logical filters. Targeted release-based owner checks pass: 8 duplicate-root tests and 101 filter tests, with scalastyle clean. No Spark PR has been opened yet: JIRA tracking, a current-master port and a normal published parquet dependency pin remain prerequisites. These are bounded experimental and targeted-module results, not released fixes or full-reactor certification. Simultaneous aliases and duplicate nested, encrypted or non-vectorized reads remain outside validated scope. Anomalous Spark values are not adopted as desired behavior. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
