0lai0 commented on PR #5932: URL: https://github.com/apache/datafusion-comet/pull/5932#issuecomment-5895475227
Thanks @andygrove. You are right, and both directions were broken. With a hint present convert_field copies its metadata and never derives the extension from the logical type, so the marker the check read was the hint's, not the file's. Fixed in EagerPageIndexReaderFactory, as you suggested: `with_reconciled_variant_markers `takes the hinted schema and the annotation-only schema (`parquet_to_arrow_schema` with and without the key-value metadata) and moves the markers to match the annotations, leaving everything else the hint decides alone. A file with no hint returns after one scan of the key-value metadata. Two tests, one per direction, each writing the data with one schema and the hint arrow-rs writes for another. Both fail with the reconciliation disabled. Also worth flagging: `probe_variant_annotation` built its own `ParquetSource` and never installed the reader factory, so no Variant test exercised the footer rewrites and none of them could have caught this. It installs the factory now. `cargo test --lib`: 530 passed. `dev/local-ci.sh spark sql_core-1` (4.1.3): 12835 passed, 0 failed. On the 4.2 diff: will do. It is not on this branch yet — #4950 landed it after my last merge — so I will merge main and regenerate it without the `IgnoreComet`. Waiting for #6398 first, since until it lands the 4.2 run stops in "Pre-compile Spark Test classes" and tells us nothing about the fix. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
