dwsmith1983 commented on PR #5365: URL: https://github.com/apache/datafusion-comet/pull/5365#issuecomment-5891089529
> Could the physical-to-requested leaf pairing choose the timestamp policy from the requested type in both directions? Yes, in 4b7ad02. The walk that pairs each physical leaf with its requested type now records the timezone-free leaves read as `TIMESTAMP` and the timezone-carrying ones read as `TIMESTAMP_NTZ`, and the policy follows that. A metadata-free `TIMESTAMP(MICROS, false)` file read as `TIMESTAMP` now fails under `EXCEPTION` like Spark (the Delta test also checks that stock Spark throws `SparkUpgradeException` for the same table), and reads the value unchanged under `CORRECTED`. Under `LEGACY` the native scan refuses that file rather than rebasing it: with no `org.apache.spark.timeZone` in the footer, Spark rebases in the JVM's default time zone, which the native side cannot reproduce. That is the same refusal adjusted files without a recorded writer zone already get. A file that records a UTC writer zone is rebased exactly. The regression you suggested is `tz_free_int64_timestamps_requested_as_ltz_follow_the_datetime_spec` in `parquet_exec.rs`, plus the Delta test "non-Spark tz-free INT64 timestamps read as TIMESTAMP follow the datetime read mode". -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
