andygrove opened a new pull request, #6266:
URL: https://github.com/apache/datafusion-comet/pull/6266

   Backport of #6116 to `branch-1.1`.
   
   Cherry-picked from `863f11de5eb78bacf11737032cdbcf496ad520b5` without 
conflicts. The ten files it changes are identical on `branch-1.1` and on `main` 
just before #6116, so the diff is byte-identical to upstream.
   
   ## Which issue does this PR close?
   
   Closes #5936 on `branch-1.1`. #6116 already closed it on `main`.
   
   ## Rationale for this change
   
   #6116 was on the list of fixes to merge before cutting 1.1.0 in #5327, but 
it merged after `branch-1.1` was cut at 36ab57c68. Without it, 1.1.0 ships the 
divergence from Spark described in #5936. The native scan checks for missing 
field ids only when `spark.sql.parquet.fieldId.read.enabled` is on, and only 
over root fields. So a read schema with ids over a file without ids returns 
rows where Spark raises, and a file whose ids sit only on nested fields is 
rejected where Spark reads it.
   
   ## What changes are included in this PR?
   
   The fix is the original one, so see #6116 for the details. No adaptations 
were needed.
   
   ## How are these changes tested?
   
   The original PR's tests, run locally on `branch-1.1` with the default 
profile (Spark 4.1, Scala 2.13, JDK 17):
   
   - Rust: the two new tests in `eager_page_index_reader_factory.rs`, 
`contains_field_ids_sees_ids_on_any_node` and 
`get_metadata_refuses_a_file_without_ids_only_when_required`, pass along with 
the rest of the core crate's `parquet` module, 226 tests in all. The new 
`parquet_external_spark_error_keeps_its_type` passes with the rest of 
`datafusion-comet-jni-bridge` (29 tests), and `datafusion-comet-common` passes 
(58 tests).
   - Scala: the six tests in `ParquetReadV1Suite` that #6116 added or changed 
pass, as does the existing `nested field ids resolve by id below struct, list 
and map, not by position`.
   - `cargo fmt --all -- --check` and `cargo clippy --all-targets --workspace 
-- -D warnings` pass on rustc 1.98.1.
   
   Spark's own `ParquetFieldIdIOSuite` covers this path in the Spark SQL job, 
which a pull request against `branch-1.1` runs only when labeled, so this 
carries `run-spark-4.1-tests` as #6116 did.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to