David-Banquet commented on issue #2749: URL: https://github.com/apache/iceberg-rust/issues/2749#issuecomment-5996168019
@VABVAT are you still working on this? #2788 was closed by the stale bot before anyone reviewed it. A note for whoever picks it up. The Java Glue catalog also writes the columns of older schemas, with `iceberg.field.current=false`, but it dedupes them by name and the first one seen wins, so a field present in several schemas shows up once. The current schema is written first, so the current definition is the one kept (`toColumns` and `addColumnWithDedupe` in [`IcebergToGlueConverter`](https://github.com/apache/iceberg/blob/main/aws/src/main/java/org/apache/iceberg/aws/glue/IcebergToGlueConverter.java)). `GlueSchemaBuilder::from_iceberg` visits the same schemas in the same order but has no dedupe, which is where the duplicates come from. Skipping historical schemas, as #2788 did, also removes the columns that were dropped from the table, which Java keeps in Glue with `iceberg.field.current=false`. Adding the same name-based dedupe would fix the duplicates and keep the output aligned with Java. If you don't plan to continue, I can open a PR with that approach. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
