David-Banquet commented on issue #2749:
URL: https://github.com/apache/iceberg-rust/issues/2749#issuecomment-5996168019

   @VABVAT are you still working on this? #2788 was closed by the stale bot 
before anyone reviewed it.
   
   A note for whoever picks it up. The Java Glue catalog also writes the 
columns of older schemas, with `iceberg.field.current=false`, but it dedupes 
them by name and the first one seen wins, so a field present in several schemas 
shows up once. The current schema is written first, so the current definition 
is the one kept (`toColumns` and `addColumnWithDedupe` in 
[`IcebergToGlueConverter`](https://github.com/apache/iceberg/blob/main/aws/src/main/java/org/apache/iceberg/aws/glue/IcebergToGlueConverter.java)).
   
   `GlueSchemaBuilder::from_iceberg` visits the same schemas in the same order 
but has no dedupe, which is where the duplicates come from. Skipping historical 
schemas, as #2788 did, also removes the columns that were dropped from the 
table, which Java keeps in Glue with `iceberg.field.current=false`. Adding the 
same name-based dedupe would fix the duplicates and keep the output aligned 
with Java.
   
   If you don't plan to continue, I can open a PR with that approach.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to