RussellSpitzer commented on issue #18309: URL: https://github.com/apache/iceberg/issues/18309#issuecomment-5917381323
The purpose of these requirements is to avoid duplicate metadata by reusing existing structures rather than assigning new IDs. Given that context, it seems clear what should be considered equivalent, and I’m not sure how an implementation would reasonably reach a different conclusion. A schema’s serialized fields determine whether it is the same schema. Different field IDs or names would refer to different columns, so that schema could not be reused. Likewise, an equivalent partition field is defined by its source IDs, transform definition, and name. Schema IDs, partition spec IDs, and partition field IDs necessarily cannot participate in equivalence because equivalence is being determined to decide whether those IDs should be reused. Is there a specific pair of schemas or partition fields for which two conforming implementations could reasonably reach different conclusions? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
