Guosmilesmile commented on issue #18098: URL: https://github.com/apache/iceberg/issues/18098#issuecomment-5862394054
I’d prefer to keep the Sink V2 behavior consistent with V1, since a stateless restart shouldn’t silently drop data. 1. IcebergFilesCommitter in Sink V1 already handles this case: when the job is not restored from state, maxSnapshot is set to -1. Sink V2 doesn’t have this handling, so the behavior has changed.https://github.com/apache/iceberg/blob/e4bcaa40cbb4de6431584cb2a5816cb626c34ca5/flink/v2.3/flink/src/main/java/org/apache/iceberg/flink/sink/IcebergFilesCommitter.java#L170-L180 2. There are cases where the job can’t be restored from state and has to be restarted directly. If the topology and HA ID stay the same, the Job ID may also remain unchanged. In that case, V2 may use the old dedup information and incorrectly treat new committables as already committed, resulting in silent data loss. 3. For SQL jobs, we can’t directly set or change the UID, so we can’t rely on changing the UID to work around this. @pvary WDYT? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
