mxm commented on code in PR #18123:
URL: https://github.com/apache/iceberg/pull/18123#discussion_r4143807496


##########
flink/v2.3/flink/src/main/java/org/apache/iceberg/flink/sink/IcebergCommitter.java:
##########
@@ -138,13 +141,24 @@ public void 
commit(Collection<CommitRequest<IcebergCommittable>> commitRequests)
     }
 
     IcebergCommittable last = 
commitRequestMap.lastEntry().getValue().getCommittable();
+    // A stateless start has no committed checkpoints; the table's mark would 
drop every commit.
     long maxCommittedCheckpointId =
-        SinkUtil.getMaxCommittedCheckpointId(table, last.jobId(), 
last.operatorId(), branch);
+        isRestored
+            ? SinkUtil.getMaxCommittedCheckpointId(table, last.jobId(), 
last.operatorId(), branch)
+            : SinkUtil.INITIAL_CHECKPOINT_ID;

Review Comment:
   What if we failed here in batch mode and re-ran the attempt? We would then 
double commit.



##########
flink/v2.3/flink/src/test/java/org/apache/iceberg/flink/sink/TestIcebergCommitter.java:
##########
@@ -238,6 +238,57 @@ public void testCommitTxn() throws Exception {
     }
   }
 
+  @TestTemplate
+  public void testCommitTxnAfterStatelessRestart() throws Exception {
+    RowData rowFromPreviousRun = SimpleDataUtil.createRowData(0, "hello0");
+    DataFile dataFileFromPreviousRun =
+        writeDataFile("data-previous-run", 
ImmutableList.of(rowFromPreviousRun));
+
+    try (OneInputStreamOperatorTestHarness<
+            CommittableMessage<IcebergCommittable>, 
CommittableMessage<IcebergCommittable>>
+        preRestartHarness = getTestHarness()) {
+      preRestartHarness.open();
+      processElement(jobId, 5, preRestartHarness, 1, OPERATOR_ID, 
dataFileFromPreviousRun);
+      preRestartHarness.notifyOfCompletedCheckpoint(5);
+    }
+
+    assertSnapshotSize(1);
+    assertMaxCommittedCheckpointId(jobId, 5);
+
+    List<RowData> rows = Lists.newArrayList(rowFromPreviousRun);
+    // A stateless restart opens a fresh operator instance without restoring 
prior state, so
+    // IcebergSink#createCommitter must see an empty 
context.getRestoredCheckpointId().
+    try (OneInputStreamOperatorTestHarness<
+            CommittableMessage<IcebergCommittable>, 
CommittableMessage<IcebergCommittable>>
+        afterRestartHarness = getTestHarness()) {
+      afterRestartHarness.open();
+
+      for (int i = 1; i <= 3; i++) {
+        RowData row = SimpleDataUtil.createRowData(i, "hello" + i);
+        DataFile dataFile = writeDataFile("data-after-restart-" + i, 
ImmutableList.of(row));
+        processElement(jobId, i, afterRestartHarness, 1, OPERATOR_ID, 
dataFile);
+        afterRestartHarness.notifyOfCompletedCheckpoint(i);
+        rows.add(row);
+        assertSnapshotSize(i + 1);
+        assertMaxCommittedCheckpointId(jobId, i);
+        SimpleDataUtil.assertTableRows(table, ImmutableList.copyOf(rows), 
branch);
+      }

Review Comment:
   Can we spell out this loop? Do we need three iterations? I think one is 
sufficient.



##########
flink/v2.3/flink/src/main/java/org/apache/iceberg/flink/sink/IcebergCommitter.java:
##########
@@ -138,13 +141,24 @@ public void 
commit(Collection<CommitRequest<IcebergCommittable>> commitRequests)
     }
 
     IcebergCommittable last = 
commitRequestMap.lastEntry().getValue().getCommittable();
+    // A stateless start has no committed checkpoints; the table's mark would 
drop every commit.
     long maxCommittedCheckpointId =
-        SinkUtil.getMaxCommittedCheckpointId(table, last.jobId(), 
last.operatorId(), branch);
+        isRestored
+            ? SinkUtil.getMaxCommittedCheckpointId(table, last.jobId(), 
last.operatorId(), branch)
+            : SinkUtil.INITIAL_CHECKPOINT_ID;

Review Comment:
   This scenario can only happen when the jobID is reused across runs.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to