mxm commented on code in PR #18348:
URL: https://github.com/apache/iceberg/pull/18348#discussion_r4217750149


##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);

Review Comment:
   Exactly, the toString() method is also better suited for log output than the 
json and contains the same info.



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);
 
     // Reset planner so it gets recreated with the cap on next call
     if (planner != null) {
       planner.stop();
       planner = null;
     }
   }
+
+  // The planners only read append snapshots. A cap at any other snapshot is 
never reached.
+  private StreamingOffset latestAppendOffset() {
+    for (Snapshot snapshot : SnapshotUtil.currentAncestors(table)) {
+      if (DataOperations.APPEND.equals(snapshot.operation())) {
+        return new StreamingOffset(
+            snapshot.snapshotId(), MicroBatchUtils.addedFilesCount(table, 
snapshot), false);
+      }
+    }
+
+    return null;
+  }

Review Comment:
   I thought about this, but it wasn't straightforward to reuse 
`shouldProcess`, so I added a test instead. However, I agree with the concern.
   
   I've moved the lookup into the planner, so it uses `shouldProcess`. As 
Huaxin mentioned, with `streaming-skip-delete-snapshots=false`, an AvailableNow 
run in which no snapshot matches the timestamp now fails at start if the newest 
snapshot is a delete, also for a new stream that wouldn't read it. I'm leaning 
towards accepting that to keep the code consistent.
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to