dramaticlly commented on code in PR #18348:
URL: https://github.com/apache/iceberg/pull/18348#discussion_r4210765678


##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########


Review Comment:
   do you need to update this?



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SyncSparkMicroBatchPlanner.java:
##########
@@ -118,7 +118,9 @@ public StreamingOffset latestOffset(StreamingOffset 
startOffset, ReadLimit limit
       return StreamingOffset.START_OFFSET;
     }
 
-    if (table().currentSnapshot().timestampMillis() < fromTimestamp) {
+    // Only a new stream starts from the timestamp. A resumed stream continues 
from its offset.
+    if (startOffset.equals(StreamingOffset.START_OFFSET)
+        && table().currentSnapshot().timestampMillis() < fromTimestamp) {

Review Comment:
   not sure if this worth reusing MicroBatchUtils.determineStartingOffset ?



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);

Review Comment:
   I think now lastOffsetForTriggerAvailableNow is nullable from return of 
latestAppendOffset()



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);
 
     // Reset planner so it gets recreated with the cap on next call
     if (planner != null) {
       planner.stop();
       planner = null;
     }
   }
+
+  // The planners only read append snapshots. A cap at any other snapshot is 
never reached.
+  private StreamingOffset latestAppendOffset() {
+    for (Snapshot snapshot : SnapshotUtil.currentAncestors(table)) {
+      if (DataOperations.APPEND.equals(snapshot.operation())) {
+        return new StreamingOffset(
+            snapshot.snapshotId(), MicroBatchUtils.addedFilesCount(table, 
snapshot), false);
+      }
+    }
+
+    return null;

Review Comment:
   wondering if we need UT to cover the OVERWRITE data operations which this 
return null and append is expired ?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to