singhpk234 commented on code in PR #18348:
URL: https://github.com/apache/iceberg/pull/18348#discussion_r4209233407


##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);

Review Comment:
   why did we remove `.json()` here ? 



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
   public void prepareForTriggerAvailableNow() {
     LOG.info("The streaming query reports to use Trigger.AvailableNow");
 
-    lastOffsetForTriggerAvailableNow =
+    StreamingOffset lastOffset =
         (StreamingOffset) latestOffset(initialOffset, 
ReadLimit.allAvailable());
+    // START_OFFSET means that no snapshot matched stream-from-timestamp. A 
new stream has nothing
+    // to read in this run, but a resumed stream continues from its offset and 
needs an actual cap.
+    this.noSnapshotMatchedForTriggerAvailableNow = 
StreamingOffset.START_OFFSET.equals(lastOffset);
+    this.lastOffsetForTriggerAvailableNow =
+        noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() : 
lastOffset;
 
-    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow.json());
+    LOG.info("lastOffset for Trigger.AvailableNow is {}", 
lastOffsetForTriggerAvailableNow);
 
     // Reset planner so it gets recreated with the cap on next call
     if (planner != null) {
       planner.stop();
       planner = null;
     }
   }
+
+  // The planners only read append snapshots. A cap at any other snapshot is 
never reached.
+  private StreamingOffset latestAppendOffset() {
+    for (Snapshot snapshot : SnapshotUtil.currentAncestors(table)) {
+      if (DataOperations.APPEND.equals(snapshot.operation())) {
+        return new StreamingOffset(
+            snapshot.snapshotId(), MicroBatchUtils.addedFilesCount(table, 
snapshot), false);
+      }
+    }
+
+    return null;
+  }

Review Comment:
   I thought we just check the addedDataFiles (check MicrobatchUtils#), as 
there were discussion on mode to allow operations on non append snapshots ... 



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:
##########
@@ -252,7 +252,9 @@ public synchronized StreamingOffset 
latestOffset(StreamingOffset startOffset, Re
       return StreamingOffset.START_OFFSET;
     }
 
-    if (table().currentSnapshot().timestampMillis() < 
readConf().streamFromTimestamp()) {
+    // Only a new stream starts from the timestamp. A resumed stream continues 
from its offset.
+    if (startOffset.equals(StreamingOffset.START_OFFSET)
+        && table().currentSnapshot().timestampMillis() < 
readConf().streamFromTimestamp()) {

Review Comment:
   can you please elaborate this case more ... is this protecting the stream 
resume but the stream from timestamp is leading to skip the snapshot / missing 
data ? 



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to