singhpk234 commented on code in PR #18348:
URL: https://github.com/apache/iceberg/pull/18348#discussion_r4209233407
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
public void prepareForTriggerAvailableNow() {
LOG.info("The streaming query reports to use Trigger.AvailableNow");
- lastOffsetForTriggerAvailableNow =
+ StreamingOffset lastOffset =
(StreamingOffset) latestOffset(initialOffset,
ReadLimit.allAvailable());
+ // START_OFFSET means that no snapshot matched stream-from-timestamp. A
new stream has nothing
+ // to read in this run, but a resumed stream continues from its offset and
needs an actual cap.
+ this.noSnapshotMatchedForTriggerAvailableNow =
StreamingOffset.START_OFFSET.equals(lastOffset);
+ this.lastOffsetForTriggerAvailableNow =
+ noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() :
lastOffset;
- LOG.info("lastOffset for Trigger.AvailableNow is {}",
lastOffsetForTriggerAvailableNow.json());
+ LOG.info("lastOffset for Trigger.AvailableNow is {}",
lastOffsetForTriggerAvailableNow);
Review Comment:
why did we remove `.json()` here ?
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/SparkMicroBatchStream.java:
##########
@@ -244,15 +253,32 @@ public ReadLimit getDefaultReadLimit() {
public void prepareForTriggerAvailableNow() {
LOG.info("The streaming query reports to use Trigger.AvailableNow");
- lastOffsetForTriggerAvailableNow =
+ StreamingOffset lastOffset =
(StreamingOffset) latestOffset(initialOffset,
ReadLimit.allAvailable());
+ // START_OFFSET means that no snapshot matched stream-from-timestamp. A
new stream has nothing
+ // to read in this run, but a resumed stream continues from its offset and
needs an actual cap.
+ this.noSnapshotMatchedForTriggerAvailableNow =
StreamingOffset.START_OFFSET.equals(lastOffset);
+ this.lastOffsetForTriggerAvailableNow =
+ noSnapshotMatchedForTriggerAvailableNow ? latestAppendOffset() :
lastOffset;
- LOG.info("lastOffset for Trigger.AvailableNow is {}",
lastOffsetForTriggerAvailableNow.json());
+ LOG.info("lastOffset for Trigger.AvailableNow is {}",
lastOffsetForTriggerAvailableNow);
// Reset planner so it gets recreated with the cap on next call
if (planner != null) {
planner.stop();
planner = null;
}
}
+
+ // The planners only read append snapshots. A cap at any other snapshot is
never reached.
+ private StreamingOffset latestAppendOffset() {
+ for (Snapshot snapshot : SnapshotUtil.currentAncestors(table)) {
+ if (DataOperations.APPEND.equals(snapshot.operation())) {
+ return new StreamingOffset(
+ snapshot.snapshotId(), MicroBatchUtils.addedFilesCount(table,
snapshot), false);
+ }
+ }
+
+ return null;
+ }
Review Comment:
I thought we just check the addedDataFiles (check MicrobatchUtils#), as
there were discussion on mode to allow operations on non append snapshots ...
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/source/AsyncSparkMicroBatchPlanner.java:
##########
@@ -252,7 +252,9 @@ public synchronized StreamingOffset
latestOffset(StreamingOffset startOffset, Re
return StreamingOffset.START_OFFSET;
}
- if (table().currentSnapshot().timestampMillis() <
readConf().streamFromTimestamp()) {
+ // Only a new stream starts from the timestamp. A resumed stream continues
from its offset.
+ if (startOffset.equals(StreamingOffset.START_OFFSET)
+ && table().currentSnapshot().timestampMillis() <
readConf().streamFromTimestamp()) {
Review Comment:
can you please elaborate this case more ... is this protecting the stream
resume but the stream from timestamp is leading to skip the snapshot / missing
data ?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]