mxm opened a new pull request, #18348:
URL: https://github.com/apache/iceberg/pull/18348

   `stream-from-timestamp` tells a new streaming query which snapshot to start 
from. Until that snapshot exists, the query uses the offset -1, which means "no 
position yet".
   
   Problem 1: The option was applied again on every restart. For example, a job 
may pass the current time every time it starts. After a restart, the newest 
snapshot was usually older than that time, so the query fell back to -1 instead 
of continuing where it stopped. Everything committed while the job was down was 
skipped (see #10156), and -1 was saved in the checkpoint.
   
   Problem 2: When Spark replays a batch that starts at -1, Iceberg looks up 
the start again with the option's current value. If the job stopped in the 
middle of such a batch, the restart used a later time, found no snapshot after 
it, and failed with "Cannot load current offset at snapshot -1". It failed the 
same way on every further restart. A new checkpoint could hit this as well, 
because it saved -1 as its starting point.
   
   Fix: A query that already has a position continues from it. The option is 
only used to find the first snapshot. -1 is never saved: the query waits until 
there is a snapshot to start from, and saves it as its starting point before 
the first batch is recorded in the checkpoint.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to