mxm opened a new pull request, #18348: URL: https://github.com/apache/iceberg/pull/18348
`stream-from-timestamp` tells a new streaming query which snapshot to start from. Until that snapshot exists, the query uses the offset -1, which means "no position yet". Problem 1: The option was applied again on every restart. For example, a job may pass the current time every time it starts. After a restart, the newest snapshot was usually older than that time, so the query fell back to -1 instead of continuing where it stopped. Everything committed while the job was down was skipped (see #10156), and -1 was saved in the checkpoint. Problem 2: When Spark replays a batch that starts at -1, Iceberg looks up the start again with the option's current value. If the job stopped in the middle of such a batch, the restart used a later time, found no snapshot after it, and failed with "Cannot load current offset at snapshot -1". It failed the same way on every further restart. A new checkpoint could hit this as well, because it saved -1 as its starting point. Fix: A query that already has a position continues from it. The option is only used to find the first snapshot. -1 is never saved: the query waits until there is a snapshot to start from, and saves it as its starting point before the first batch is recorded in the checkpoint. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
