rdblue commented on code in PR #16025: URL: https://github.com/apache/iceberg/pull/16025#discussion_r4161566634
########## format/spec.md: ########## @@ -977,18 +1116,18 @@ For other optional snapshot summary fields, see [Appendix F](#optional-snapshot- Data and delete files for a snapshot can be stored in more than one manifest. This enables: * Appends can add a new manifest to minimize the amount of data written, instead of adding new records by rewriting and appending to an existing manifest. (This is called a “fast append”.) -* Tables can use multiple partition specs. A table’s partition configuration can evolve if, for example, its data volume changes. Each manifest uses a single partition spec, and queries do not need to change because partition filters are derived from data predicates. +* Tables can use multiple partition specs. A table’s partition configuration can evolve if, for example, its data volume changes. Queries do not need to change because partition filters are derived from data predicates. In v1-v3, each manifest uses a single partition spec. * Large tables can be split across multiple manifests so that implementations can parallelize job planning or reduce the cost of rewriting a manifest. -Manifests for a snapshot are tracked by a manifest list. +Manifests for a snapshot are tracked by the snapshot root. Valid snapshots are stored as a list in table metadata. For serialization, see Appendix C. #### Snapshot Row IDs A snapshot's `first-row-id` is assigned to the table's current `next-row-id` on each commit attempt. If a commit is retried, the `first-row-id` must be reassigned based on the table's current `next-row-id`. The `first-row-id` field is required even if a commit does not assign any ID space. -The snapshot's `first-row-id` is the starting `first_row_id` assigned to manifests in the snapshot's manifest list. +The snapshot's `first-row-id` is the starting `first_row_id` assigned to manifests in the snapshot root. In v4, this includes data files in the root manifest. Review Comment: I've suggested this a few times, but I think it may be generally more clear if we use "snapshot root file" so that people don't confuse "snapshot root" with the snapshot metadata in metadata.json. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
