amogh-jahagirdar commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4199720498
##########
format/spec.md:
##########
@@ -731,39 +744,151 @@ The `data_file` struct consists of the following fields:
| _optional_ | _optional_ | _optional_ | **`110 null_value_counts`**
| `map<121: int, 122: long>` |
Map from column id to number of null values in the column |
| _optional_ | _optional_ | _optional_ | **`137 nan_value_counts`**
| `map<138: int, 139: long>` |
Map from column id to number of NaN values in the column |
| _optional_ | _optional_ | | ~~**`111 distinct_counts`**~~
| `map<123: int, 124: long>` |
**Deprecated. Do not write.** |
- | _optional_ | _optional_ | _optional_ | **`125 lower_bounds`**
| `map<126: int, 127: binary>` |
Map from column id to lower bound in the column serialized as binary [1]. Each
value must be less than or equal to all non-null, non-NaN values in the column
for the file [2] |
- | _optional_ | _optional_ | _optional_ | **`128 upper_bounds`**
| `map<129: int, 130: binary>` |
Map from column id to upper bound in the column serialized as binary [1]. Each
value must be greater than or equal to all non-null, non-Nan values in the
column for the file [2] |
+ | _optional_ | _optional_ | _optional_ | **`125 lower_bounds`**
| `map<126: int, 127: binary>` |
Map from column id to lower bound in the column serialized as binary. Each
value must be less than or equal to all non-null, non-NaN values in the column
for the file. See [Field-level Metrics and
Statistics](#field-level-metrics-and-statistics) |
+ | _optional_ | _optional_ | _optional_ | **`128 upper_bounds`**
| `map<129: int, 130: binary>` |
Map from column id to upper bound in the column serialized as binary. Each
value must be greater than or equal to all non-null, non-Nan values in the
column for the file. See [Field-level Metrics and
Statistics](#field-level-metrics-and-statistics) |
| _optional_ | _optional_ | _optional_ | **`131 key_metadata`**
| `binary` |
Implementation-specific key metadata for encryption |
| _optional_ | _optional_ | _optional_ | **`132 split_offsets`**
| `list<133: long>` |
Split offsets for the data file. For example, all row group offsets in a
Parquet file. Must be sorted ascending |
| | _optional_ | _optional_ | **`135 equality_ids`**
| `list<136: int>` |
Field ids used to determine row equality in equality delete files. Required
when `content=2` and should be null otherwise. Fields with ids listed in this
column must be present in the delete file |
- | _optional_ | _optional_ | _optional_ | **`140 sort_order_id`**
| `int` |
ID representing sort order for this file [3]. |
+ | _optional_ | _optional_ | _optional_ | **`140 sort_order_id`**
| `int` |
ID representing sort order for this file [1]. |
| | | _optional_ | **`142 first_row_id`**
| `long` |
The `_row_id` for the first row in the data file. See [First Row ID
Inheritance](#first-row-id-inheritance) |
- | | _optional_ | _optional_ | **`143 referenced_data_file`**
| `string` |
Fully qualified location (URI with FS scheme) of a data file that all deletes
reference [4] |
- | | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
- | | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
+ | | _optional_ | _optional_ | **`143 referenced_data_file`**
| `string` |
Fully qualified location (URI with FS scheme) of a data file that all deletes
reference [2] |
+ | | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [3] |
+ | | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [3] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 2. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 3. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 4. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ The `tracked_file` struct has the following fields:
+
+ | On write | Field id | Name | Type
| Description |
+
|------------|----------|--------------------------|-------------------------------------------------------|-------------|
+ | _required_ | 134 | **`content_type`** | `int` (0: DATA, 3:
DATA_MANIFEST, 4: DELETE_MANIFEST) | Type of content stored in the entry. |
+ | _required_ | 157 | **`format_version`** | `int` (0: PRE-V4, 4:
V4) | Writer format version. |
+ | _required_ | 147 | **`tracking`** | `tracking` struct
| Tracking metadata like status, snapshot ID,
and sequence number. See tracking struct below. |
+ | _required_ | 100 | **`location`** | `string`
| Location of the file. |
+ | _required_ | 101 | **`file_format`** | `string`
| String file format name: `avro`, `orc`, or
`parquet` |
+ | _required_ | 104 | **`file_size_in_bytes`** | `long`
| Total file size in bytes. |
+ | _required_ | 103 | **`record_count`** | `long`
| Number of records in this file. |
+ | _optional_ | 131 | **`key_metadata`** | `binary`
| Implementation-specific key metadata for
encryption. |
+ | _optional_ | 132 | **`split_offsets`** | `list<133: long>`
| Split offsets for the data file. Must be
sorted ascending. |
+ | _optional_ | 141 | **`spec_id`** | `int`
| ID of the partition spec used to partition
the file; null if unpartitioned |
+ | _optional_ | 102 | **`partition`** | `struct<...>`
| Partition data tuple for the file; null if
unpartitioned. |
+ | _optional_ | 140 | **`sort_order_id`** | `int`
| ID representing sort order for this file. If
missing or unknown, the order is assumed to be unsorted. |
+ | _optional_ | 146 | **`content_stats`** | `content_stats`
struct | Field-level stats. See [Content
Stats](#content-stats). |
+ | _optional_ | 150 | **`manifest_info`** | `manifest_info`
struct | Manifest-specific stats. See [Manifest
Info](#manifest-info) |
+ | _optional_ | 148 | **`deletion_vector`** | `deletion_vector`
struct | Row-level deletion vector for a data
file. |
+ | _optional_ | 158 | **`column_files`** | `list<159:
column_file>` | Column files associated with this
file. |
+
+ The `tracking` struct has the following fields:
+
+ | On write | Field id | Name | Type
| Description |
+
|------------|----------|-------------------------------|---------------------------------------------------------------------|-------------|
+ | _required_ | 0 | **`status`** | `int` (0:
EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | Used to track
additions, deletions, replacements, and modifications. |
+ | _optional_ | 1 | **`snapshot_id`** | `long`
| Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | _optional_ | 5 | **`dv_snapshot_id`** | `long`
| Snapshot ID where the
deletion vector or manifest DV last changed. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | _optional_ | 8 | **`column_file_snapshot_id`** | `long`
| Snapshot ID where the
latest column file was added. |
+ | _optional_ | 3 | **`sequence_number`** | `long`
| Data sequence number of
the file. Inherited when null. See [Sequence Number
Inheritance](#sequence-number-inheritance). |
+ | _optional_ | 4 | **`file_sequence_number`** | `long`
| File sequence number
indicating when the file was added. Inherited when null. See [Sequence Number
Inheritance](#sequence-number-inheritance). |
+ | _optional_ | 142 | **`first_row_id`** | `long`
| Base row ID for assigning
`_row_id` values. See [First Row ID Inheritance](#first-row-id-inheritance). |
+ | _optional_ | 6 | **`deleted_positions`** | `binary`
| Positions deleted via
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | _optional_ | 7 | **`replaced_positions`** | `binary`
| Positions replaced via
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ The `deletion_vector` struct has the following fields:
+
+ | On write | Field id | Name | Type | Description |
+ |------------|----------|---------------------|----------|-------------|
+ | _required_ | 155 | **`location`** | `string` | Location of the
file that stores the DV. |
+ | _required_ | 144 | **`offset`** | `long` | Offset in the
file where the content starts. |
+ | _required_ | 145 | **`size_in_bytes`** | `long` | Length of the
referenced content stored in the file. |
+ | _required_ | 156 | **`cardinality`** | `long` | Number of set
bits (deleted rows) in the deletion vector. |
+ | _optional_ | 149 | **`key_metadata`** | `binary` | Key metadata
for encryption; specific to the encryption scheme. |
+
+ ##### Manifest Info
+
+ The `manifest_info` struct has the following fields:
+
+ | On write | Field id | Name | Type |
Description |
+
|------------|----------|----------------------------|----------|-------------|
+ | _optional_ | 522 | **`dv`** | `binary` |
Positions in the referenced leaf manifest that are not live. See [Manifest
Deletion Vectors](#manifest-deletion-vectors). |
+ | _required_ | 504 | **`added_files_count`** | `int` | Count of
entries with status ADDED in the manifest. |
+ | _required_ | 505 | **`existing_files_count`** | `int` | Count of
entries with status EXISTING in the manifest. |
+ | _required_ | 525 | **`modified_files_count`** | `int` | Count of
entries with status MODIFIED in the manifest. |
+ | _required_ | 506 | **`deleted_files_count`** | `int` | Count of
entries with status DELETED in the manifest. |
+ | _required_ | 523 | **`replaced_files_count`** | `int` | Count of
entries with status REPLACED in the manifest. |
+ | _required_ | 512 | **`added_rows_count`** | `long` | Total
number of rows in ADDED entries. |
+ | _required_ | 513 | **`existing_rows_count`** | `long` | Total
number of rows in EXISTING entries. |
+ | _required_ | 526 | **`modified_rows_count`** | `long` | Total
number of rows in MODIFIED entries. |
+ | _required_ | 514 | **`deleted_rows_count`** | `long` | Total
number of rows in DELETED entries. |
+ | _required_ | 524 | **`replaced_rows_count`** | `long` | Total
number of rows in REPLACED entries. |
+ | _required_ | 516 | **`min_sequence_number`** | `long` | Minimum
data sequence number of all live entries in the manifest. |
+
+ The `column_file` struct has the following fields:
+
+ | On write | Field id | Name | Type |
Description |
+
|------------|----------|--------------------------|------------------|-------------|
+ | _required_ | 161 | **`format_version`** | `int` (4: V4) |
Format version of this column file. |
+ | _required_ | 162 | **`field_ids`** | `list<163: int>` |
Live field IDs stored in this column file. |
+ | _required_ | 164 | **`location`** | `string` |
Location of the column file. |
+ | _required_ | 165 | **`file_format`** | `string` |
String file format name: `avro`, `orc`, or `parquet`. |
+ | _required_ | 166 | **`file_size_in_bytes`** | `long` |
Total column file size in bytes. |
+ | _optional_ | 167 | **`key_metadata`** | `binary` |
Implementation-specific key metadata for encryption. |
+
+ **Tracked File Requirements**
+
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest written in v4 may only contain data files.
+ - Row-level deletes may only be written in v4 as deletion vectors in the
data file's `deletion_vector`.
+ - Delete files from pre-v4 tables are valid in upgraded tables and are
tracked in delete manifests written before the upgrade.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - For manifests, `manifest_info.added_files_count`,
`existing_files_count`, `deleted_files_count`, `replaced_files_count`, and
`modified_files_count` must sum to `record_count`.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
v4 leaf manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
+ - Writers should not write a null `tracking.snapshot_id`.
+ - For manifests, `tracking.sequence_number` must equal
`tracking.file_sequence_number`.
+ - For manifests, `spec_id` must be set to the `spec_id` of the manifest's
entries if all entries have the same `spec_id`, and must be null otherwise.
+ - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or
`manifest_info.dv` is set.
+ - `tracking.column_file_snapshot_id` may only be set if `column_files` is
set.
+
+ When a file is added to the dataset, its tracked file must set status to
ADDED and store the snapshot ID in which the file was added.
+
+ When a data file's deletion vector or column files are updated, the writer
must record a MODIFIED entry for the live version and must mark the prior
version as replaced with a REPLACED entry or in a [manifest deletion
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the
writer must set the position in the leaf manifest's
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries'
`dv_snapshot_id` or `column_file_snapshot_id` must record the snapshot in which
their deletion vector, manifest deletion vector, or column files last changed.
+
+ When a file is deleted from the dataset, the deletion must be recorded in
the snapshot that deletes the file with a DELETED entry that stores the
snapshot ID in which the file was deleted or, for an entry in a leaf manifest,
alternatively by setting its position in the leaf manifest's
`tracking.deleted_positions` and `manifest_info.dv` and updating
`tracking.dv_snapshot_id` to the new snapshot ID.
+
+ A leaf manifest whose `manifest_info.dv` changed must have status
MODIFIED. `tracking.deleted_positions` and `tracking.replaced_positions` should
only be set in the snapshot that changes `manifest_info.dv`.
+
+A file that is no longer live may be deleted from the file system when the
snapshot in which it was deleted is garbage collected, assuming that older
snapshots have also been garbage collected [1].
+
+Iceberg v2 adds data and file sequence numbers to the entry and makes the
snapshot ID optional. Values for these fields are inherited from manifest
metadata when `null`. That is, if the field is `null` for an entry, then the
entry must inherit its value from the manifest file's metadata, stored in the
snapshot root.
+The `sequence_number` field represents the data sequence number and must never
change after a file is added to the dataset. The data sequence number
represents a relative age of the file content and should be used for planning
which delete files apply to a data file.
+The `file_sequence_number` field represents the sequence number of the
snapshot that added the file and must also remain unchanged upon assigning at
commit. The file sequence number can't be used for pruning delete files as the
data within the file may have an older data sequence number.
+The data and file sequence numbers are inherited only if the entry status is 1
(added). If the entry status is 0 (existing) or 2 (deleted), the entry must
include both sequence numbers explicitly. In v4, a MODIFIED entry that adds a
column file also inherits its data sequence number.
Notes:
-1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
-2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the
IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
-3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
-4. Position delete metadata can use `referenced_data_file` when all deletes
tracked by the entry are in a single data file. Setting the referenced file is
required for deletion vectors.
-5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
-6. The following field ids are reserved on `data_file`: 141.
+1. Technically, data files can be deleted when the last snapshot that contains
the file as "live" data is garbage collected. But this is harder to detect and
requires finding the diff of multiple snapshots. It is easier to track what
files are deleted in a snapshot and delete them when that snapshot expires. It
is not recommended to add a deleted file back to a table. Adding a deleted file
can lead to edge cases where incremental deletes can break table snapshots.
+2. Manifest lists are required in v2, so that the `sequence_number` and
`snapshot_id` to inherit are always available.
##### Field-level Metrics and Statistics
Field statistics (or, interchangeably, metrics) are used when filtering to
select data and delete files.
-In v3 and earlier, metrics are stored in maps keyed by field id:
`value_counts`, `null_value_counts`, `nan_value_counts`, `lower_bounds` and
`upper_bounds`.
+In v3 and earlier, metrics are stored in maps keyed by field id:
`value_counts`, `null_value_counts`, `nan_value_counts`, `lower_bounds` and
`upper_bounds`. Bounds in these maps use the single-value serialization
detailed in [Appendix D](#appendix-d-single-value-serialization).
In v4, metrics are stored as typed values in the `content_stats` struct,
documented in the [Content Stats](#content-stats) section.
-Both representations store equivalent information. If a map or id in a map is
missing in v3, it is equivalent to a `null` value or missing field struct in
v4. Lower bounds must be less than or equal to all non-null and non-NaN values
and upper bounds must be greater than or equal to all non-null and non-NaN
values.
+Both representations store equivalent information. If a map or id in a map is
missing in v3, it is equivalent to a `null` value or missing field struct in
v4. Lower bounds must be less than or equal to all non-null and non-NaN values
and upper bounds must be greater than or equal to all non-null and non-NaN
values. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
Review Comment:
Done!
##########
format/spec.md:
##########
@@ -928,12 +1053,12 @@ A simple (and recommended) way for writers to adapt
existing metadata for table
Manifests track the sequence number when a data or delete file was added to
the table.
-When adding a new file, its data and file sequence numbers are set to `null`
because the snapshot's sequence number is not assigned until the snapshot is
successfully committed. When reading, sequence numbers are inherited by
replacing `null` with the manifest's sequence number from the manifest list.
+When adding a new file, its data and file sequence numbers are set to `null`
because the snapshot's sequence number is not assigned until the snapshot is
successfully committed. When reading, sequence numbers are inherited by
replacing `null` with the manifest's sequence number from the snapshot root.
It is also possible to add a new file with data that logically belongs to an
older sequence number. In that case, the data sequence number must be provided
explicitly and not inherited. However, the file sequence number must be always
assigned when the snapshot is successfully committed.
When writing an existing file to a new manifest or marking an existing file as
deleted, the data and file sequence numbers must be non-null and set to the
original values that were either inherited or provided at the commit time.
-Inheriting sequence numbers through the metadata tree allows writing a new
manifest without a known sequence number, so that a manifest can be written
once and reused in commit retries. To change a sequence number for a retry,
only the manifest list must be rewritten.
+Inheriting sequence numbers through the metadata tree allows writing a new
manifest without a known sequence number, so that a manifest can be written
once and reused in commit retries. To change a sequence number for a retry,
only the snapshot root must be rewritten.
Review Comment:
Done!
##########
format/spec.md:
##########
@@ -977,18 +1116,18 @@ For other optional snapshot summary fields, see
[Appendix F](#optional-snapshot-
Data and delete files for a snapshot can be stored in more than one manifest.
This enables:
* Appends can add a new manifest to minimize the amount of data written,
instead of adding new records by rewriting and appending to an existing
manifest. (This is called a “fast append”.)
-* Tables can use multiple partition specs. A table’s partition configuration
can evolve if, for example, its data volume changes. Each manifest uses a
single partition spec, and queries do not need to change because partition
filters are derived from data predicates.
+* Tables can use multiple partition specs. A table’s partition configuration
can evolve if, for example, its data volume changes. Queries do not need to
change because partition filters are derived from data predicates. In v1-v3,
each manifest uses a single partition spec.
* Large tables can be split across multiple manifests so that implementations
can parallelize job planning or reduce the cost of rewriting a manifest.
-Manifests for a snapshot are tracked by a manifest list.
+Manifests for a snapshot are tracked by the snapshot root.
Valid snapshots are stored as a list in table metadata. For serialization, see
Appendix C.
#### Snapshot Row IDs
A snapshot's `first-row-id` is assigned to the table's current `next-row-id`
on each commit attempt. If a commit is retried, the `first-row-id` must be
reassigned based on the table's current `next-row-id`. The `first-row-id` field
is required even if a commit does not assign any ID space.
-The snapshot's `first-row-id` is the starting `first_row_id` assigned to
manifests in the snapshot's manifest list.
+The snapshot's `first-row-id` is the starting `first_row_id` assigned to
manifests in the snapshot root. In v4, this includes data files in the root
manifest.
Review Comment:
Done! Updated this everywhere.
##########
format/spec.md:
##########
@@ -965,6 +1093,20 @@ A snapshot consists of the following fields:
| | | _required_ | **`added-rows`** |
The upper bound of the number of rows with assigned row IDs, see [Row
Lineage](#row-lineage) |
| | | _optional_ | **`key-id`** | ID
of the encryption key that encrypts the manifest list key metadata |
+=== "v4"
+ | v4 | Field | Description |
+ | ---------- |------------------------------|-------------|
+ | _required_ | **`snapshot-id`** | A unique long ID |
+ | _optional_ | **`parent-snapshot-id`** | The snapshot ID of the
snapshot's parent. Omitted for any snapshot with no parent |
+ | _required_ | **`sequence-number`** | A monotonically increasing
long that tracks the order of changes to a table |
+ | _required_ | **`timestamp-ms`** | A timestamp when the
snapshot was created, used for garbage collection and table inspection |
+ | _required_ | **`root-manifest`** | The location of the root
manifest for this snapshot |
Review Comment:
Done, and I also generalized the first sentence of that section to say
snapshot root file.
##########
format/spec.md:
##########
@@ -1043,19 +1182,21 @@ Notes:
#### First Row ID Assignment
-The `first_row_id` for existing manifests must be preserved when writing a new
manifest list. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read.
+The `first_row_id` for existing manifests must be preserved when writing a new
snapshot root. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read. In v4, data files in the root
manifest must also have a `first_row_id`: existing values must be preserved,
and data files without one are assigned a `first_row_id` in the same way as
data manifests.
Review Comment:
Done!
##########
format/spec.md:
##########
@@ -1043,19 +1182,21 @@ Notes:
#### First Row ID Assignment
-The `first_row_id` for existing manifests must be preserved when writing a new
manifest list. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read.
+The `first_row_id` for existing manifests must be preserved when writing a new
snapshot root. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read. In v4, data files in the root
manifest must also have a `first_row_id`: existing values must be preserved,
and data files without one are assigned a `first_row_id` in the same way as
data manifests.
-The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest.
+The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest. In v4, when the last assigned entry is a data file, each
`first_row_id` must be greater than or equal to the last assigned
`first_row_id` plus that data file's `record_count`.
Review Comment:
Reworded to start from the first file without a `first_row_id`, then split
the manifest and data file cases into bullets.
##########
format/spec.md:
##########
@@ -1043,19 +1182,21 @@ Notes:
#### First Row ID Assignment
-The `first_row_id` for existing manifests must be preserved when writing a new
manifest list. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read.
+The `first_row_id` for existing manifests must be preserved when writing a new
snapshot root. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read. In v4, data files in the root
manifest must also have a `first_row_id`: existing values must be preserved,
and data files without one are assigned a `first_row_id` in the same way as
data manifests.
-The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest.
+The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest. In v4, when the last assigned entry is a data file, each
`first_row_id` must be greater than or equal to the last assigned
`first_row_id` plus that data file's `record_count`.
A simple and valid approach is to estimate the number of rows in data files
that will be assigned a `first_row_id` using the manifest's `added_rows_count`
and `existing_rows_count`: `first_row_id = last_assigned.first_row_id +
last_assigned.added_rows_count + last_assigned.existing_rows_count`.
### Scan Planning
-Scans are planned by reading the manifest files for the current snapshot.
Deleted entries in data and delete manifests (those marked with status
"DELETED") are not used in a scan.
+Scans are planned by reading the manifests referenced by the snapshot root for
the current snapshot; starting in v4, the snapshot root may also contain data
files.
+
+A scan uses only [live](#manifest-schema) entries.
-Manifests that contain no matching files, determined using either file counts
or partition summaries, may be skipped.
+Manifests that contain no matching files, determined using file counts,
partition summaries (v1-v3), or column stats (v4), may be skipped.
Review Comment:
Done!
##########
format/spec.md:
##########
@@ -1043,19 +1182,21 @@ Notes:
#### First Row ID Assignment
-The `first_row_id` for existing manifests must be preserved when writing a new
manifest list. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read.
+The `first_row_id` for existing manifests must be preserved when writing a new
snapshot root. The value of `first_row_id` for delete manifests is always
`null`. The `first_row_id` is only assigned for data manifests that do not have
a `first_row_id`. Assignment must account for data files that will be assigned
`first_row_id` values when the manifest is read. In v4, data files in the root
manifest must also have a `first_row_id`: existing values must be preserved,
and data files without one are assigned a `first_row_id` in the same way as
data manifests.
-The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest.
+The first manifest without a `first_row_id` is assigned a value that is
greater than or equal to the `first_row_id` of the snapshot. Subsequent
manifests without a `first_row_id` are assigned one based on the previous
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must
increase by the row count of all files that will be assigned a `first_row_id`
via inheritance in the last assigned manifest. That is, each `first_row_id`
must be greater than or equal to the last assigned `first_row_id` plus the
total record count of data files with a null `first_row_id` in the last
assigned manifest. In v4, when the last assigned entry is a data file, each
`first_row_id` must be greater than or equal to the last assigned
`first_row_id` plus that data file's `record_count`.
A simple and valid approach is to estimate the number of rows in data files
that will be assigned a `first_row_id` using the manifest's `added_rows_count`
and `existing_rows_count`: `first_row_id = last_assigned.first_row_id +
last_assigned.added_rows_count + last_assigned.existing_rows_count`.
### Scan Planning
-Scans are planned by reading the manifest files for the current snapshot.
Deleted entries in data and delete manifests (those marked with status
"DELETED") are not used in a scan.
+Scans are planned by reading the manifests referenced by the snapshot root for
the current snapshot; starting in v4, the snapshot root may also contain data
files.
+
+A scan uses only [live](#manifest-schema) entries.
-Manifests that contain no matching files, determined using either file counts
or partition summaries, may be skipped.
+Manifests that contain no matching files, determined using file counts,
partition summaries (v1-v3), or column stats (v4), may be skipped.
-For each manifest, scan predicates, which filter data rows, are converted to
partition predicates, which filter partition tuples. These partition predicates
are used to select relevant data files, delete files, and deletion vector
metadata. Conversion uses the partition spec that was used to write the
manifest file regardless of the current partition spec.
+In v1-v3, for each manifest, scan predicates, which filter data rows, are
converted to partition predicates, which filter partition tuples. These
partition predicates are used to select relevant data files, delete files, and
deletion vector metadata. Conversion uses the partition spec that was used to
write the manifest file regardless of the current partition spec.
Review Comment:
Discussed offline, we wanted to keep wording wherever possible to be general
rather than have too much branching logic on version. I updated this to lead
with content stats and describe partition predicates as what was used before
content stats were introduced in v4.
##########
format/spec.md:
##########
@@ -1071,7 +1212,8 @@ Duplicate live manifest entries for the same content file
violate [content file
Delete files and deletion vector metadata that match the filters must be
applied to data files at read time, limited by the following scope rules.
-* A deletion vector must be applied to a data file when all of the following
are true:
+* In v4, a deletion vector must be applied to the data file tracked by the
same entry. No path, sequence number, or partition comparison applies because
the vector is colocated with the data file.
Review Comment:
Done, the rules now just say whether the DV is colocated or in a delete
manifest instead of using versions.
##########
format/spec.md:
##########
@@ -1071,7 +1212,8 @@ Duplicate live manifest entries for the same content file
violate [content file
Delete files and deletion vector metadata that match the filters must be
applied to data files at read time, limited by the following scope rules.
-* A deletion vector must be applied to a data file when all of the following
are true:
+* In v4, a deletion vector must be applied to the data file tracked by the
same entry. No path, sequence number, or partition comparison applies because
the vector is colocated with the data file.
+* In v1-v3, a deletion vector must be applied to a data file when all of the
following are true:
Review Comment:
Done, covered by the change above.
##########
format/spec.md:
##########
@@ -1367,7 +1510,7 @@ When removing a data file, writers must also remove any
deletion vector that app
Row-level delete files (both equality and position delete files) are valid
Iceberg data files: files must use valid Iceberg formats, schemas, and column
projection. It is recommended that these delete files are written using the
table's default file format.
-Row-level delete files and deletion vectors are tracked by manifests. A
separate set of manifests is used for delete files and DVs, but the same
manifest schema is used for both data and delete manifests. Deletion vectors
are tracked individually by file location, offset, and length within the
containing file. Deletion vector metadata must include the referenced data file.
+Row-level delete files and deletion vectors are tracked by manifests. A
separate set of manifests is used for delete files and DVs, but the same
manifest schema is used for both data and delete manifests. Deletion vectors
are tracked individually by file location, offset, and length within the
containing file. Deletion vector metadata must include the referenced data
file. Starting in v4, a deletion vector is instead colocated with its data file.
Review Comment:
Updated to say a DV may be colocated, writers must colocate new DVs in v4,
and DVs in pre-upgrade delete manifests remain valid. The invariants that
there's at most one DV per data file and that equality deletes still apply are
already established in the Delete Formats and Scan Planning sections.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]