amogh-jahagirdar commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4199718426


##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
+
+    Notes:
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Tracking 
metadata like status, snapshot ID, and sequence number. See tracking struct 
below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to partition the file; null if unpartitioned |
+    | 102 | **`partition`** | `struct<...>` | *optional* | Partition data 
tuple for the file; null if unpartitioned. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | 
Field-level stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column files associated with this file. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null. See [Sequence 
Number Inheritance](#sequence-number-inheritance). |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |
+    | 149 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+
+    **`manifest_info` struct (field 150)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 504 | **`added_files_count`** | `int` | *required* | Count of entries 
with status ADDED in the manifest. |
+    | 505 | **`existing_files_count`** | `int` | *required* | Count of entries 
with status EXISTING in the manifest. |
+    | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries 
with status DELETED in the manifest. |
+    | 523 | **`replaced_files_count`** | `int` | *required* | Count of entries 
with status REPLACED in the manifest. |
+    | 525 | **`modified_files_count`** | `int` | *required* | Count of entries 
with status MODIFIED in the manifest. |
+    | 512 | **`added_rows_count`** | `long` | *required* | Total number of 
rows in ADDED entries. |
+    | 513 | **`existing_rows_count`** | `long` | *required* | Total number of 
rows in EXISTING entries. |
+    | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of 
rows in DELETED entries. |
+    | 524 | **`replaced_rows_count`** | `long` | *required* | Total number of 
rows in REPLACED entries. |
+    | 526 | **`modified_rows_count`** | `long` | *required* | Total number of 
rows in MODIFIED entries. |
+    | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data 
sequence number of all live entries in the manifest. |
+    | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced 
leaf manifest that are not live. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`column_file` struct (element 159 of `column_files`, field 158)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 161 | **`format_version`** | `int` (4: V4) | *required* | Format version 
of this column file. |
+    | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs 
stored in this column file. |
+    | 164 | **`location`** | `string` | *required* | Location of the column 
file. |
+    | 165 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet`. |
+    | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file 
size in bytes. |
+    | 167 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split 
offsets for the column file. Must be sorted ascending. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3 
delete files.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.
+    - `manifest_info` must be set if and only if the tracked file is a 
manifest.
+    - `deletion_vector` may only be set if the tracked file is a data file.
+    - `column_files` may only be set if the tracked file is a data file or a 
data manifest.
+    - `tracking.deleted_positions` and `tracking.replaced_positions` may only 
be set if the tracked file is a manifest.
+    - `tracking.snapshot_id` and `tracking.sequence_number` are required for 
the tracked file in the root manifest.
+    - For manifests, `tracking.sequence_number` must equal 
`tracking.file_sequence_number`.
+    - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or 
`manifest_info.dv` is set.
+    - `tracking.latest_column_file_snapshot_id` may only be set if 
`column_files` is set.
+
+    When a file is added to the dataset, its tracked file must set status to 
ADDED and store the snapshot ID in which the file was added.
+
+    When a data file's deletion vector or column files are updated, the writer 
must record a MODIFIED entry for the live version and must mark the prior 
version as replaced with a REPLACED entry or in a [manifest deletion 
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the 
writer must set the position in the leaf manifest's 
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries' 
`dv_snapshot_id` or `latest_column_file_snapshot_id` must record the snapshot 
in which their deletion vector, manifest deletion vector, or column files last 
changed.

Review Comment:
   Applied this, though I kept the original "must" wording because this is a 
strict requirement. I also merged `dv_snapshot_id` and 
`column_file_snapshot_id` into a single `modified_snapshot_id` so the REPLACED 
snapshot is unambiguous.



##########
format/spec.md:
##########
@@ -1387,6 +1530,20 @@ At most one deletion vector is allowed per data file in 
a snapshot. If a DV is w
 
 [puffin-spec]: https://iceberg.apache.org/puffin-spec/
 
+#### Manifest Deletion Vectors
+
+A manifest deletion vector marks entries in a leaf manifest as deleted or 
replaced by encoding their positions in a bitmap. A set bit at position P 
indicates that the entry at position P in the referenced leaf manifest is 
deleted or replaced.
+
+Manifest deletion vectors are encoded using the [Mumbling bitmap 
spec][mumbling-spec] and stored inline on the root manifest entry that 
references the leaf manifest. The snapshot in which the vector last changed is 
recorded in `tracking.dv_snapshot_id`; the three bitmaps are:
+
+* `manifest_info.dv`: every position in the leaf manifest that is not live.
+* `tracking.deleted_positions`: the positions deleted in the `dv_snapshot_id` 
snapshot.
+* `tracking.replaced_positions`: the positions replaced in the 
`dv_snapshot_id` snapshot.

Review Comment:
   Done, consolidated onto a single `modified_snapshot_id`.



##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
+
+    Notes:
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Tracking 
metadata like status, snapshot ID, and sequence number. See tracking struct 
below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to partition the file; null if unpartitioned |
+    | 102 | **`partition`** | `struct<...>` | *optional* | Partition data 
tuple for the file; null if unpartitioned. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | 
Field-level stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column files associated with this file. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null. See [Sequence 
Number Inheritance](#sequence-number-inheritance). |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |
+    | 149 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+
+    **`manifest_info` struct (field 150)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 504 | **`added_files_count`** | `int` | *required* | Count of entries 
with status ADDED in the manifest. |
+    | 505 | **`existing_files_count`** | `int` | *required* | Count of entries 
with status EXISTING in the manifest. |
+    | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries 
with status DELETED in the manifest. |
+    | 523 | **`replaced_files_count`** | `int` | *required* | Count of entries 
with status REPLACED in the manifest. |
+    | 525 | **`modified_files_count`** | `int` | *required* | Count of entries 
with status MODIFIED in the manifest. |
+    | 512 | **`added_rows_count`** | `long` | *required* | Total number of 
rows in ADDED entries. |
+    | 513 | **`existing_rows_count`** | `long` | *required* | Total number of 
rows in EXISTING entries. |
+    | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of 
rows in DELETED entries. |
+    | 524 | **`replaced_rows_count`** | `long` | *required* | Total number of 
rows in REPLACED entries. |
+    | 526 | **`modified_rows_count`** | `long` | *required* | Total number of 
rows in MODIFIED entries. |
+    | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data 
sequence number of all live entries in the manifest. |
+    | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced 
leaf manifest that are not live. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`column_file` struct (element 159 of `column_files`, field 158)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 161 | **`format_version`** | `int` (4: V4) | *required* | Format version 
of this column file. |
+    | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs 
stored in this column file. |
+    | 164 | **`location`** | `string` | *required* | Location of the column 
file. |
+    | 165 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet`. |
+    | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file 
size in bytes. |
+    | 167 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split 
offsets for the column file. Must be sorted ascending. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3 
delete files.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.
+    - `manifest_info` must be set if and only if the tracked file is a 
manifest.
+    - `deletion_vector` may only be set if the tracked file is a data file.
+    - `column_files` may only be set if the tracked file is a data file or a 
data manifest.
+    - `tracking.deleted_positions` and `tracking.replaced_positions` may only 
be set if the tracked file is a manifest.
+    - `tracking.snapshot_id` and `tracking.sequence_number` are required for 
the tracked file in the root manifest.
+    - For manifests, `tracking.sequence_number` must equal 
`tracking.file_sequence_number`.
+    - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or 
`manifest_info.dv` is set.
+    - `tracking.latest_column_file_snapshot_id` may only be set if 
`column_files` is set.
+
+    When a file is added to the dataset, its tracked file must set status to 
ADDED and store the snapshot ID in which the file was added.
+
+    When a data file's deletion vector or column files are updated, the writer 
must record a MODIFIED entry for the live version and must mark the prior 
version as replaced with a REPLACED entry or in a [manifest deletion 
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the 
writer must set the position in the leaf manifest's 
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries' 
`dv_snapshot_id` or `latest_column_file_snapshot_id` must record the snapshot 
in which their deletion vector, manifest deletion vector, or column files last 
changed.
+
+    When a file is deleted from the dataset, the deletion must be recorded in 
the snapshot that deletes the file with a DELETED entry that stores the 
snapshot ID in which the file was deleted or, for an entry in a leaf manifest, 
alternatively by setting its position in the leaf manifest's 
`tracking.deleted_positions` and `manifest_info.dv` and updating 
`tracking.dv_snapshot_id` to the new snapshot ID.
+
+    A leaf manifest whose `manifest_info.dv` changed must have status 
MODIFIED. `tracking.deleted_positions` and `tracking.replaced_positions` should 
only be set in the snapshot that changes `manifest_info.dv`.
+
+The file may be deleted from the file system when the snapshot in which it was 
deleted is garbage collected, assuming that older snapshots have also been 
garbage collected [1].

Review Comment:
   Done, I applied your wording and moved this and its footnote into a 
"Deleting Files" subsection under Snapshot Retention Policy. I didn't change 
anything else though since it's important to preserve the semantics, and I 
really want to avoid any subtle issues that come from us moving stuff around.



##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file or 
manifest. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, `parquet`, or `puffin` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column update files associated with this entry. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. Deletes are not used in scans. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null and status is 1 (ADDED). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null and status is 
ADDED. |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |
+    | 149 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+
+    **`manifest_info` struct (field 150)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 504 | **`added_files_count`** | `int` | *required* | Count of entries 
with status ADDED in the manifest. |
+    | 505 | **`existing_files_count`** | `int` | *required* | Count of entries 
with status EXISTING in the manifest. |
+    | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries 
with status DELETED in the manifest. |
+    | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries 
with status REPLACED in the manifest. |
+    | 524 | **`modified_files_count`** | `int` | *required* | Count of entries 
with status MODIFIED in the manifest. |
+    | 512 | **`added_rows_count`** | `long` | *required* | Total number of 
rows in ADDED entries. |
+    | 513 | **`existing_rows_count`** | `long` | *required* | Total number of 
rows in EXISTING entries. |
+    | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of 
rows in DELETED entries. |
+    | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of 
rows in REPLACED entries. |
+    | 525 | **`modified_rows_count`** | `long` | *required* | Total number of 
rows in MODIFIED entries. |
+    | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data 
sequence number of all live entries in the manifest. |
+    | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced 
leaf manifest that are not live. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the 
manifest deletion vector. |
+
+    **`column_file` struct (element 159 of `column_files`, field 158)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 161 | **`format_version`** | `int` | *required* | Format version of this 
column file. |
+    | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs 
stored in this column file. |
+    | 164 | **`location`** | `string` | *required* | Location of the column 
file. |
+    | 165 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet`. |
+    | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file 
size in bytes. |
+    | 167 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split 
offsets for the column file. Must be sorted ascending. |
+
+    **Tracked File Requirements**
+
+    - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES).
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest may only contain data files.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.
+    - `manifest_info` must be set if and only if the tracked file is a 
manifest.
+    - `deletion_vector` may only be set if the tracked file is a data file.
+    - `column_files` may only be set if the tracked file is a data file or a 
data manifest.
+    - `tracking.deleted_positions` and `tracking.replaced_positions` may only 
be set if the tracked file is a manifest.
+    - `tracking.snapshot_id` and `tracking.sequence_number` are required for 
the tracked file in the root manifest.
+    - For manifests, `tracking.sequence_number` must equal 
`tracking.file_sequence_number`.
+    - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or 
`manifest_info.dv` is set.
+    - `tracking.latest_column_file_snapshot_id` may only be set if 
`column_files` is set.
+    - `manifest_info.dv_cardinality` must be set if and only if 
`manifest_info.dv` is non-null.
+
+    When a file is added to the dataset, its tracked file must set status to 
ADDED and store the snapshot ID in which the file was added.
+
+    When a data file's deletion vector or column files are updated, the writer 
records a MODIFIED entry for the live version and marks the prior version as 
replaced, either with a REPLACED entry or in a [manifest deletion 
vector](#manifest-deletion-vectors). The resulting entries' `dv_snapshot_id` or 
`latest_column_file_snapshot_id` must record the snapshot in which the deletion 
vector or column files, respectively, last changed. For leaf manifest entries, 
MODIFIED marks a live manifest whose `dv` changed.
+
+    When a file is deleted from the dataset, its tracked file must set status 
to DELETED and store the snapshot ID in which the file was deleted. Writers 
must include DELETED entries in the manifest for the snapshot that deletes the 
file. The next manifest written for those entries must omit the DELETED entries.
+
+The file may be deleted from the file system when the snapshot in which it was 
deleted is garbage collected, assuming that older snapshots have also been 
garbage collected [1].
+
+Iceberg v2 adds data and file sequence numbers to the entry and makes the 
snapshot ID optional. Values for these fields are inherited from manifest 
metadata when `null`. That is, if the field is `null` for an entry, then the 
entry must inherit its value from the manifest file's metadata, stored in the 
snapshot root.
+The `sequence_number` field represents the data sequence number and must never 
change after a file is added to the dataset, except during the addition of a 
column file. The data sequence number represents a relative age of the file 
content and should be used for planning which delete files apply to a data file.
+The `file_sequence_number` field represents the sequence number of the 
snapshot that added the file and must also remain unchanged upon assigning at 
commit. The file sequence number can't be used for pruning delete files as the 
data within the file may have an older data sequence number.
+The data and file sequence numbers are inherited only if the entry status is 1 
(added). If the entry status is 0 (existing) or 2 (deleted), the entry must 
include both sequence numbers explicitly.
 
 Notes:
 
-1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
-2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the 
IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
-3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
-4. Position delete metadata can use `referenced_data_file` when all deletes 
tracked by the entry are in a single data file. Setting the referenced file is 
required for deletion vectors.
-5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
-6. The following field ids are reserved on `data_file`: 141.
+1. Technically, data files can be deleted when the last snapshot that contains 
the file as "live" data is garbage collected. But this is harder to detect and 
requires finding the diff of multiple snapshots. It is easier to track what 
files are deleted in a snapshot and delete them when that snapshot expires.  It 
is not recommended to add a deleted file back to a table. Adding a deleted file 
can lead to edge cases where incremental deletes can break table snapshots.

Review Comment:
   I moved it with that sentence into the snapshot retention section instead of 
removing it, since that's where it's relevant.



##########
format/spec.md:
##########
@@ -676,13 +700,19 @@ A manifest file must store the partition spec and other 
metadata as properties i
     | _optional_ | _required_ | `format-version`    | Table format version 
number of the manifest as a string                                              
                                       |
     |            | _required_ | `content`           | Type of content files 
tracked by the manifest: "data" or "deletes"                                    
                                      |
 
+=== "v4"
+    | Requirement | Key                 | Value                                
                                                                                
                       |
+    
|-------------|---------------------|---------------------------------------------------------------------------------------------------------------------------------------------|
+    | _optional_  | `schema-id`         | ID of the schema used to write the 
manifest as a string                                                            
                         |
+    | _optional_  | `format-version`    | Table format version number of the 
manifest as a string                                                            
                         |
+
 #### Content file uniqueness
 
 Within a snapshot, each content file must be referenced by at most one live 
manifest entry across all manifests; otherwise, the snapshot has undefined 
behavior. Writers should not produce multiple manifest entries for the same 
content file in a snapshot (for example, both ADDED and DELETED entries for the 
same file). Writers are not required to validate uniqueness at commit time.
 
-#### Manifest Entry Fields
+#### Entries in Manifests
 
-The schema of a manifest file is defined by the `manifest_entry` struct, which 
consists of the following fields:
+In v1-v3, manifest entries are described by the `manifest_entry` struct. In 
v4, entries are called tracked files and are described by the `tracked_file` 
struct. In v4, `data_file` struct fields are flattened directly into the 
tracked file, and tracking fields are grouped into a nested `tracking` struct. 
An entry is **live** in a snapshot if its `status` is ADDED, EXISTING, or 
MODIFIED and its position is not set in the containing manifest's 
[`manifest_info.dv`](#manifest-deletion-vectors).

Review Comment:
   Done!



##########
format/spec.md:
##########
@@ -731,39 +744,151 @@ The `data_file` struct consists of the following fields:
     | _optional_ | _optional_ | _optional_ | **`110  null_value_counts`**      
| `map<121: int, 122: long>`                                                  | 
Map from column id to number of null values in the column |
     | _optional_ | _optional_ | _optional_ | **`137  nan_value_counts`**       
| `map<138: int, 139: long>`                                                  | 
Map from column id to number of NaN values in the column |
     | _optional_ | _optional_ |            | ~~**`111  distinct_counts`**~~    
| `map<123: int, 124: long>`                                                  | 
**Deprecated. Do not write.** |
-    | _optional_ | _optional_ | _optional_ | **`125  lower_bounds`**           
| `map<126: int, 127: binary>`                                                | 
Map from column id to lower bound in the column serialized as binary [1]. Each 
value must be less than or equal to all non-null, non-NaN values in the column 
for the file [2] |
-    | _optional_ | _optional_ | _optional_ | **`128  upper_bounds`**           
| `map<129: int, 130: binary>`                                                | 
Map from column id to upper bound in the column serialized as binary [1]. Each 
value must be greater than or equal to all non-null, non-Nan values in the 
column for the file [2] |
+    | _optional_ | _optional_ | _optional_ | **`125  lower_bounds`**           
| `map<126: int, 127: binary>`                                                | 
Map from column id to lower bound in the column serialized as binary. Each 
value must be less than or equal to all non-null, non-NaN values in the column 
for the file. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |
+    | _optional_ | _optional_ | _optional_ | **`128  upper_bounds`**           
| `map<129: int, 130: binary>`                                                | 
Map from column id to upper bound in the column serialized as binary. Each 
value must be greater than or equal to all non-null, non-Nan values in the 
column for the file. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |
     | _optional_ | _optional_ | _optional_ | **`131  key_metadata`**           
| `binary`                                                                    | 
Implementation-specific key metadata for encryption |
     | _optional_ | _optional_ | _optional_ | **`132  split_offsets`**          
| `list<133: long>`                                                           | 
Split offsets for the data file. For example, all row group offsets in a 
Parquet file. Must be sorted ascending |
     |            | _optional_ | _optional_ | **`135  equality_ids`**           
| `list<136: int>`                                                            | 
Field ids used to determine row equality in equality delete files. Required 
when `content=2` and should be null otherwise. Fields with ids listed in this 
column must be present in the delete file |
-    | _optional_ | _optional_ | _optional_ | **`140  sort_order_id`**          
| `int`                                                                       | 
ID representing sort order for this file [3]. |
+    | _optional_ | _optional_ | _optional_ | **`140  sort_order_id`**          
| `int`                                                                       | 
ID representing sort order for this file [1]. |
     |            |            | _optional_ | **`142  first_row_id`**           
| `long`                                                                      | 
The `_row_id` for the first row in the data file. See [First Row ID 
Inheritance](#first-row-id-inheritance) |
-    |            | _optional_ | _optional_ | **`143  referenced_data_file`**   
| `string`                                                                    | 
Fully qualified location (URI with FS scheme) of a data file that all deletes 
reference [4] |
-    |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
-    |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
+    |            | _optional_ | _optional_ | **`143  referenced_data_file`**   
| `string`                                                                    | 
Fully qualified location (URI with FS scheme) of a data file that all deletes 
reference [2] |
+    |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [3] |
+    |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [3] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    2. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    3. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    4. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    The `tracked_file` struct has the following fields:
+
+    | On write   | Field id | Name                     | Type                  
                                | Description |
+    
|------------|----------|--------------------------|-------------------------------------------------------|-------------|
+    | _required_ | 134      | **`content_type`**       | `int` (0: DATA, 3: 
DATA_MANIFEST, 4: DELETE_MANIFEST) | Type of content stored in the entry. |
+    | _required_ | 157      | **`format_version`**     | `int` (0: PRE-V4, 4: 
V4)                              | Writer format version. |
+    | _required_ | 147      | **`tracking`**           | `tracking` struct     
                                | Tracking metadata like status, snapshot ID, 
and sequence number. See tracking struct below. |
+    | _required_ | 100      | **`location`**           | `string`              
                                | Location of the file. |
+    | _required_ | 101      | **`file_format`**        | `string`              
                                | String file format name: `avro`, `orc`, or 
`parquet` |
+    | _required_ | 104      | **`file_size_in_bytes`** | `long`                
                                | Total file size in bytes. |
+    | _required_ | 103      | **`record_count`**       | `long`                
                                | Number of records in this file. |
+    | _optional_ | 131      | **`key_metadata`**       | `binary`              
                                | Implementation-specific key metadata for 
encryption. |
+    | _optional_ | 132      | **`split_offsets`**      | `list<133: long>`     
                                | Split offsets for the data file. Must be 
sorted ascending. |
+    | _optional_ | 141      | **`spec_id`**            | `int`                 
                                | ID of the partition spec used to partition 
the file; null if unpartitioned |
+    | _optional_ | 102      | **`partition`**          | `struct<...>`         
                                | Partition data tuple for the file; null if 
unpartitioned. |
+    | _optional_ | 140      | **`sort_order_id`**      | `int`                 
                                | ID representing sort order for this file. If 
missing or unknown, the order is assumed to be unsorted. |
+    | _optional_ | 146      | **`content_stats`**      | `content_stats` 
struct                                | Field-level stats. See [Content 
Stats](#content-stats). |
+    | _optional_ | 150      | **`manifest_info`**      | `manifest_info` 
struct                                | Manifest-specific stats. See [Manifest 
Info](#manifest-info) |
+    | _optional_ | 148      | **`deletion_vector`**    | `deletion_vector` 
struct                              | Row-level deletion vector for a data 
file. |
+    | _optional_ | 158      | **`column_files`**       | `list<159: 
column_file>`                              | Column files associated with this 
file. |
+
+    The `tracking` struct has the following fields:
+
+    | On write   | Field id | Name                          | Type             
                                                   | Description |
+    
|------------|----------|-------------------------------|---------------------------------------------------------------------|-------------|
+    | _required_ | 0        | **`status`**                  | `int` (0: 
EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | Used to track 
additions, deletions, replacements, and modifications. |
+    | _optional_ | 1        | **`snapshot_id`**             | `long`           
                                                   | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | _optional_ | 5        | **`dv_snapshot_id`**          | `long`           
                                                   | Snapshot ID where the 
deletion vector or manifest DV last changed. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | _optional_ | 8        | **`column_file_snapshot_id`** | `long`           
                                                   | Snapshot ID where the 
latest column file was added. |
+    | _optional_ | 3        | **`sequence_number`**         | `long`           
                                                   | Data sequence number of 
the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | _optional_ | 4        | **`file_sequence_number`**    | `long`           
                                                   | File sequence number 
indicating when the file was added. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | _optional_ | 142      | **`first_row_id`**            | `long`           
                                                   | Base row ID for assigning 
`_row_id` values. See [First Row ID Inheritance](#first-row-id-inheritance). |
+    | _optional_ | 6        | **`deleted_positions`**       | `binary`         
                                                   | Positions deleted via 
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | _optional_ | 7        | **`replaced_positions`**      | `binary`         
                                                   | Positions replaced via 
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    The `deletion_vector` struct has the following fields:
+
+    | On write   | Field id | Name                | Type     | Description |
+    |------------|----------|---------------------|----------|-------------|
+    | _required_ | 155      | **`location`**      | `string` | Location of the 
file that stores the DV. |
+    | _required_ | 144      | **`offset`**        | `long`   | Offset in the 
file where the content starts. |
+    | _required_ | 145      | **`size_in_bytes`** | `long`   | Length of the 
referenced content stored in the file. |
+    | _required_ | 156      | **`cardinality`**   | `long`   | Number of set 
bits (deleted rows) in the deletion vector. |
+    | _optional_ | 149      | **`key_metadata`**  | `binary` | Key metadata 
for encryption; specific to the encryption scheme. |
+
+    ##### Manifest Info
+
+    The `manifest_info` struct has the following fields:
+
+    | On write   | Field id | Name                       | Type     | 
Description |
+    
|------------|----------|----------------------------|----------|-------------|
+    | _optional_ | 522      | **`dv`**                   | `binary` | 
Positions in the referenced leaf manifest that are not live. See [Manifest 
Deletion Vectors](#manifest-deletion-vectors). |
+    | _required_ | 504      | **`added_files_count`**    | `int`    | Count of 
entries with status ADDED in the manifest. |
+    | _required_ | 505      | **`existing_files_count`** | `int`    | Count of 
entries with status EXISTING in the manifest. |
+    | _required_ | 525      | **`modified_files_count`** | `int`    | Count of 
entries with status MODIFIED in the manifest. |
+    | _required_ | 506      | **`deleted_files_count`**  | `int`    | Count of 
entries with status DELETED in the manifest. |
+    | _required_ | 523      | **`replaced_files_count`** | `int`    | Count of 
entries with status REPLACED in the manifest. |
+    | _required_ | 512      | **`added_rows_count`**     | `long`   | Total 
number of rows in ADDED entries. |
+    | _required_ | 513      | **`existing_rows_count`**  | `long`   | Total 
number of rows in EXISTING entries. |
+    | _required_ | 526      | **`modified_rows_count`**  | `long`   | Total 
number of rows in MODIFIED entries. |
+    | _required_ | 514      | **`deleted_rows_count`**   | `long`   | Total 
number of rows in DELETED entries. |
+    | _required_ | 524      | **`replaced_rows_count`**  | `long`   | Total 
number of rows in REPLACED entries. |
+    | _required_ | 516      | **`min_sequence_number`**  | `long`   | Minimum 
data sequence number of all live entries in the manifest. |
+
+    The `column_file` struct has the following fields:
+
+    | On write   | Field id | Name                     | Type             | 
Description |
+    
|------------|----------|--------------------------|------------------|-------------|
+    | _required_ | 161      | **`format_version`**     | `int` (4: V4)    | 
Format version of this column file. |
+    | _required_ | 162      | **`field_ids`**          | `list<163: int>` | 
Live field IDs stored in this column file. |
+    | _required_ | 164      | **`location`**           | `string`         | 
Location of the column file. |
+    | _required_ | 165      | **`file_format`**        | `string`         | 
String file format name: `avro`, `orc`, or `parquet`. |
+    | _required_ | 166      | **`file_size_in_bytes`** | `long`           | 
Total column file size in bytes. |
+    | _optional_ | 167      | **`key_metadata`**       | `binary`         | 
Implementation-specific key metadata for encryption. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - Row-level deletes may only be written in v4 as deletion vectors in the 
data file's `deletion_vector`.
+    - Delete files from pre-v4 tables are valid in upgraded tables and are 
tracked in delete manifests written before the upgrade.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.
+    - `manifest_info` must be set if and only if the tracked file is a 
manifest.
+    - For manifests, `manifest_info.added_files_count`, 
`existing_files_count`, `deleted_files_count`, `replaced_files_count`, and 
`modified_files_count` must sum to `record_count`.
+    - `deletion_vector` may only be set if the tracked file is a data file.
+    - `column_files` may only be set if the tracked file is a data file or a 
v4 leaf manifest.
+    - `tracking.deleted_positions` and `tracking.replaced_positions` may only 
be set if the tracked file is a manifest.
+    - `tracking.snapshot_id` and `tracking.sequence_number` are required for 
the tracked file in the root manifest.
+    - Writers should not write a null `tracking.snapshot_id`.
+    - For manifests, `tracking.sequence_number` must equal 
`tracking.file_sequence_number`.
+    - For manifests, `spec_id` must be set to the `spec_id` of the manifest's 
entries if all entries have the same `spec_id`, and must be null otherwise.
+    - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or 
`manifest_info.dv` is set.
+    - `tracking.column_file_snapshot_id` may only be set if `column_files` is 
set.
+
+    When a file is added to the dataset, its tracked file must set status to 
ADDED and store the snapshot ID in which the file was added.
+
+    When a data file's deletion vector or column files are updated, the writer 
must record a MODIFIED entry for the live version and must mark the prior 
version as replaced with a REPLACED entry or in a [manifest deletion 
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the 
writer must set the position in the leaf manifest's 
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries' 
`dv_snapshot_id` or `column_file_snapshot_id` must record the snapshot in which 
their deletion vector, manifest deletion vector, or column files last changed.
+
+    When a file is deleted from the dataset, the deletion must be recorded in 
the snapshot that deletes the file with a DELETED entry that stores the 
snapshot ID in which the file was deleted or, for an entry in a leaf manifest, 
alternatively by setting its position in the leaf manifest's 
`tracking.deleted_positions` and `manifest_info.dv` and updating 
`tracking.dv_snapshot_id` to the new snapshot ID.
+
+    A leaf manifest whose `manifest_info.dv` changed must have status 
MODIFIED. `tracking.deleted_positions` and `tracking.replaced_positions` should 
only be set in the snapshot that changes `manifest_info.dv`.
+
+A file that is no longer live may be deleted from the file system when the 
snapshot in which it was deleted is garbage collected, assuming that older 
snapshots have also been garbage collected [1].
+
+Iceberg v2 adds data and file sequence numbers to the entry and makes the 
snapshot ID optional. Values for these fields are inherited from manifest 
metadata when `null`. That is, if the field is `null` for an entry, then the 
entry must inherit its value from the manifest file's metadata, stored in the 
snapshot root.
+The `sequence_number` field represents the data sequence number and must never 
change after a file is added to the dataset. The data sequence number 
represents a relative age of the file content and should be used for planning 
which delete files apply to a data file.
+The `file_sequence_number` field represents the sequence number of the 
snapshot that added the file and must also remain unchanged upon assigning at 
commit. The file sequence number can't be used for pruning delete files as the 
data within the file may have an older data sequence number.
+The data and file sequence numbers are inherited only if the entry status is 1 
(added). If the entry status is 0 (existing) or 2 (deleted), the entry must 
include both sequence numbers explicitly. In v4, a MODIFIED entry that adds a 
column file also inherits its data sequence number.

Review Comment:
   Done:
   - Combined into an Inheritance section with the sequence number and first 
row ID sections
   - Dropped "v2 adds" and noted v1 at the end
   - MODIFIED entries that add a column file only inherit the data sequence 
number
   - Inheritance only applies to files in leaf manifests, the root requires 
these to be set



##########
format/spec.md:
##########
@@ -731,39 +744,151 @@ The `data_file` struct consists of the following fields:
     | _optional_ | _optional_ | _optional_ | **`110  null_value_counts`**      
| `map<121: int, 122: long>`                                                  | 
Map from column id to number of null values in the column |
     | _optional_ | _optional_ | _optional_ | **`137  nan_value_counts`**       
| `map<138: int, 139: long>`                                                  | 
Map from column id to number of NaN values in the column |
     | _optional_ | _optional_ |            | ~~**`111  distinct_counts`**~~    
| `map<123: int, 124: long>`                                                  | 
**Deprecated. Do not write.** |
-    | _optional_ | _optional_ | _optional_ | **`125  lower_bounds`**           
| `map<126: int, 127: binary>`                                                | 
Map from column id to lower bound in the column serialized as binary [1]. Each 
value must be less than or equal to all non-null, non-NaN values in the column 
for the file [2] |
-    | _optional_ | _optional_ | _optional_ | **`128  upper_bounds`**           
| `map<129: int, 130: binary>`                                                | 
Map from column id to upper bound in the column serialized as binary [1]. Each 
value must be greater than or equal to all non-null, non-Nan values in the 
column for the file [2] |
+    | _optional_ | _optional_ | _optional_ | **`125  lower_bounds`**           
| `map<126: int, 127: binary>`                                                | 
Map from column id to lower bound in the column serialized as binary. Each 
value must be less than or equal to all non-null, non-NaN values in the column 
for the file. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |
+    | _optional_ | _optional_ | _optional_ | **`128  upper_bounds`**           
| `map<129: int, 130: binary>`                                                | 
Map from column id to upper bound in the column serialized as binary. Each 
value must be greater than or equal to all non-null, non-Nan values in the 
column for the file. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |
     | _optional_ | _optional_ | _optional_ | **`131  key_metadata`**           
| `binary`                                                                    | 
Implementation-specific key metadata for encryption |
     | _optional_ | _optional_ | _optional_ | **`132  split_offsets`**          
| `list<133: long>`                                                           | 
Split offsets for the data file. For example, all row group offsets in a 
Parquet file. Must be sorted ascending |
     |            | _optional_ | _optional_ | **`135  equality_ids`**           
| `list<136: int>`                                                            | 
Field ids used to determine row equality in equality delete files. Required 
when `content=2` and should be null otherwise. Fields with ids listed in this 
column must be present in the delete file |
-    | _optional_ | _optional_ | _optional_ | **`140  sort_order_id`**          
| `int`                                                                       | 
ID representing sort order for this file [3]. |
+    | _optional_ | _optional_ | _optional_ | **`140  sort_order_id`**          
| `int`                                                                       | 
ID representing sort order for this file [1]. |
     |            |            | _optional_ | **`142  first_row_id`**           
| `long`                                                                      | 
The `_row_id` for the first row in the data file. See [First Row ID 
Inheritance](#first-row-id-inheritance) |
-    |            | _optional_ | _optional_ | **`143  referenced_data_file`**   
| `string`                                                                    | 
Fully qualified location (URI with FS scheme) of a data file that all deletes 
reference [4] |
-    |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
-    |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
+    |            | _optional_ | _optional_ | **`143  referenced_data_file`**   
| `string`                                                                    | 
Fully qualified location (URI with FS scheme) of a data file that all deletes 
reference [2] |
+    |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [3] |
+    |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [3] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    2. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    3. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    4. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    The `tracked_file` struct has the following fields:
+
+    | On write   | Field id | Name                     | Type                  
                                | Description |
+    
|------------|----------|--------------------------|-------------------------------------------------------|-------------|
+    | _required_ | 134      | **`content_type`**       | `int` (0: DATA, 3: 
DATA_MANIFEST, 4: DELETE_MANIFEST) | Type of content stored in the entry. |
+    | _required_ | 157      | **`format_version`**     | `int` (0: PRE-V4, 4: 
V4)                              | Writer format version. |
+    | _required_ | 147      | **`tracking`**           | `tracking` struct     
                                | Tracking metadata like status, snapshot ID, 
and sequence number. See tracking struct below. |
+    | _required_ | 100      | **`location`**           | `string`              
                                | Location of the file. |
+    | _required_ | 101      | **`file_format`**        | `string`              
                                | String file format name: `avro`, `orc`, or 
`parquet` |
+    | _required_ | 104      | **`file_size_in_bytes`** | `long`                
                                | Total file size in bytes. |
+    | _required_ | 103      | **`record_count`**       | `long`                
                                | Number of records in this file. |
+    | _optional_ | 131      | **`key_metadata`**       | `binary`              
                                | Implementation-specific key metadata for 
encryption. |
+    | _optional_ | 132      | **`split_offsets`**      | `list<133: long>`     
                                | Split offsets for the data file. Must be 
sorted ascending. |
+    | _optional_ | 141      | **`spec_id`**            | `int`                 
                                | ID of the partition spec used to partition 
the file; null if unpartitioned |
+    | _optional_ | 102      | **`partition`**          | `struct<...>`         
                                | Partition data tuple for the file; null if 
unpartitioned. |
+    | _optional_ | 140      | **`sort_order_id`**      | `int`                 
                                | ID representing sort order for this file. If 
missing or unknown, the order is assumed to be unsorted. |
+    | _optional_ | 146      | **`content_stats`**      | `content_stats` 
struct                                | Field-level stats. See [Content 
Stats](#content-stats). |
+    | _optional_ | 150      | **`manifest_info`**      | `manifest_info` 
struct                                | Manifest-specific stats. See [Manifest 
Info](#manifest-info) |
+    | _optional_ | 148      | **`deletion_vector`**    | `deletion_vector` 
struct                              | Row-level deletion vector for a data 
file. |
+    | _optional_ | 158      | **`column_files`**       | `list<159: 
column_file>`                              | Column files associated with this 
file. |
+
+    The `tracking` struct has the following fields:
+
+    | On write   | Field id | Name                          | Type             
                                                   | Description |
+    
|------------|----------|-------------------------------|---------------------------------------------------------------------|-------------|
+    | _required_ | 0        | **`status`**                  | `int` (0: 
EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | Used to track 
additions, deletions, replacements, and modifications. |
+    | _optional_ | 1        | **`snapshot_id`**             | `long`           
                                                   | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | _optional_ | 5        | **`dv_snapshot_id`**          | `long`           
                                                   | Snapshot ID where the 
deletion vector or manifest DV last changed. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | _optional_ | 8        | **`column_file_snapshot_id`** | `long`           
                                                   | Snapshot ID where the 
latest column file was added. |
+    | _optional_ | 3        | **`sequence_number`**         | `long`           
                                                   | Data sequence number of 
the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | _optional_ | 4        | **`file_sequence_number`**    | `long`           
                                                   | File sequence number 
indicating when the file was added. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | _optional_ | 142      | **`first_row_id`**            | `long`           
                                                   | Base row ID for assigning 
`_row_id` values. See [First Row ID Inheritance](#first-row-id-inheritance). |
+    | _optional_ | 6        | **`deleted_positions`**       | `binary`         
                                                   | Positions deleted via 
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | _optional_ | 7        | **`replaced_positions`**      | `binary`         
                                                   | Positions replaced via 
manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    The `deletion_vector` struct has the following fields:
+
+    | On write   | Field id | Name                | Type     | Description |
+    |------------|----------|---------------------|----------|-------------|
+    | _required_ | 155      | **`location`**      | `string` | Location of the 
file that stores the DV. |
+    | _required_ | 144      | **`offset`**        | `long`   | Offset in the 
file where the content starts. |
+    | _required_ | 145      | **`size_in_bytes`** | `long`   | Length of the 
referenced content stored in the file. |
+    | _required_ | 156      | **`cardinality`**   | `long`   | Number of set 
bits (deleted rows) in the deletion vector. |
+    | _optional_ | 149      | **`key_metadata`**  | `binary` | Key metadata 
for encryption; specific to the encryption scheme. |
+
+    ##### Manifest Info
+
+    The `manifest_info` struct has the following fields:
+
+    | On write   | Field id | Name                       | Type     | 
Description |
+    
|------------|----------|----------------------------|----------|-------------|
+    | _optional_ | 522      | **`dv`**                   | `binary` | 
Positions in the referenced leaf manifest that are not live. See [Manifest 
Deletion Vectors](#manifest-deletion-vectors). |
+    | _required_ | 504      | **`added_files_count`**    | `int`    | Count of 
entries with status ADDED in the manifest. |
+    | _required_ | 505      | **`existing_files_count`** | `int`    | Count of 
entries with status EXISTING in the manifest. |
+    | _required_ | 525      | **`modified_files_count`** | `int`    | Count of 
entries with status MODIFIED in the manifest. |
+    | _required_ | 506      | **`deleted_files_count`**  | `int`    | Count of 
entries with status DELETED in the manifest. |
+    | _required_ | 523      | **`replaced_files_count`** | `int`    | Count of 
entries with status REPLACED in the manifest. |
+    | _required_ | 512      | **`added_rows_count`**     | `long`   | Total 
number of rows in ADDED entries. |
+    | _required_ | 513      | **`existing_rows_count`**  | `long`   | Total 
number of rows in EXISTING entries. |
+    | _required_ | 526      | **`modified_rows_count`**  | `long`   | Total 
number of rows in MODIFIED entries. |
+    | _required_ | 514      | **`deleted_rows_count`**   | `long`   | Total 
number of rows in DELETED entries. |
+    | _required_ | 524      | **`replaced_rows_count`**  | `long`   | Total 
number of rows in REPLACED entries. |
+    | _required_ | 516      | **`min_sequence_number`**  | `long`   | Minimum 
data sequence number of all live entries in the manifest. |
+
+    The `column_file` struct has the following fields:
+
+    | On write   | Field id | Name                     | Type             | 
Description |
+    
|------------|----------|--------------------------|------------------|-------------|
+    | _required_ | 161      | **`format_version`**     | `int` (4: V4)    | 
Format version of this column file. |
+    | _required_ | 162      | **`field_ids`**          | `list<163: int>` | 
Live field IDs stored in this column file. |
+    | _required_ | 164      | **`location`**           | `string`         | 
Location of the column file. |
+    | _required_ | 165      | **`file_format`**        | `string`         | 
String file format name: `avro`, `orc`, or `parquet`. |
+    | _required_ | 166      | **`file_size_in_bytes`** | `long`           | 
Total column file size in bytes. |
+    | _optional_ | 167      | **`key_metadata`**       | `binary`         | 
Implementation-specific key metadata for encryption. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - Row-level deletes may only be written in v4 as deletion vectors in the 
data file's `deletion_vector`.
+    - Delete files from pre-v4 tables are valid in upgraded tables and are 
tracked in delete manifests written before the upgrade.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.
+    - `manifest_info` must be set if and only if the tracked file is a 
manifest.
+    - For manifests, `manifest_info.added_files_count`, 
`existing_files_count`, `deleted_files_count`, `replaced_files_count`, and 
`modified_files_count` must sum to `record_count`.
+    - `deletion_vector` may only be set if the tracked file is a data file.
+    - `column_files` may only be set if the tracked file is a data file or a 
v4 leaf manifest.
+    - `tracking.deleted_positions` and `tracking.replaced_positions` may only 
be set if the tracked file is a manifest.
+    - `tracking.snapshot_id` and `tracking.sequence_number` are required for 
the tracked file in the root manifest.
+    - Writers should not write a null `tracking.snapshot_id`.
+    - For manifests, `tracking.sequence_number` must equal 
`tracking.file_sequence_number`.
+    - For manifests, `spec_id` must be set to the `spec_id` of the manifest's 
entries if all entries have the same `spec_id`, and must be null otherwise.
+    - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or 
`manifest_info.dv` is set.
+    - `tracking.column_file_snapshot_id` may only be set if `column_files` is 
set.
+
+    When a file is added to the dataset, its tracked file must set status to 
ADDED and store the snapshot ID in which the file was added.
+
+    When a data file's deletion vector or column files are updated, the writer 
must record a MODIFIED entry for the live version and must mark the prior 
version as replaced with a REPLACED entry or in a [manifest deletion 
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the 
writer must set the position in the leaf manifest's 
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries' 
`dv_snapshot_id` or `column_file_snapshot_id` must record the snapshot in which 
their deletion vector, manifest deletion vector, or column files last changed.
+
+    When a file is deleted from the dataset, the deletion must be recorded in 
the snapshot that deletes the file with a DELETED entry that stores the 
snapshot ID in which the file was deleted or, for an entry in a leaf manifest, 
alternatively by setting its position in the leaf manifest's 
`tracking.deleted_positions` and `manifest_info.dv` and updating 
`tracking.dv_snapshot_id` to the new snapshot ID.
+
+    A leaf manifest whose `manifest_info.dv` changed must have status 
MODIFIED. `tracking.deleted_positions` and `tracking.replaced_positions` should 
only be set in the snapshot that changes `manifest_info.dv`.
+
+A file that is no longer live may be deleted from the file system when the 
snapshot in which it was deleted is garbage collected, assuming that older 
snapshots have also been garbage collected [1].
+
+Iceberg v2 adds data and file sequence numbers to the entry and makes the 
snapshot ID optional. Values for these fields are inherited from manifest 
metadata when `null`. That is, if the field is `null` for an entry, then the 
entry must inherit its value from the manifest file's metadata, stored in the 
snapshot root.
+The `sequence_number` field represents the data sequence number and must never 
change after a file is added to the dataset. The data sequence number 
represents a relative age of the file content and should be used for planning 
which delete files apply to a data file.
+The `file_sequence_number` field represents the sequence number of the 
snapshot that added the file and must also remain unchanged upon assigning at 
commit. The file sequence number can't be used for pruning delete files as the 
data within the file may have an older data sequence number.
+The data and file sequence numbers are inherited only if the entry status is 1 
(added). If the entry status is 0 (existing) or 2 (deleted), the entry must 
include both sequence numbers explicitly. In v4, a MODIFIED entry that adds a 
column file also inherits its data sequence number.
 
 Notes:
 
-1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
-2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the 
IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
-3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
-4. Position delete metadata can use `referenced_data_file` when all deletes 
tracked by the entry are in a single data file. Setting the referenced file is 
required for deletion vectors.
-5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
-6. The following field ids are reserved on `data_file`: 141.
+1. Technically, data files can be deleted when the last snapshot that contains 
the file as "live" data is garbage collected. But this is harder to detect and 
requires finding the diff of multiple snapshots. It is easier to track what 
files are deleted in a snapshot and delete them when that snapshot expires.  It 
is not recommended to add a deleted file back to a table. Adding a deleted file 
can lead to edge cases where incremental deletes can break table snapshots.
+2. Manifest lists are required in v2, so that the `sequence_number` and 
`snapshot_id` to inherit are always available.

Review Comment:
   Done!



##########
format/spec.md:
##########
@@ -820,8 +945,8 @@ Each stats struct holds statistics for one table field. It 
may contain the follo
 
 | Requirement | Offset | Name                      | Type                      
| Included for                                  | Description |
 
|-------------|--------|---------------------------|---------------------------|-----------------------------------------------|-------------|
-| _optional_  | 1      | `lower_bound`             | Field type or `geo_lower` 
| all primitives or `variant`                   | Lower bound stored as the 
field's type, or `geo_lower` for geo types |
-| _optional_  | 2      | `upper_bound`             | Field type or `geo_upper` 
| all primitives or `variant`                   | Upper bound stored as the 
field's type, or `geo_upper` for geo types |
+| _optional_  | 1      | `lower_bound`             | Field type or `geo_lower` 
| all primitives or `variant`                   | Lower bound stored as the 
field's type, or `geo_lower` for geo types. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |
+| _optional_  | 2      | `upper_bound`             | Field type or `geo_upper` 
| all primitives or `variant`                   | Upper bound stored as the 
field's type, or `geo_upper` for geo types. See [Field-level Metrics and 
Statistics](#field-level-metrics-and-statistics) |

Review Comment:
   It was pointing to the float/double bound rules after I moved them out of 
the table notes, but that's not needed anymore so I removed it.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to