rdblue commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4224475156


##########
format/spec.md:
##########
@@ -1043,35 +1189,41 @@ Notes:
 
 #### First Row ID Assignment
 
-The `first_row_id` for existing manifests must be preserved when writing a new 
manifest list. The value of `first_row_id` for delete manifests is always 
`null`. The `first_row_id` is only assigned for data manifests that do not have 
a `first_row_id`. Assignment must account for data files that will be assigned 
`first_row_id` values when the manifest is read.
+The `first_row_id` for existing manifests must be preserved when writing a new 
snapshot root file. The value of `first_row_id` for delete manifests is always 
`null`. The `first_row_id` is only assigned for data manifests that do not have 
a `first_row_id`. Assignment must account for data files that will be assigned 
`first_row_id` values when the manifest is read. Data files in the root 
manifest (v4) must also have a `first_row_id`: existing values must be 
preserved and `first_row_id` is assigned when a data file is ADDED.
+
+The first file in the snapshot root file without a `first_row_id` is assigned 
a value that is greater than or equal to the `first_row_id` of the snapshot. 
Subsequent files without a `first_row_id` are assigned one based on the 
previous file to be assigned a `first_row_id`. Each assigned `first_row_id` 
must be greater than or equal to the last assigned `first_row_id` plus the row 
count of the last assigned file, where the row count is:
 
-The first manifest without a `first_row_id` is assigned a value that is 
greater than or equal to the `first_row_id` of the snapshot. Subsequent 
manifests without a `first_row_id` are assigned one based on the previous 
manifest to be assigned a `first_row_id`. Each assigned `first_row_id` must 
increase by the row count of all files that will be assigned a `first_row_id` 
via inheritance in the last assigned manifest. That is, each `first_row_id` 
must be greater than or equal to the last assigned `first_row_id` plus the 
total record count of data files with a null `first_row_id` in the last 
assigned manifest.
+* For a manifest, the total record count of data files with a null 
`first_row_id` in the manifest.
+* For a data file, its `record_count`.
 
 A simple and valid approach is to estimate the number of rows in data files 
that will be assigned a `first_row_id` using the manifest's `added_rows_count` 
and `existing_rows_count`: `first_row_id = last_assigned.first_row_id + 
last_assigned.added_rows_count + last_assigned.existing_rows_count`.
 
 ### Scan Planning
 
-Scans are planned by reading the manifest files for the current snapshot. 
Deleted entries in data and delete manifests (those marked with status 
"DELETED") are not used in a scan.
+A reader plans a scan by producing live data files from the snapshot root file 
and any leaf manifests referenced by the root.
 
-Manifests that contain no matching files, determined using either file counts 
or partition summaries, may be skipped.
+A scan uses only [live](#manifest-schema) entries.
 
-For each manifest, scan predicates, which filter data rows, are converted to 
partition predicates, which filter partition tuples. These partition predicates 
are used to select relevant data files, delete files, and deletion vector 
metadata. Conversion uses the partition spec that was used to write the 
manifest file regardless of the current partition spec.
+Manifests that contain no matching files, determined using file counts, 
partition summaries, or column stats, may be skipped.
 
-Scan predicates are converted to partition predicates using an _inclusive 
projection_: if a scan predicate matches a row, then the partition predicate 
must match that row’s partition. This is called _inclusive_ [1] because rows 
that do not match the scan predicate may be included in the scan by the 
partition predicate.
+Using content stats, a manifest is filtered by evaluating scan predicates 
against the column bounds and counts of each tracked file. The same filter 
logic can be used for both data and delete files because both store metrics of 
the rows either inserted or deleted. If metrics show that a delete file has no 
rows that match a scan predicate, it may be ignored just as a data file would 
be ignored [1].
+
+Before content stats were introduced in v4, each manifest was filtered using 
scan predicates converted to partition predicates, which filter partition 
tuples. These partition predicates are used to select relevant data files, 
delete files, and deletion vector metadata. Conversion uses the partition spec 
that was used to write the manifest file regardless of the current partition 
spec. Column bounds and counts stored by field id in metrics maps are used in 
the same way as content stats.

Review Comment:
   ```suggestion
   v1-v3 manifest are filtered using scan predicates converted to partition 
predicates, which filter partition tuples. These partition predicates are used 
to select relevant data files, delete files, and deletion vector metadata. 
Conversion uses the partition spec that was used to write the manifest file 
regardless of the current partition spec. Column bounds and counts stored by 
field id in metrics maps are used in the same way as content stats.
   ```



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to