amogh-jahagirdar commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4160922668
##########
format/spec.md:
##########
@@ -731,39 +744,150 @@ The `data_file` struct consists of the following fields:
| _optional_ | _optional_ | _optional_ | **`110 null_value_counts`**
| `map<121: int, 122: long>` |
Map from column id to number of null values in the column |
| _optional_ | _optional_ | _optional_ | **`137 nan_value_counts`**
| `map<138: int, 139: long>` |
Map from column id to number of NaN values in the column |
| _optional_ | _optional_ | | ~~**`111 distinct_counts`**~~
| `map<123: int, 124: long>` |
**Deprecated. Do not write.** |
- | _optional_ | _optional_ | _optional_ | **`125 lower_bounds`**
| `map<126: int, 127: binary>` |
Map from column id to lower bound in the column serialized as binary [1]. Each
value must be less than or equal to all non-null, non-NaN values in the column
for the file [2] |
- | _optional_ | _optional_ | _optional_ | **`128 upper_bounds`**
| `map<129: int, 130: binary>` |
Map from column id to upper bound in the column serialized as binary [1]. Each
value must be greater than or equal to all non-null, non-Nan values in the
column for the file [2] |
+ | _optional_ | _optional_ | _optional_ | **`125 lower_bounds`**
| `map<126: int, 127: binary>` |
Map from column id to lower bound in the column serialized as binary. Each
value must be less than or equal to all non-null, non-NaN values in the column
for the file. See [Field-level Metrics and
Statistics](#field-level-metrics-and-statistics) |
+ | _optional_ | _optional_ | _optional_ | **`128 upper_bounds`**
| `map<129: int, 130: binary>` |
Map from column id to upper bound in the column serialized as binary. Each
value must be greater than or equal to all non-null, non-Nan values in the
column for the file. See [Field-level Metrics and
Statistics](#field-level-metrics-and-statistics) |
| _optional_ | _optional_ | _optional_ | **`131 key_metadata`**
| `binary` |
Implementation-specific key metadata for encryption |
| _optional_ | _optional_ | _optional_ | **`132 split_offsets`**
| `list<133: long>` |
Split offsets for the data file. For example, all row group offsets in a
Parquet file. Must be sorted ascending |
| | _optional_ | _optional_ | **`135 equality_ids`**
| `list<136: int>` |
Field ids used to determine row equality in equality delete files. Required
when `content=2` and should be null otherwise. Fields with ids listed in this
column must be present in the delete file |
- | _optional_ | _optional_ | _optional_ | **`140 sort_order_id`**
| `int` |
ID representing sort order for this file [3]. |
+ | _optional_ | _optional_ | _optional_ | **`140 sort_order_id`**
| `int` |
ID representing sort order for this file [1]. |
| | | _optional_ | **`142 first_row_id`**
| `long` |
The `_row_id` for the first row in the data file. See [First Row ID
Inheritance](#first-row-id-inheritance) |
- | | _optional_ | _optional_ | **`143 referenced_data_file`**
| `string` |
Fully qualified location (URI with FS scheme) of a data file that all deletes
reference [4] |
- | | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
- | | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
+ | | _optional_ | _optional_ | **`143 referenced_data_file`**
| `string` |
Fully qualified location (URI with FS scheme) of a data file that all deletes
reference [2] |
+ | | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [3] |
+ | | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [3] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 2. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 3. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 4. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ The `tracked_file` struct has the following fields:
+
+ | On write | Field id | Name | Type
| Description |
+
|------------|----------|--------------------------|-------------------------------------------------------|-------------|
+ | _required_ | 134 | **`content_type`** | `int` (0: DATA, 3:
DATA_MANIFEST, 4: DELETE_MANIFEST) | Type of content stored in the entry. |
+ | _required_ | 157 | **`format_version`** | `int` (0: PRE-V4, 4:
V4) | Writer format version. |
+ | _required_ | 147 | **`tracking`** | `tracking` struct
| Tracking metadata like status, snapshot ID,
and sequence number. See tracking struct below. |
+ | _required_ | 100 | **`location`** | `string`
| Location of the file. |
+ | _required_ | 101 | **`file_format`** | `string`
| String file format name: `avro`, `orc`, or
`parquet` |
+ | _required_ | 104 | **`file_size_in_bytes`** | `long`
| Total file size in bytes. |
+ | _required_ | 103 | **`record_count`** | `long`
| Number of records in this file. |
+ | _optional_ | 131 | **`key_metadata`** | `binary`
| Implementation-specific key metadata for
encryption. |
+ | _optional_ | 132 | **`split_offsets`** | `list<133: long>`
| Split offsets for the data file. Must be
sorted ascending. |
+ | _optional_ | 141 | **`spec_id`** | `int`
| ID of the partition spec used to partition
the file; null if unpartitioned |
+ | _optional_ | 102 | **`partition`** | `struct<...>`
| Partition data tuple for the file; null if
unpartitioned. |
+ | _optional_ | 140 | **`sort_order_id`** | `int`
| ID representing sort order for this file. If
missing or unknown, the order is assumed to be unsorted. |
+ | _optional_ | 146 | **`content_stats`** | `content_stats`
struct | Field-level stats. See [Content
Stats](#content-stats). |
+ | _optional_ | 150 | **`manifest_info`** | `manifest_info`
struct | Manifest-specific stats. See [Manifest
Info](#manifest-info) |
+ | _optional_ | 148 | **`deletion_vector`** | `deletion_vector`
struct | Row-level deletion vector for a data
file. |
+ | _optional_ | 158 | **`column_files`** | `list<159:
column_file>` | Column files associated with this
file. |
+
+ The `tracking` struct has the following fields:
+
+ | On write | Field id | Name | Type
| Description |
+
|------------|----------|--------------------------------------|---------------------------------------------------------------------|-------------|
+ | _required_ | 0 | **`status`** | `int` (0:
EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | Used to track
additions, deletions, replacements, and modifications. |
+ | _optional_ | 1 | **`snapshot_id`** | `long`
| Snapshot ID where
the file was added or deleted. Inherited when null. |
+ | _optional_ | 5 | **`dv_snapshot_id`** | `long`
| Snapshot ID where
the deletion vector or manifest DV last changed. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | _optional_ | 160 | **`latest_column_file_snapshot_id`** | `long`
| Snapshot ID where
the latest column file was added. |
+ | _optional_ | 3 | **`sequence_number`** | `long`
| Data sequence
number of the file. Inherited when null. See [Sequence Number
Inheritance](#sequence-number-inheritance). |
+ | _optional_ | 4 | **`file_sequence_number`** | `long`
| File sequence
number indicating when the file was added. Inherited when null. See [Sequence
Number Inheritance](#sequence-number-inheritance). |
+ | _optional_ | 142 | **`first_row_id`** | `long`
| Base row ID for
assigning `_row_id` values. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | _optional_ | 6 | **`deleted_positions`** | `binary`
| Positions deleted
via manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | _optional_ | 7 | **`replaced_positions`** | `binary`
| Positions replaced
via manifest DV in the `dv_snapshot_id` snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ The `deletion_vector` struct has the following fields:
+
+ | On write | Field id | Name | Type | Description |
+ |------------|----------|---------------------|----------|-------------|
+ | _required_ | 155 | **`location`** | `string` | Location of the
file that stores the DV. |
+ | _required_ | 144 | **`offset`** | `long` | Offset in the
file where the content starts. |
+ | _required_ | 145 | **`size_in_bytes`** | `long` | Length of the
referenced content stored in the file. |
+ | _required_ | 156 | **`cardinality`** | `long` | Number of set
bits (deleted rows) in the deletion vector. |
+ | _optional_ | 149 | **`key_metadata`** | `binary` | Key metadata
for encryption; specific to the encryption scheme. |
+
+ ##### Manifest Info
+
+ The `manifest_info` struct has the following fields:
+
+ | On write | Field id | Name | Type |
Description |
+
|------------|----------|----------------------------|----------|-------------|
+ | _required_ | 504 | **`added_files_count`** | `int` | Count of
entries with status ADDED in the manifest. |
+ | _required_ | 505 | **`existing_files_count`** | `int` | Count of
entries with status EXISTING in the manifest. |
+ | _required_ | 506 | **`deleted_files_count`** | `int` | Count of
entries with status DELETED in the manifest. |
+ | _required_ | 523 | **`replaced_files_count`** | `int` | Count of
entries with status REPLACED in the manifest. |
+ | _required_ | 525 | **`modified_files_count`** | `int` | Count of
entries with status MODIFIED in the manifest. |
+ | _required_ | 512 | **`added_rows_count`** | `long` | Total
number of rows in ADDED entries. |
+ | _required_ | 513 | **`existing_rows_count`** | `long` | Total
number of rows in EXISTING entries. |
+ | _required_ | 514 | **`deleted_rows_count`** | `long` | Total
number of rows in DELETED entries. |
+ | _required_ | 524 | **`replaced_rows_count`** | `long` | Total
number of rows in REPLACED entries. |
+ | _required_ | 526 | **`modified_rows_count`** | `long` | Total
number of rows in MODIFIED entries. |
+ | _required_ | 516 | **`min_sequence_number`** | `long` | Minimum
data sequence number of all live entries in the manifest. |
+ | _optional_ | 522 | **`dv`** | `binary` |
Positions in the referenced leaf manifest that are not live. See [Manifest
Deletion Vectors](#manifest-deletion-vectors). |
+
+ The `column_file` struct has the following fields:
+
+ | On write | Field id | Name | Type |
Description |
+
|------------|----------|--------------------------|-------------------|-------------|
+ | _required_ | 161 | **`format_version`** | `int` (4: V4) |
Format version of this column file. |
+ | _required_ | 162 | **`field_ids`** | `list<163: int>` |
Live field IDs stored in this column file. |
+ | _required_ | 164 | **`location`** | `string` |
Location of the column file. |
+ | _required_ | 165 | **`file_format`** | `string` |
String file format name: `avro`, `orc`, or `parquet`. |
+ | _required_ | 166 | **`file_size_in_bytes`** | `long` |
Total column file size in bytes. |
+ | _optional_ | 167 | **`key_metadata`** | `binary` |
Implementation-specific key metadata for encryption. |
+ | _optional_ | 168 | **`split_offsets`** | `list<169: long>` |
Split offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest written in v4 may only contain data files.
+ - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3
delete files.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
data manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
+ - Writers should not write a null `tracking.snapshot_id`.
Review Comment:
Should is intentional. We allow for inheritance on read, but want to
discourage writers from producing null. so I added this "should" requirement,
see https://github.com/apache/iceberg/pull/16025#discussion_r4150486693
If we discourage by making it a must requirement I think it'd require some
more discussion because I think there are people out there who do actively use
this.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]