mkroll-db commented on code in PR #17822:
URL: https://github.com/apache/iceberg/pull/17822#discussion_r4129561188
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
Review Comment:
I think `timestamp-ms` should be mentioned here as well, since the
constraint is **modified**.
##########
format/spec.md:
##########
@@ -962,6 +1084,20 @@ A snapshot consists of the following fields:
| | | _required_ | **`first-row-id`** |
The first `_row_id` assigned to the first row in the first data file in the
first manifest, see [Row Lineage](#row-lineage) |
| | | _required_ | **`added-rows`** |
The upper bound of the number of rows with assigned row IDs, see [Row
Lineage](#row-lineage) |
| | | _optional_ | **`key-id`** | ID
of the encryption key that encrypts the manifest list key metadata |
+=== "v4"
+ | v4 | Field | Description |
+ |------------|------------------------------|-------------|
+ | _required_ | **`snapshot-id`** | A unique long ID |
Review Comment:
NIT: The table is not properly rendered.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
Review Comment:
Looking through the [expression
spec](https://github.com/huaxingao/iceberg/blob/195121960df1b8611a8108e866ce93c68ff9eb6d/format/expressions-spec.md)
I'm unable to find any limitation on determinism.
This results in the issue that constraints may behave different for
different engines (or at different times) as [pointed
out](https://github.com/apache/iceberg/pull/17822/changes#r4125414156) by
@mbutrovich .
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is
treated as 0 when it is not present. Writers must assign a new constraint an ID
that is higher than the table's current `last-constraint-id` and must update
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be
reused after the constraint that used an ID is removed, because retained
snapshots may still reference the removed ID. Readers must not assume that
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the
[Iceberg expressions spec](expressions-spec.md) and must use ID references so
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An
expression may reference more than one field of the row, such as `start_date <=
end_date`. Expressions that depend on more than one row, such as aggregates and
window functions, and expressions that depend on another table, such as
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or
false and never produces null, so a comparison with a null operand produces
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless
the predicate produces false and a null value therefore satisfies the
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an
optional `price` field is stored as the expression for `price >= 0 OR price IS
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with
constraints as follows. The referenced fields of a `check` constraint are the
field IDs in its `expression`; the referenced fields of a `unique` or
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop
is allowed and those constraints are removed automatically. If a dropped field
is referenced by a multi-column constraint, the writer must reject the drop
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule
above: only `name` and `enforced` may be changed in place; changing an
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must
verify the rows that they add is a property of a constraint, tracked by
`enforced`. Whether a table is known to satisfy a constraint is a property of a
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status | Description |
+|---------------|-------------|
+| `validated` | The constraint was checked and holds for all rows in the
snapshot |
+| `valid` | The constraint holds for all rows in the snapshot because it
was enforced for the commit that produced the snapshot and the parent
snapshot's status is `validated` or `valid` |
Review Comment:
What does `valid` mean for an empty table?
I think in this case we should force either `validated`0or `unvalidated`.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
Review Comment:
My understanding was that this is one of the cases where the writer could
fallback to `valid` instead of `validated`.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is
treated as 0 when it is not present. Writers must assign a new constraint an ID
that is higher than the table's current `last-constraint-id` and must update
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be
reused after the constraint that used an ID is removed, because retained
snapshots may still reference the removed ID. Readers must not assume that
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the
[Iceberg expressions spec](expressions-spec.md) and must use ID references so
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An
expression may reference more than one field of the row, such as `start_date <=
end_date`. Expressions that depend on more than one row, such as aggregates and
window functions, and expressions that depend on another table, such as
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or
false and never produces null, so a comparison with a null operand produces
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless
the predicate produces false and a null value therefore satisfies the
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an
optional `price` field is stored as the expression for `price >= 0 OR price IS
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with
constraints as follows. The referenced fields of a `check` constraint are the
field IDs in its `expression`; the referenced fields of a `unique` or
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop
is allowed and those constraints are removed automatically. If a dropped field
is referenced by a multi-column constraint, the writer must reject the drop
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule
above: only `name` and `enforced` may be changed in place; changing an
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must
verify the rows that they add is a property of a constraint, tracked by
`enforced`. Whether a table is known to satisfy a constraint is a property of a
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status | Description |
+|---------------|-------------|
+| `validated` | The constraint was checked and holds for all rows in the
snapshot |
+| `valid` | The constraint holds for all rows in the snapshot because it
was enforced for the commit that produced the snapshot and the parent
snapshot's status is `validated` or `valid` |
+| `invalid` | The constraint was checked and at least one row in the
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the
constraints that have that status for the snapshot:
+
+| Requirement | Field name | Type | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`** | `list<int>` | IDs of constraints that are
`validated` for the snapshot |
+| _optional_ | **`valid`** | `list<int>` | IDs of constraints that are
`valid` for the snapshot |
+| _optional_ | **`invalid`** | `list<int>` | IDs of constraints that are
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that
status for the snapshot. Every constraint that exists when the snapshot is
created must be listed in exactly one of the four lists, and a `constraint-id`
must not appear in more than one list. A list with no constraints may be
omitted. A constraint whose ID is not present in any list did not exist when
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded
independently, so the size of `constraint-statuses` grows with the number of
constraints in a table. This keeps the encoding simple; more compact
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`,
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that
has constraints, and must place every constraint that exists when the snapshot
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the
commit and the parent snapshot's status for the constraint is `validated` or
`valid`
Review Comment:
Following constraint: `COUNT(*) > 100` could becomes `false` in case of
delete. But for `compaction` I agree. Engines should be allowed to 'carry-over'
the statuses if the operation performed does not modify data.
Assuming deterministic constraints.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is
treated as 0 when it is not present. Writers must assign a new constraint an ID
that is higher than the table's current `last-constraint-id` and must update
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be
reused after the constraint that used an ID is removed, because retained
snapshots may still reference the removed ID. Readers must not assume that
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the
[Iceberg expressions spec](expressions-spec.md) and must use ID references so
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An
expression may reference more than one field of the row, such as `start_date <=
end_date`. Expressions that depend on more than one row, such as aggregates and
window functions, and expressions that depend on another table, such as
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or
false and never produces null, so a comparison with a null operand produces
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless
the predicate produces false and a null value therefore satisfies the
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an
optional `price` field is stored as the expression for `price >= 0 OR price IS
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with
constraints as follows. The referenced fields of a `check` constraint are the
field IDs in its `expression`; the referenced fields of a `unique` or
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop
is allowed and those constraints are removed automatically. If a dropped field
is referenced by a multi-column constraint, the writer must reject the drop
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule
above: only `name` and `enforced` may be changed in place; changing an
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must
verify the rows that they add is a property of a constraint, tracked by
`enforced`. Whether a table is known to satisfy a constraint is a property of a
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status | Description |
+|---------------|-------------|
+| `validated` | The constraint was checked and holds for all rows in the
snapshot |
+| `valid` | The constraint holds for all rows in the snapshot because it
was enforced for the commit that produced the snapshot and the parent
snapshot's status is `validated` or `valid` |
+| `invalid` | The constraint was checked and at least one row in the
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the
constraints that have that status for the snapshot:
+
+| Requirement | Field name | Type | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`** | `list<int>` | IDs of constraints that are
`validated` for the snapshot |
+| _optional_ | **`valid`** | `list<int>` | IDs of constraints that are
`valid` for the snapshot |
+| _optional_ | **`invalid`** | `list<int>` | IDs of constraints that are
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that
status for the snapshot. Every constraint that exists when the snapshot is
created must be listed in exactly one of the four lists, and a `constraint-id`
must not appear in more than one list. A list with no constraints may be
omitted. A constraint whose ID is not present in any list did not exist when
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded
independently, so the size of `constraint-statuses` grows with the number of
constraints in a table. This keeps the encoding simple; more compact
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`,
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that
has constraints, and must place every constraint that exists when the snapshot
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the
commit and the parent snapshot's status for the constraint is `validated` or
`valid`
+* A constraint must not be listed as `invalid` unless a row in the snapshot is
known to violate it
+* `unvalidated` is the status of a constraint that cannot be listed in any
other status
+
+Enforcing a constraint for a commit is not sufficient to list it as `valid`.
When the parent snapshot's status is not `validated` or `valid`, rows added by
earlier commits were never checked, so the status is `unvalidated` even though
the writer verified the rows that it added.
+
+When a constraint becomes enforced, either by being added with `enforced` set
to true or by `enforced` changing from false to true, writers should validate
the table and record `validated`. A writer that does not validate records
`unvalidated`, and the constraint remains `unvalidated` until a later
validation records `validated`.
+
+A writer does not have to check every row in a single scan. After checking
every row in an ancestor snapshot, a writer may check only the rows added
between that ancestor and the current snapshot and record `validated` for the
current snapshot. This allows a validation to finish on a table that is written
concurrently, without blocking writes or restarting the scan.
+
+A snapshot's `constraint-statuses` must not be modified after the snapshot is
created. Recording a different status for a constraint requires a new snapshot.
A snapshot that changes only constraint statuses may reuse its parent's
manifest list.
+
+A constraint that a snapshot reports as `valid` may later be found not to hold
for that snapshot. A `valid` status depends on every commit in the snapshot's
history having correctly enforced the constraint, and Iceberg records the
status a writer reports without re-checking the data. So if any of those
writers was buggy or non-compliant, the `valid` status can be wrong even though
nothing detected it at commit time. A `validated` status, which reflects an
actual scan of all rows, does not depend on that chain. Writers should then
commit a snapshot that records `invalid` and should expire the snapshots that
state that the constraint holds, because queries against those snapshots would
otherwise continue to rely on a constraint that does not hold.
Review Comment:
Wouldn't that put a lot of burden on the readers?
For a given old snapshot they would need to read all the new snapshots to
verify that there was no `invalid` marker if they want to make sure that the
data is correctly constraint.
In addition we expect readers to implement optimizations on top of
constraints. Not expiring invalid snapshots could if it hasn't result in wrong
data being returned by the readers.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
Review Comment:
Regarding equality I'm in favor of equality by value not by
bit-representation.
Is in line with the other DBs:
- Postgres
- MySQL
- Sql Server
- Postgres
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is
treated as 0 when it is not present. Writers must assign a new constraint an ID
that is higher than the table's current `last-constraint-id` and must update
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be
reused after the constraint that used an ID is removed, because retained
snapshots may still reference the removed ID. Readers must not assume that
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the
[Iceberg expressions spec](expressions-spec.md) and must use ID references so
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An
expression may reference more than one field of the row, such as `start_date <=
end_date`. Expressions that depend on more than one row, such as aggregates and
window functions, and expressions that depend on another table, such as
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or
false and never produces null, so a comparison with a null operand produces
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless
the predicate produces false and a null value therefore satisfies the
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an
optional `price` field is stored as the expression for `price >= 0 OR price IS
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with
constraints as follows. The referenced fields of a `check` constraint are the
field IDs in its `expression`; the referenced fields of a `unique` or
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop
is allowed and those constraints are removed automatically. If a dropped field
is referenced by a multi-column constraint, the writer must reject the drop
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule
above: only `name` and `enforced` may be changed in place; changing an
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must
verify the rows that they add is a property of a constraint, tracked by
`enforced`. Whether a table is known to satisfy a constraint is a property of a
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status | Description |
+|---------------|-------------|
+| `validated` | The constraint was checked and holds for all rows in the
snapshot |
+| `valid` | The constraint holds for all rows in the snapshot because it
was enforced for the commit that produced the snapshot and the parent
snapshot's status is `validated` or `valid` |
+| `invalid` | The constraint was checked and at least one row in the
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the
constraints that have that status for the snapshot:
+
+| Requirement | Field name | Type | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`** | `list<int>` | IDs of constraints that are
`validated` for the snapshot |
+| _optional_ | **`valid`** | `list<int>` | IDs of constraints that are
`valid` for the snapshot |
+| _optional_ | **`invalid`** | `list<int>` | IDs of constraints that are
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that
status for the snapshot. Every constraint that exists when the snapshot is
created must be listed in exactly one of the four lists, and a `constraint-id`
must not appear in more than one list. A list with no constraints may be
omitted. A constraint whose ID is not present in any list did not exist when
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded
independently, so the size of `constraint-statuses` grows with the number of
constraints in a table. This keeps the encoding simple; more compact
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`,
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that
has constraints, and must place every constraint that exists when the snapshot
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the
commit and the parent snapshot's status for the constraint is `validated` or
`valid`
+* A constraint must not be listed as `invalid` unless a row in the snapshot is
known to violate it
+* `unvalidated` is the status of a constraint that cannot be listed in any
other status
+
+Enforcing a constraint for a commit is not sufficient to list it as `valid`.
When the parent snapshot's status is not `validated` or `valid`, rows added by
earlier commits were never checked, so the status is `unvalidated` even though
the writer verified the rows that it added.
+
+When a constraint becomes enforced, either by being added with `enforced` set
to true or by `enforced` changing from false to true, writers should validate
the table and record `validated`. A writer that does not validate records
`unvalidated`, and the constraint remains `unvalidated` until a later
validation records `validated`.
+
+A writer does not have to check every row in a single scan. After checking
every row in an ancestor snapshot, a writer may check only the rows added
between that ancestor and the current snapshot and record `validated` for the
current snapshot. This allows a validation to finish on a table that is written
concurrently, without blocking writes or restarting the scan.
Review Comment:
We had a long discussion about `uuid` and `validated` vs `valid`.
If I remember correctly we landed on `uuid`s being `validated` if they the
parent is `validated` and the current commit `valid`.
We should mention this explicitly here.
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
Review Comment:
+1 on listing allowed constraint types.s
+1 on `primary-key` fields must have `required`
##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the
following behavior: `-NaN` < `
A data or delete file is associated with a sort order by the sort order's id
within [a manifest](#manifests). Therefore, the table must declare all the sort
orders for lookup. A table could also be configured with a default sort order
id, indicating how the new data should be sorted by default. Writers should use
this default sort order to sort the data on write, but are not required to if
the default order is prohibitively expensive, as it would be for streaming
writes.
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to
satisfy. A constraint's definition is stored in table metadata. Whether a
constraint holds is recorded for each snapshot, see [Constraint
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are
performed by engines that write to a table. Iceberg stores constraint
definitions and records the status that a writer reports for a commit without
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple
fields and evolve independently. Every constraint references the fields that it
applies to by field ID, so a constraint continues to apply to the same columns
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name | Type | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`** | `int` | ID of the constraint;
unique within the table |
+| _required_ | **`type`** | `string` | The constraint type:
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`** | `string` | A name for the
constraint that is unique within the table. Names are for human consumption and
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`** | `boolean` | Whether writers must
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`** | `long` | Timestamp in
milliseconds from the unix epoch when the constraint was created or last
modified. The timestamp is informational and must not be used to determine
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`** | `expression` | The predicate that
every row must satisfy, see [Check Constraint
Expressions](#check-constraint-expressions). Required for a `check` constraint
and must not be set for other types |
+| _optional_ | **`field-ids`** | `list<int>` | The list of field
IDs that the constraint applies to. Required for a `unique` or `primary-key`
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the
constraint based on its `type`. Each type carries only the metadata that it
requires: a `check` constraint has an `expression` and must not declare
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and
must not declare an `expression`. This keeps a single source of truth for the
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference
primitive fields that are either top-level fields or nested in required
structs, and must not reference fields within a `list` or a `map`. These are
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add
satisfy the constraint and must fail the write if they do not. A writer that
cannot verify an enforced constraint must reject writes to the table. When a
constraint is not enforced, writers are not required to verify the rows they
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans
several fields is expressed as a single `primary-key` constraint over multiple
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field
IDs](#identifier-field-ids), which express the same concept: a set of fields
that identifies a row, without a uniqueness guarantee. When a table is upgraded
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place.
Changing the `expression` of a `check` constraint or the `field-ids` of a
`unique` or `primary-key` constraint changes what the constraint requires, so
it must be done by removing the constraint and adding a new one with a new
`constraint-id`, so that statuses recorded for the old definition are not read
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is
treated as 0 when it is not present. Writers must assign a new constraint an ID
that is higher than the table's current `last-constraint-id` and must update
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be
reused after the constraint that used an ID is removed, because retained
snapshots may still reference the removed ID. Readers must not assume that
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the
[Iceberg expressions spec](expressions-spec.md) and must use ID references so
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An
expression may reference more than one field of the row, such as `start_date <=
end_date`. Expressions that depend on more than one row, such as aggregates and
window functions, and expressions that depend on another table, such as
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or
false and never produces null, so a comparison with a null operand produces
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless
the predicate produces false and a null value therefore satisfies the
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an
optional `price` field is stored as the expression for `price >= 0 OR price IS
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with
constraints as follows. The referenced fields of a `check` constraint are the
field IDs in its `expression`; the referenced fields of a `unique` or
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop
is allowed and those constraints are removed automatically. If a dropped field
is referenced by a multi-column constraint, the writer must reject the drop
unless that constraint is removed in the same change.
Review Comment:
I understand why we are doing it but it feels a asymmetric.
Dropping a single column will drop single-column constraints only of that
column 'silently'.
Dropping a single column with a multi-column constraint will fail 'loud'.
I think we should apply the same rule to single-column constraints as well,
that the column can only be dropped if the constraints removal is in the same
change. This also feels more deliberate.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]