mkroll-db commented on code in PR #17822:
URL: https://github.com/apache/iceberg/pull/17822#discussion_r4129561188


##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.

Review Comment:
   I think `timestamp-ms` should be mentioned here as well, since the 
constraint is **modified**.



##########
format/spec.md:
##########
@@ -962,6 +1084,20 @@ A snapshot consists of the following fields:
     |            |            | _required_ | **`first-row-id`**           | 
The first `_row_id` assigned to the first row in the first data file in the 
first manifest, see [Row Lineage](#row-lineage) |
     |            |            | _required_ | **`added-rows`**             | 
The upper bound of the number of rows with assigned row IDs, see [Row 
Lineage](#row-lineage) |
     |            |            | _optional_ | **`key-id`**                 | ID 
of the encryption key that encrypts the manifest list key metadata |
+=== "v4"
+    | v4         | Field                        | Description |
+    |------------|------------------------------|-------------|
+    | _required_ | **`snapshot-id`**            | A unique long ID |

Review Comment:
   NIT: The table is not properly rendered.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |

Review Comment:
   Looking through the [expression 
spec](https://github.com/huaxingao/iceberg/blob/195121960df1b8611a8108e866ce93c68ff9eb6d/format/expressions-spec.md)
 I'm unable to find any limitation on determinism.
   
   This results in the issue that constraints may behave different for 
different engines (or at different times) as [pointed 
out](https://github.com/apache/iceberg/pull/17822/changes#r4125414156) by 
@mbutrovich .



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions 
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule 
above: only `name` and `enforced` may be changed in place; changing an 
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must 
verify the rows that they add is a property of a constraint, tracked by 
`enforced`. Whether a table is known to satisfy a constraint is a property of a 
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status        | Description |
+|---------------|-------------|
+| `validated`   | The constraint was checked and holds for all rows in the 
snapshot |
+| `valid`       | The constraint holds for all rows in the snapshot because it 
was enforced for the commit that produced the snapshot and the parent 
snapshot's status is `validated` or `valid` |

Review Comment:
   What does `valid` mean for an empty table?
   I think in this case we should force either `validated`0or `unvalidated`.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.

Review Comment:
   My understanding was that this is one of the cases where the writer could 
fallback to `valid` instead of `validated`.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions 
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule 
above: only `name` and `enforced` may be changed in place; changing an 
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must 
verify the rows that they add is a property of a constraint, tracked by 
`enforced`. Whether a table is known to satisfy a constraint is a property of a 
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status        | Description |
+|---------------|-------------|
+| `validated`   | The constraint was checked and holds for all rows in the 
snapshot |
+| `valid`       | The constraint holds for all rows in the snapshot because it 
was enforced for the commit that produced the snapshot and the parent 
snapshot's status is `validated` or `valid` |
+| `invalid`     | The constraint was checked and at least one row in the 
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is 
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the 
constraints that have that status for the snapshot:
+
+| Requirement | Field name        | Type        | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`**   | `list<int>` | IDs of constraints that are 
`validated` for the snapshot |
+| _optional_ | **`valid`**       | `list<int>` | IDs of constraints that are 
`valid` for the snapshot |
+| _optional_ | **`invalid`**     | `list<int>` | IDs of constraints that are 
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are 
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that 
status for the snapshot. Every constraint that exists when the snapshot is 
created must be listed in exactly one of the four lists, and a `constraint-id` 
must not appear in more than one list. A list with no constraints may be 
omitted. A constraint whose ID is not present in any list did not exist when 
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded 
independently, so the size of `constraint-statuses` grows with the number of 
constraints in a table. This keeps the encoding simple; more compact 
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim 
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`, 
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and 
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that 
has constraints, and must place every constraint that exists when the snapshot 
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for 
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the 
commit and the parent snapshot's status for the constraint is `validated` or 
`valid`

Review Comment:
   Following constraint: `COUNT(*) > 100` could becomes `false` in case of 
delete. But for `compaction` I agree. Engines should be allowed to 'carry-over' 
the statuses if the operation performed does not modify data.
   Assuming deterministic constraints.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions 
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule 
above: only `name` and `enforced` may be changed in place; changing an 
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must 
verify the rows that they add is a property of a constraint, tracked by 
`enforced`. Whether a table is known to satisfy a constraint is a property of a 
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status        | Description |
+|---------------|-------------|
+| `validated`   | The constraint was checked and holds for all rows in the 
snapshot |
+| `valid`       | The constraint holds for all rows in the snapshot because it 
was enforced for the commit that produced the snapshot and the parent 
snapshot's status is `validated` or `valid` |
+| `invalid`     | The constraint was checked and at least one row in the 
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is 
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the 
constraints that have that status for the snapshot:
+
+| Requirement | Field name        | Type        | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`**   | `list<int>` | IDs of constraints that are 
`validated` for the snapshot |
+| _optional_ | **`valid`**       | `list<int>` | IDs of constraints that are 
`valid` for the snapshot |
+| _optional_ | **`invalid`**     | `list<int>` | IDs of constraints that are 
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are 
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that 
status for the snapshot. Every constraint that exists when the snapshot is 
created must be listed in exactly one of the four lists, and a `constraint-id` 
must not appear in more than one list. A list with no constraints may be 
omitted. A constraint whose ID is not present in any list did not exist when 
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded 
independently, so the size of `constraint-statuses` grows with the number of 
constraints in a table. This keeps the encoding simple; more compact 
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim 
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`, 
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and 
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that 
has constraints, and must place every constraint that exists when the snapshot 
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for 
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the 
commit and the parent snapshot's status for the constraint is `validated` or 
`valid`
+* A constraint must not be listed as `invalid` unless a row in the snapshot is 
known to violate it
+* `unvalidated` is the status of a constraint that cannot be listed in any 
other status
+
+Enforcing a constraint for a commit is not sufficient to list it as `valid`. 
When the parent snapshot's status is not `validated` or `valid`, rows added by 
earlier commits were never checked, so the status is `unvalidated` even though 
the writer verified the rows that it added.
+
+When a constraint becomes enforced, either by being added with `enforced` set 
to true or by `enforced` changing from false to true, writers should validate 
the table and record `validated`. A writer that does not validate records 
`unvalidated`, and the constraint remains `unvalidated` until a later 
validation records `validated`.
+
+A writer does not have to check every row in a single scan. After checking 
every row in an ancestor snapshot, a writer may check only the rows added 
between that ancestor and the current snapshot and record `validated` for the 
current snapshot. This allows a validation to finish on a table that is written 
concurrently, without blocking writes or restarting the scan.
+
+A snapshot's `constraint-statuses` must not be modified after the snapshot is 
created. Recording a different status for a constraint requires a new snapshot. 
A snapshot that changes only constraint statuses may reuse its parent's 
manifest list.
+
+A constraint that a snapshot reports as `valid` may later be found not to hold 
for that snapshot. A `valid` status depends on every commit in the snapshot's 
history having correctly enforced the constraint, and Iceberg records the 
status a writer reports without re-checking the data. So if any of those 
writers was buggy or non-compliant, the `valid` status can be wrong even though 
nothing detected it at commit time. A `validated` status, which reflects an 
actual scan of all rows, does not depend on that chain. Writers should then 
commit a snapshot that records `invalid` and should expire the snapshots that 
state that the constraint holds, because queries against those snapshots would 
otherwise continue to rely on a constraint that does not hold.

Review Comment:
   Wouldn't that put a lot of burden on the readers?
   For a given old snapshot they would need to read all the new snapshots to 
verify that there was no `invalid` marker if they want to make sure that the 
data is correctly constraint.
   
   In addition we expect readers to implement optimizations on top of 
constraints. Not expiring invalid snapshots could if it hasn't result in wrong 
data being returned by the readers.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value

Review Comment:
   Regarding equality I'm in favor of equality by value not by 
bit-representation. 
   Is in line with the other DBs:
   - Postgres
   - MySQL
   - Sql Server
   - Postgres



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.
+* The type of a referenced field must not be changed, even for type promotions 
that are otherwise allowed. This restriction may be relaxed in a later version.
+
+Changing a constraint's own definition is governed by the mutability rule 
above: only `name` and `enforced` may be changed in place; changing an 
`expression` or `field-ids` requires a new constraint.
+
+#### Constraint Validation
+
+Enforcement and validation are separate properties. Whether writers must 
verify the rows that they add is a property of a constraint, tracked by 
`enforced`. Whether a table is known to satisfy a constraint is a property of a 
table's data, tracked per snapshot by `constraint-statuses`.
+
+The status of a constraint for a snapshot is one of:
+
+| Status        | Description |
+|---------------|-------------|
+| `validated`   | The constraint was checked and holds for all rows in the 
snapshot |
+| `valid`       | The constraint holds for all rows in the snapshot because it 
was enforced for the commit that produced the snapshot and the parent 
snapshot's status is `validated` or `valid` |
+| `invalid`     | The constraint was checked and at least one row in the 
snapshot violates it |
+| `unvalidated` | Whether the constraint holds for all rows in the snapshot is 
not known |
+
+A snapshot's `constraint-statuses` records, for each status, the IDs of the 
constraints that have that status for the snapshot:
+
+| Requirement | Field name        | Type        | Description |
+|-------------|-------------------|-------------|-------------|
+| _optional_ | **`validated`**   | `list<int>` | IDs of constraints that are 
`validated` for the snapshot |
+| _optional_ | **`valid`**       | `list<int>` | IDs of constraints that are 
`valid` for the snapshot |
+| _optional_ | **`invalid`**     | `list<int>` | IDs of constraints that are 
`invalid` for the snapshot |
+| _optional_ | **`unvalidated`** | `list<int>` | IDs of constraints that are 
`unvalidated` for the snapshot |
+
+Each list contains the `constraint-id` of every constraint that has that 
status for the snapshot. Every constraint that exists when the snapshot is 
created must be listed in exactly one of the four lists, and a `constraint-id` 
must not appear in more than one list. A list with no constraints may be 
omitted. A constraint whose ID is not present in any list did not exist when 
the snapshot was created, so the snapshot makes no claim about it.
+
+This is an explicit representation: each constraint's status is recorded 
independently, so the size of `constraint-statuses` grows with the number of 
constraints in a table. This keeps the encoding simple; more compact 
representations may be added in a later version if it becomes a problem.
+
+Readers must determine the status of a constraint for a snapshot as follows:
+
+1. If the snapshot has no `constraint-statuses`, the snapshot makes no claim 
about any constraint
+2. If the constraint's `constraint-id` is listed in `validated`, `valid`, 
`invalid`, or `unvalidated`, that is its status
+3. Otherwise, the constraint did not exist when the snapshot was created and 
the snapshot makes no claim about it
+
+Writers must record `constraint-statuses` in every snapshot of a table that 
has constraints, and must place every constraint that exists when the snapshot 
is created into exactly one status list, following these rules:
+
+* A constraint must not be listed as `validated` unless it was checked for 
every row in the snapshot
+* A constraint must not be listed as `valid` unless it was enforced for the 
commit and the parent snapshot's status for the constraint is `validated` or 
`valid`
+* A constraint must not be listed as `invalid` unless a row in the snapshot is 
known to violate it
+* `unvalidated` is the status of a constraint that cannot be listed in any 
other status
+
+Enforcing a constraint for a commit is not sufficient to list it as `valid`. 
When the parent snapshot's status is not `validated` or `valid`, rows added by 
earlier commits were never checked, so the status is `unvalidated` even though 
the writer verified the rows that it added.
+
+When a constraint becomes enforced, either by being added with `enforced` set 
to true or by `enforced` changing from false to true, writers should validate 
the table and record `validated`. A writer that does not validate records 
`unvalidated`, and the constraint remains `unvalidated` until a later 
validation records `validated`.
+
+A writer does not have to check every row in a single scan. After checking 
every row in an ancestor snapshot, a writer may check only the rows added 
between that ancestor and the current snapshot and record `validated` for the 
current snapshot. This allows a validation to finish on a table that is written 
concurrently, without blocking writes or restarting the scan.

Review Comment:
   We had a long discussion about `uuid` and `validated` vs `valid`.
   If I remember correctly we landed on `uuid`s being `validated` if they the 
parent is `validated` and the current commit `valid`.
   We should mention this explicitly here.



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).

Review Comment:
   +1 on listing allowed constraint types.s
   +1 on `primary-key` fields must have `required` 



##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.

Review Comment:
   I understand why we are doing it but it feels a asymmetric.
   Dropping a single column will drop single-column constraints only of that 
column 'silently'.
   Dropping a single column with a multi-column constraint will fail 'loud'.
   
   I think we should apply the same rule to single-column constraints as well, 
that the column can only be dropped if the constraints removal is in the same 
change. This also feels more deliberate.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to