dejankrak-db commented on code in PR #17918:
URL: https://github.com/apache/iceberg/pull/17918#discussion_r4224370922


##########
format/spec.md:
##########
@@ -322,6 +325,35 @@ For `geography` types, an additional parameter A specifies 
an algorithm for inte
 * `andoyer`: Thomas, Paul D. Mathematical models for navigation systems. US 
Naval Oceanographic Office, 1965.
 * `karney`: [Karney, Charles FF. "Algorithms for geodesics." Journal of 
Geodesy 87 (2013): 
43-55](https://link.springer.com/content/pdf/10.1007/s00190-012-0578-z.pdf), 
and [GeographicLib](https://geographiclib.sourceforge.io/)
 
+#### File Type
+
+A **`file`** represents a range of bytes that may be stored inline as a value 
or as a reference to an external file. The `file` type and its value semantics 
are defined by the `FILE` logical type in the [Parquet 
project](https://github.com/apache/parquet-format/blob/master/LogicalTypes.md#file).
+
+A `file` value has a fixed set of sub-fields. The sub-fields are implicit: 
they are not represented in the Iceberg schema and cannot be added, removed, 
reordered, or promoted. Their names, types, and field-ID offsets are:
+
+| Sub-field      | ID offset | Type     |
+|----------------|-----------|----------|
+| `uri`          | +1        | `string` |
+| `offset`       | +2        | `long`   |
+| `size`         | +3        | `long`   |
+| `content_type` | +4        | `string` |
+| `checksum`     | +5        | `string` |
+| `inline`       | +6        | `binary` |
+
+A `file` field reserves the root field's ID plus six consecutive IDs for its 
sub-fields, assigned by the offsets above. Adding a `file` field must advance 
`last-column-id` to account for all 7 IDs.
+
+The `uri` field may contain absolute or relative references. Implementations 
that receive a relative path should resolve the path against the table location 
(see [Path Resolution](#path-resolution)).

Review Comment:
   On the should/must question for relative `uri`: Delta's `file` RFC 
(delta-io/delta#7148, with a follow-up extension in #7724) resolves a relative 
`uri` by joining it to the table root with a single `/`, with no normalization, 
so the same bytes resolve identically when the table is read as Delta or 
Iceberg. That only works if every reader applies the same rule, so I'd support 
"must" and a pointer to Path Resolution rather than leaving the interpretation 
to each engine.



##########
format/spec.md:
##########
@@ -1566,6 +1599,7 @@ Lists must use the [3-level 
representation](https://github.com/apache/parquet-fo
 | **`variant`**      | `group` with `metadata` and `value` fields. `metadata` 
and `value` must not be assigned field IDs and the fields are accessed through 
names. | `VARIANT`                                   | See Parquet docs for 
[Variant 
encoding](https://github.com/apache/parquet-format/blob/master/VariantEncoding.md)
 and [Variant shredding 
encoding](https://github.com/apache/parquet-format/blob/master/VariantShredding.md).
 |
 | **`geometry`**     | `binary`                                                
                                                                                
     | `GEOMETRY`                                  | WKB format, see [Appendix 
G](#appendix-g-geospatial-notes).                             |
 | **`geography`**    | `binary`                                                
                                                                                
     | `GEOGRAPHY`                                 | WKB format, see [Appendix 
G](#appendix-g-geospatial-notes).                             |
+| **`file`**         | `group` with the `file` sub-fields. Sub-fields must be 
assigned field IDs.                                                             
      | `FILE`                                      | See Parquet docs for the 
[`FILE` 
type](https://github.com/apache/parquet-format/blob/master/LogicalTypes.md#file)
 and [File Type](#file-type). |

Review Comment:
   The Parquet and ORC format rows require the reserved field IDs on the 
sub-fields. Is an Iceberg reader expected to resolve the sub-fields by ID only, 
or may it fall back to the sub-field names when the IDs are absent? Delta 
writes the IDs only when column mapping is enabled, and resolves the sub-fields 
by name otherwise, so it matters which data files a reader must accept.



##########
format/spec.md:
##########
@@ -322,6 +325,35 @@ For `geography` types, an additional parameter A specifies 
an algorithm for inte
 * `andoyer`: Thomas, Paul D. Mathematical models for navigation systems. US 
Naval Oceanographic Office, 1965.
 * `karney`: [Karney, Charles FF. "Algorithms for geodesics." Journal of 
Geodesy 87 (2013): 
43-55](https://link.springer.com/content/pdf/10.1007/s00190-012-0578-z.pdf), 
and [GeographicLib](https://geographiclib.sourceforge.io/)
 
+#### File Type
+
+A **`file`** represents a range of bytes that may be stored inline as a value 
or as a reference to an external file. The `file` type and its value semantics 
are defined by the `FILE` logical type in the [Parquet 
project](https://github.com/apache/parquet-format/blob/master/LogicalTypes.md#file).
+
+A `file` value has a fixed set of sub-fields. The sub-fields are implicit: 
they are not represented in the Iceberg schema and cannot be added, removed, 
reordered, or promoted. Their names, types, and field-ID offsets are:
+
+| Sub-field      | ID offset | Type     |
+|----------------|-----------|----------|
+| `uri`          | +1        | `string` |
+| `offset`       | +2        | `long`   |
+| `size`         | +3        | `long`   |
+| `content_type` | +4        | `string` |
+| `checksum`     | +5        | `string` |
+| `inline`       | +6        | `binary` |
+
+A `file` field reserves the root field's ID plus six consecutive IDs for its 
sub-fields, assigned by the offsets above. Adding a `file` field must advance 
`last-column-id` to account for all 7 IDs.
+
+The `uri` field may contain absolute or relative references. Implementations 
that receive a relative path should resolve the path against the table location 
(see [Path Resolution](#path-resolution)).
+
+Statistics for `file` are tracked using separate field stats for each 
sub-field. Writers should produce statistics for `uri`, `content_type`, and 
`inline` fields; other fields may be omitted.
+
+A `file` column is subject to the following restrictions:
+
+* No type promotion to or from `file` is defined.
+* Equality, ordering, and hashing are not defined for `file` objects
+* A `file` column cannot be an identifier field or a source for partition or 
sort transforms.
+* A `file` is not interchangeable with a `struct`; a `struct` with the same 
sub-fields is not equivalent to a `file`.
+* The behavior for a `file` entry that references a non-existent file is not 
defined.

Review Comment:
   Daniel's wording that a `file` is "purely a reference" with no guarantee of 
existence helps. Should the spec also say that a relative `uri` is not 
table-managed data, so orphan-file removal under the table location must not 
treat bytes it points to as removable? Otherwise cleanup tools could delete 
bytes a table still references.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to