szehon-ho commented on code in PR #18057:
URL: https://github.com/apache/iceberg/pull/18057#discussion_r4140007101
##########
format/udf-spec.md:
##########
@@ -118,19 +122,25 @@ following fields required. Any other fields must be
ignored.
e.g., `{ "type": "struct", "fields": [ { "name": "id", "type": "int" }, {
"name": "name", "type": "string" } ] }`
#### Definition ID
-The `definition-id` is a canonical string derived from the parameter types,
formatted as a comma-separated list with no
-spaces. Each type uses the following string representation:
+The `definition-id` is a canonical string derived from the parameter types,
formatted as a comma-separated list. The
+separators that this format adds must not be followed by a space. Each type
uses the following string representation:
* Primitives and semi-structured: the type name (e.g., `int`, `variant`)
* List: `list<element-type>` (e.g., `list<int>`)
* Map: `map<key-type,value-type>` (e.g., `map<string,int>`)
* Struct: `struct<name1:type1,name2:type2,...>` with field names and types
(e.g., `struct<id:int,name:string>`)
+In a struct field name, `\`, `:`, `,`, `<`, and `>` must each be escaped with
a preceding `\`. Without escaping, a
Review Comment:
maybe simplify by getting rid of 'without escaping'. I think the spec can
be more minimalistic
##########
format/udf-spec.md:
##########
@@ -107,7 +107,11 @@ Notes:
Types are based on the [Iceberg
Type](https://iceberg.apache.org/spec/#schemas-and-data-types).
Primitive and semi-structured type strings are encoded based on [Iceberg Type
JSON Representation][iceberg-type-json]
-(e.g., `int`, `string`, `timestamp`, `decimal(9,2)`, `variant`). Type strings
must contain no spaces or quote characters.
+(e.g., `int`, `string`, `timestamp`, `decimal(9, 2)`, `variant`). Type strings
must contain no quote characters.
+
+Writers must use Iceberg's canonical serialized form. Readers should accept
optional whitespace around parameters and
Review Comment:
can you specify what do you mean in terms of spec ? 'canonical serialized
form'?
##########
format/udf-spec.md:
##########
@@ -118,19 +122,25 @@ following fields required. Any other fields must be
ignored.
e.g., `{ "type": "struct", "fields": [ { "name": "id", "type": "int" }, {
"name": "name", "type": "string" } ] }`
#### Definition ID
-The `definition-id` is a canonical string derived from the parameter types,
formatted as a comma-separated list with no
-spaces. Each type uses the following string representation:
+The `definition-id` is a canonical string derived from the parameter types,
formatted as a comma-separated list. The
+separators that this format adds must not be followed by a space. Each type
uses the following string representation:
Review Comment:
It's a bit confusing (what's this format, whats separators). Would this
make sense?
The definition-ID must not insert spaces after its separators (commas or
colons). Embedded type strings retain their canonical formatting.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]