szehon-ho commented on code in PR #18057:
URL: https://github.com/apache/iceberg/pull/18057#discussion_r4140007101


##########
format/udf-spec.md:
##########
@@ -118,19 +122,25 @@ following fields required. Any other fields must be 
ignored.
   e.g., `{ "type": "struct", "fields": [ { "name": "id", "type": "int" }, { 
"name": "name", "type": "string" } ] }`
 
 #### Definition ID
-The `definition-id` is a canonical string derived from the parameter types, 
formatted as a comma-separated list with no
-spaces. Each type uses the following string representation:
+The `definition-id` is a canonical string derived from the parameter types, 
formatted as a comma-separated list. The
+separators that this format adds must not be followed by a space. Each type 
uses the following string representation:
 
 * Primitives and semi-structured: the type name (e.g., `int`, `variant`)
 * List: `list<element-type>` (e.g., `list<int>`)
 * Map: `map<key-type,value-type>` (e.g., `map<string,int>`)
 * Struct: `struct<name1:type1,name2:type2,...>` with field names and types 
(e.g., `struct<id:int,name:string>`)
 
+In a struct field name, `\`, `:`, `,`, `<`, and `>` must each be escaped with 
a preceding `\`. Without escaping, a

Review Comment:
   maybe simplify by getting rid of 'without escaping'.  I think the spec can 
be more minimalistic



##########
format/udf-spec.md:
##########
@@ -107,7 +107,11 @@ Notes:
 Types are based on the [Iceberg 
Type](https://iceberg.apache.org/spec/#schemas-and-data-types).
 
 Primitive and semi-structured type strings are encoded based on [Iceberg Type 
JSON Representation][iceberg-type-json]
-(e.g., `int`, `string`, `timestamp`, `decimal(9,2)`, `variant`). Type strings 
must contain no spaces or quote characters.
+(e.g., `int`, `string`, `timestamp`, `decimal(9, 2)`, `variant`). Type strings 
must contain no quote characters.
+
+Writers must use Iceberg's canonical serialized form. Readers should accept 
optional whitespace around parameters and

Review Comment:
   can you specify what do you mean in terms of spec ?  'canonical serialized 
form'?



##########
format/udf-spec.md:
##########
@@ -118,19 +122,25 @@ following fields required. Any other fields must be 
ignored.
   e.g., `{ "type": "struct", "fields": [ { "name": "id", "type": "int" }, { 
"name": "name", "type": "string" } ] }`
 
 #### Definition ID
-The `definition-id` is a canonical string derived from the parameter types, 
formatted as a comma-separated list with no
-spaces. Each type uses the following string representation:
+The `definition-id` is a canonical string derived from the parameter types, 
formatted as a comma-separated list. The
+separators that this format adds must not be followed by a space. Each type 
uses the following string representation:

Review Comment:
   It's a bit confusing (what's this format, whats separators).  Would this 
make sense?
   
   The definition-ID must not insert spaces after its separators (commas or 
colons). Embedded type strings retain their canonical formatting.
   
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to