stevenzwu commented on code in PR #18308: URL: https://github.com/apache/iceberg/pull/18308#discussion_r4137047429
########## format/spec.md: ########## @@ -840,7 +840,7 @@ For example, stats for a `required` `int` field named `id` with field-id `2` are // null_value_count is only used for optional fields // nan_value_count is only used for float and double - // avg_value_size_in_bytes is only used for variable length types + // total_bytes is only used for variable length types Review Comment: if we drop this comment, we also need to adjust the description cell in the table above, which right now limits to variable length types. > Though, for non-variable length types I don't know if it's super useful and can just be computed from constants and record count. But again, spec probably doesn't need to have an opinion here. Amogh's argument seems reasonable for spec. What should be the reference implementation behavior? Here are some pros and cons of always storing it regardless field type - pro: engines don't need to extract the metric differently based on the field type * con: store the metric on disk that can be easily derived -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
