szehon-ho commented on code in PR #18250:
URL: https://github.com/apache/iceberg/pull/18250#discussion_r4109630213
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -221,17 +219,16 @@ public Table createTable(
}
@Override
- public StagedTable stageCreate(
- Identifier ident, StructType schema, Transform[] transforms, Map<String,
String> properties)
- throws TableAlreadyExistsException {
- Schema icebergSchema = SparkSchemaUtil.convert(schema);
+ public StagedTable stageCreate(Identifier ident, TableInfo tableInfo)
+ throws TableAlreadyExistsException, NoSuchNamespaceException {
+ Schema icebergSchema = SparkSchemaUtil.convert(tableInfo);
Review Comment:
This would make CTAS fail when the selected schema contains column defaults.
Should we use `SparkSchemaUtil.convert(tableInfo.schema())` here to ignore
defaults in the selected schema?
Spark reconstructs these defaults with a SQL string but no expression, so
selecting a column with `DEFAULT 7` would reach `Unsupported default value
expression: 7`.
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkTypeToType.java:
##########
@@ -201,4 +228,43 @@ private static EdgeAlgorithm
convertAlgorithm(EdgeInterpolationAlgorithm algorit
"Iceberg does not support Spark geography edge algorithm: " +
algorithm);
}
}
+
+ private void convertDefaultValue(Types.NestedField.Builder icebergField,
StructField sparkField) {
+ Column column = nameToColumnMap.get(sparkField.name());
+
+ if (column == null) {
+ return;
+ }
+
+ ColumnDefaultValue columnDefaultValue = column.defaultValue();
+
+ if (columnDefaultValue == null) {
+ return;
+ }
+
+ if (columnDefaultValue.getExpression() == null &&
columnDefaultValue.getSql() != null) {
+ throw new UnsupportedOperationException(
+ "Unsupported default value expression: " +
columnDefaultValue.getSql());
+ }
+
+ if (columnDefaultValue.getExpression() != null
+ && !SparkV2Filters.isLiteral(columnDefaultValue.getExpression())) {
+ throw new UnsupportedOperationException(
+ "Default value expressions are not supported in Iceberg");
+ }
+
+ // the value is equivalent to the initial value in Iceberg
+ Literal<?> initialValue = columnDefaultValue.getValue();
+ if (initialValue != null && initialValue.value() != null) {
+ icebergField.withInitialDefault(
+
Expressions.lit(SparkV2Filters.convertLiteral(columnDefaultValue.getValue())));
+ }
+
+ // the expression is evaluated for future writes
+ LiteralValue<?> writeDefault = (LiteralValue<?>)
columnDefaultValue.getExpression();
Review Comment:
Could we cast to `Literal<?>` here?
Spark translates Boolean defaults such as `CREATE TABLE t (enabled BOOLEAN
DEFAULT TRUE) USING iceberg TBLPROPERTIES ('format-version' = '3')` into
`AlwaysTrue` or `AlwaysFalse`. These implement `Literal` but not
`LiteralValue`, so this cast would throw `ClassCastException`.
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkTypeToType.java:
##########
@@ -201,4 +228,43 @@ private static EdgeAlgorithm
convertAlgorithm(EdgeInterpolationAlgorithm algorit
"Iceberg does not support Spark geography edge algorithm: " +
algorithm);
}
}
+
+ private void convertDefaultValue(Types.NestedField.Builder icebergField,
StructField sparkField) {
+ Column column = nameToColumnMap.get(sparkField.name());
+
+ if (column == null) {
+ return;
+ }
+
+ ColumnDefaultValue columnDefaultValue = column.defaultValue();
+
+ if (columnDefaultValue == null) {
+ return;
+ }
+
+ if (columnDefaultValue.getExpression() == null &&
columnDefaultValue.getSql() != null) {
+ throw new UnsupportedOperationException(
+ "Unsupported default value expression: " +
columnDefaultValue.getSql());
+ }
+
+ if (columnDefaultValue.getExpression() != null
+ && !SparkV2Filters.isLiteral(columnDefaultValue.getExpression())) {
+ throw new UnsupportedOperationException(
+ "Default value expressions are not supported in Iceberg");
+ }
+
+ // the value is equivalent to the initial value in Iceberg
+ Literal<?> initialValue = columnDefaultValue.getValue();
+ if (initialValue != null && initialValue.value() != null) {
+ icebergField.withInitialDefault(
+
Expressions.lit(SparkV2Filters.convertLiteral(columnDefaultValue.getValue())));
Review Comment:
Could we normalize `Byte` and `Short` defaults to `Integer`, matching the
column-type conversion?
For example, `CREATE TABLE t (n TINYINT DEFAULT 1) USING iceberg
TBLPROPERTIES ('format-version' = '3')` passes the default as a boxed `Byte`.
`convertLiteral` leaves it unchanged, and `Expressions.lit` rejects it, causing
table creation to fail. `SMALLINT DEFAULT 1` has the same issue with `Short`.
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -259,15 +255,15 @@ public StagedTable stageReplace(
}
@Override
- public StagedTable stageCreateOrReplace(
- Identifier ident, StructType schema, Transform[] transforms, Map<String,
String> properties) {
- Schema icebergSchema = SparkSchemaUtil.convert(schema);
+ public StagedTable stageCreateOrReplace(Identifier ident, TableInfo
tableInfo)
+ throws NoSuchNamespaceException {
+ Schema icebergSchema = SparkSchemaUtil.convert(tableInfo.schema());
Review Comment:
We should preserve defaults explicitly declared in `REPLACE TABLE` and
`CREATE OR REPLACE TABLE`.
Both replacement methods currently use schema-only conversion, so a
declaration such as `id INT DEFAULT 7` silently loses its default. These
methods also handle RTAS, so we need to distinguish explicit defaults from
inherited query metadata.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]