szehon-ho commented on code in PR #18250:
URL: https://github.com/apache/iceberg/pull/18250#discussion_r4109630213


##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -221,17 +219,16 @@ public Table createTable(
   }
 
   @Override
-  public StagedTable stageCreate(
-      Identifier ident, StructType schema, Transform[] transforms, Map<String, 
String> properties)
-      throws TableAlreadyExistsException {
-    Schema icebergSchema = SparkSchemaUtil.convert(schema);
+  public StagedTable stageCreate(Identifier ident, TableInfo tableInfo)
+      throws TableAlreadyExistsException, NoSuchNamespaceException {
+    Schema icebergSchema = SparkSchemaUtil.convert(tableInfo);

Review Comment:
   This would make CTAS fail when the selected schema contains column defaults. 
Should we use `SparkSchemaUtil.convert(tableInfo.schema())` here to ignore 
defaults in the selected schema?
   
   Spark reconstructs these defaults with a SQL string but no expression, so 
selecting a column with `DEFAULT 7` would reach `Unsupported default value 
expression: 7`.



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkTypeToType.java:
##########
@@ -201,4 +228,43 @@ private static EdgeAlgorithm 
convertAlgorithm(EdgeInterpolationAlgorithm algorit
             "Iceberg does not support Spark geography edge algorithm: " + 
algorithm);
     }
   }
+
+  private void convertDefaultValue(Types.NestedField.Builder icebergField, 
StructField sparkField) {
+    Column column = nameToColumnMap.get(sparkField.name());
+
+    if (column == null) {
+      return;
+    }
+
+    ColumnDefaultValue columnDefaultValue = column.defaultValue();
+
+    if (columnDefaultValue == null) {
+      return;
+    }
+
+    if (columnDefaultValue.getExpression() == null && 
columnDefaultValue.getSql() != null) {
+      throw new UnsupportedOperationException(
+          "Unsupported default value expression: " + 
columnDefaultValue.getSql());
+    }
+
+    if (columnDefaultValue.getExpression() != null
+        && !SparkV2Filters.isLiteral(columnDefaultValue.getExpression())) {
+      throw new UnsupportedOperationException(
+          "Default value expressions are not supported in Iceberg");
+    }
+
+    // the value is equivalent to the initial value in Iceberg
+    Literal<?> initialValue = columnDefaultValue.getValue();
+    if (initialValue != null && initialValue.value() != null) {
+      icebergField.withInitialDefault(
+          
Expressions.lit(SparkV2Filters.convertLiteral(columnDefaultValue.getValue())));
+    }
+
+    // the expression is evaluated for future writes
+    LiteralValue<?> writeDefault = (LiteralValue<?>) 
columnDefaultValue.getExpression();

Review Comment:
   Could we cast to `Literal<?>` here?
   
   Spark translates Boolean defaults such as `CREATE TABLE t (enabled BOOLEAN 
DEFAULT TRUE) USING iceberg TBLPROPERTIES ('format-version' = '3')` into 
`AlwaysTrue` or `AlwaysFalse`. These implement `Literal` but not 
`LiteralValue`, so this cast would throw `ClassCastException`.



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkTypeToType.java:
##########
@@ -201,4 +228,43 @@ private static EdgeAlgorithm 
convertAlgorithm(EdgeInterpolationAlgorithm algorit
             "Iceberg does not support Spark geography edge algorithm: " + 
algorithm);
     }
   }
+
+  private void convertDefaultValue(Types.NestedField.Builder icebergField, 
StructField sparkField) {
+    Column column = nameToColumnMap.get(sparkField.name());
+
+    if (column == null) {
+      return;
+    }
+
+    ColumnDefaultValue columnDefaultValue = column.defaultValue();
+
+    if (columnDefaultValue == null) {
+      return;
+    }
+
+    if (columnDefaultValue.getExpression() == null && 
columnDefaultValue.getSql() != null) {
+      throw new UnsupportedOperationException(
+          "Unsupported default value expression: " + 
columnDefaultValue.getSql());
+    }
+
+    if (columnDefaultValue.getExpression() != null
+        && !SparkV2Filters.isLiteral(columnDefaultValue.getExpression())) {
+      throw new UnsupportedOperationException(
+          "Default value expressions are not supported in Iceberg");
+    }
+
+    // the value is equivalent to the initial value in Iceberg
+    Literal<?> initialValue = columnDefaultValue.getValue();
+    if (initialValue != null && initialValue.value() != null) {
+      icebergField.withInitialDefault(
+          
Expressions.lit(SparkV2Filters.convertLiteral(columnDefaultValue.getValue())));

Review Comment:
   Could we normalize `Byte` and `Short` defaults to `Integer`, matching the 
column-type conversion?
   
   For example, `CREATE TABLE t (n TINYINT DEFAULT 1) USING iceberg 
TBLPROPERTIES ('format-version' = '3')` passes the default as a boxed `Byte`. 
`convertLiteral` leaves it unchanged, and `Expressions.lit` rejects it, causing 
table creation to fail. `SMALLINT DEFAULT 1` has the same issue with `Short`.



##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -259,15 +255,15 @@ public StagedTable stageReplace(
   }
 
   @Override
-  public StagedTable stageCreateOrReplace(
-      Identifier ident, StructType schema, Transform[] transforms, Map<String, 
String> properties) {
-    Schema icebergSchema = SparkSchemaUtil.convert(schema);
+  public StagedTable stageCreateOrReplace(Identifier ident, TableInfo 
tableInfo)
+      throws NoSuchNamespaceException {
+    Schema icebergSchema = SparkSchemaUtil.convert(tableInfo.schema());

Review Comment:
   We should preserve defaults explicitly declared in `REPLACE TABLE` and 
`CREATE OR REPLACE TABLE`.
   
   Both replacement methods currently use schema-only conversion, so a 
declaration such as `id INT DEFAULT 7` silently loses its default. These 
methods also handle RTAS, so we need to distinguish explicit defaults from 
inherited query metadata.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to