szehon-ho commented on code in PR #17954:
URL: https://github.com/apache/iceberg/pull/17954#discussion_r4170104834
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkSessionCatalog.java:
##########
@@ -252,6 +253,23 @@ public Table createTable(
}
}
+ @Override
+ public Table createTableLike(Identifier ident, TableInfo tableInfo, Table
sourceTable)
+ throws TableAlreadyExistsException, NoSuchNamespaceException {
+ checkViewNotExists(ident);
+
+ String provider = tableInfo.properties().get("provider");
+ if (provider == null) {
+ provider = sourceTable.properties().get("provider");
+ }
+
+ if (useIceberg(provider)) {
+ return icebergCatalog.createTableLike(ident, tableInfo, sourceTable);
+ } else {
+ return getSessionCatalog().createTableLike(ident, tableInfo,
sourceTable);
Review Comment:
Please use the session catalog's supported `createTable` API here, carrying
over the resolved provider. Spark 4.2's `V2SessionCatalog` does not implement
`createTableLike`, so an Iceberg source cloned into `spark_catalog` with `USING
parquet` throws `UnsupportedOperationException`. Please add an integration test
using an Iceberg source; the current provider test follows Spark's V1 path.
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -220,6 +294,93 @@ public Table createTable(
}
}
+ private static PartitionSpec copyPartitionSpec(Schema schema, PartitionSpec
sourceSpec) {
+ UnboundPartitionSpec spec = sourceSpec.toUnbound();
+ for (int index = sourceSpec.fields().size() - 1; index >= 0; index -= 1) {
+ if (schema.findField(sourceSpec.fields().get(index).sourceId()) == null)
{
+ spec.fields().remove(index);
+ }
+ }
+
+ return spec.bind(schema);
+ }
+
+ private static SortOrder copySortOrder(Schema schema, SortOrder
sourceSortOrder) {
+ if (sourceSortOrder.isUnsorted()) {
+ return SortOrder.unsorted();
+ }
+
+ SortOrder.Builder builder = SortOrder.builderFor(schema);
+ for (SortField field : sourceSortOrder.fields()) {
+ String sourceName = schema.findColumnName(field.sourceId());
Review Comment:
Please omit sort fields whose source IDs are absent from the pinned schema.
If a column is added and included in the sort order after a snapshot, cloning
that older snapshot returns null here and `Expressions.transform` throws “Name
cannot be null”. Please add coverage for this case.
##########
spark/v4.2/spark/src/main/java/org/apache/iceberg/spark/SparkCatalog.java:
##########
@@ -206,11 +222,69 @@ public Table createTable(
Identifier ident, StructType schema, Transform[] transforms, Map<String,
String> properties)
throws TableAlreadyExistsException {
Schema icebergSchema = SparkSchemaUtil.convert(schema);
+ return createTable(
+ ident,
+ icebergSchema,
+ Spark3Util.toPartitionSpec(icebergSchema, transforms),
+ properties,
+ SortOrder.unsorted());
+ }
+
+ @Override
+ public Table createTableLike(Identifier ident, TableInfo tableInfo, Table
sourceTable)
+ throws TableAlreadyExistsException, NoSuchNamespaceException {
+ // Spark intentionally excludes the source table's properties from
tableInfo and leaves it to
+ // the connector to decide which to clone via sourceTable. Clone the
source Iceberg table's
+ // schema, partition spec, properties and sort order, then let
user-specified LIKE options (in
+ // tableInfo) take precedence.
+ Schema icebergSchema;
+ PartitionSpec spec;
+ Map<String, String> properties = Maps.newHashMap();
+ SortOrder sortOrder = SortOrder.unsorted();
+
+ if (sourceTable instanceof SparkTable sparkTable) {
+ org.apache.iceberg.Table sourceIcebergTable = sparkTable.table();
+ icebergSchema = sparkTable.icebergSchema();
+ spec = copyPartitionSpec(icebergSchema, sourceIcebergTable.spec());
+ properties.putAll(sourceIcebergTable.properties());
+ properties.remove(TableProperties.WRITE_METADATA_LOCATION);
+ properties.remove(TableProperties.WRITE_DATA_LOCATION);
+ properties.remove(TableProperties.OBJECT_STORE_PATH);
+ properties.remove(TableProperties.WRITE_FOLDER_STORAGE_LOCATION);
+ properties.remove(TableProperties.DEFAULT_NAME_MAPPING);
+ properties.put(
+ TableProperties.FORMAT_VERSION,
+ String.valueOf(TableUtil.formatVersion(sourceIcebergTable)));
+ sortOrder = copySortOrder(icebergSchema, sourceIcebergTable.sortOrder());
+ } else {
+ icebergSchema =
+ copyColumnDefaults(tableInfo.schema(),
SparkSchemaUtil.convert(tableInfo.schema()));
Review Comment:
Please rebase on #18250 and reuse
`SparkSchemaUtil.convertWithDefaults(tableInfo)` for non-Iceberg sources. This
handles initial and write defaults through the typed API and makes `CREATE
TABLE LIKE` consistent with `CREATE TABLE`. The custom
`copyColumnDefaults`/`columnDefault` helpers and SQL parsing can then be
removed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]