res-life opened a new issue, #18131:
URL: https://github.com/apache/iceberg/issues/18131
### Apache Iceberg version
1.10.1 and 1.11.0
### Query engine
Apache Spark 3.5.8, Scala 2.12
### Please describe the bug
When the JVM system property `spark.testing` is present, a copy-on-write
`DELETE` on an Iceberg
V3 table fails Spark's per-rule plan validation. The
`GroupBasedRowLevelOperationScanPlanning` optimizer rule produces an
intermediate `ReplaceData`
plan that Spark considers unresolved after the scan relation contains
`_row_id` and
`_last_updated_sequence_number`.
The same query succeeds when the `spark.testing` property is omitted. Note
that
`-Dspark.testing=false` does not disable this validation because Spark
checks whether the property
is present.
This reproduces with both Iceberg 1.10.1 and 1.11.0 on Spark 3.5.8. The
equivalent cases pass with
`spark.testing=true` on these combinations:
- Spark 4.0.1 / Iceberg 1.10.1
- Spark 4.0.2 / Iceberg 1.10.1
- Spark 4.0.2 / Iceberg 1.11.0
- Spark 4.1.1 / Iceberg 1.11.0
This appears specific to Iceberg's Spark 3.5 row-lineage compatibility path.
In particular,
`SparkCopyOnWriteScan.rowLineageAsDataCols` removes the `__metadata_col`
marker from the row ID and
sequence-number fields. After Spark replaces the `DataSourceV2Relation` with
a scan relation,
`ReplaceData.outputResolved` treats those fields as data columns, so the
command is temporarily
unresolved. Normal execution does not expose the problem because validation
after every optimizer
rule is enabled only in Spark testing mode.
### Reproduction
Save the following as `repro.py`:
```python
import shutil
import tempfile
from pyspark.sql import SparkSession
warehouse = tempfile.mkdtemp(prefix="iceberg-row-lineage-")
spark = (
SparkSession.builder
.master("local[2]")
.config(
"spark.sql.extensions",
"org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions")
.config(
"spark.sql.catalog.spark_catalog",
"org.apache.iceberg.spark.SparkSessionCatalog")
.config("spark.sql.catalog.spark_catalog.type", "hadoop")
.config("spark.sql.catalog.spark_catalog.warehouse", warehouse)
.getOrCreate()
)
try:
spark.sql(
"CREATE TABLE t (id BIGINT, v INT) USING ICEBERG "
"TBLPROPERTIES ('format-version'='3', "
"'write.delete.mode'='copy-on-write')")
spark.sql("INSERT INTO t SELECT id, CAST(id AS INT) FROM range(0, 2048)")
spark.sql("DELETE FROM t WHERE v % 3 = 0")
finally:
spark.stop()
shutil.rmtree(warehouse, ignore_errors=True)
```
Run it with Spark 3.5.8:
```bash
spark-submit \
--driver-java-options=-Dspark.testing=true \
--packages org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.11.0 \
repro.py
```
The failure is:
```text
org.apache.spark.SparkException: [PLAN_VALIDATION_FAILED_RULE_IN_BATCH]
Rule
org.apache.spark.sql.execution.datasources.v2.GroupBasedRowLevelOperationScanPlanning
in batch Early Filter and Projection Push-Down generated an invalid plan:
The plan becomes unresolved: 'ReplaceData ...
+- Filter ...
+- RelationV2[id, v, _file, _pos, _row_id, _last_updated_sequence_number]
...
```
### Expected behavior
Each optimizer rule should leave the plan resolved when `spark.testing`
validation is enabled, and
the V3 copy-on-write `DELETE` should complete.
### Additional context
The Spark 3.5 row-lineage support was introduced in
[apache/iceberg#12736](https://github.com/apache/iceberg/pull/12736). Spark
4 uses the native
metadata-column semantics introduced by
[SPARK-50820](https://github.com/apache/spark/pull/49493) and does not
reproduce this failure.
A regression test should run a V3 copy-on-write row-level operation with the
`spark.testing` JVM
property enabled.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]