andygrove commented on issue #5121: URL: https://github.com/apache/datafusion-comet/issues/5121#issuecomment-5786236373
Hi @Visorgood, sorry for the slow reply, and thanks for offering to help. This issue is out of date. Since it was filed, stages 1 and 2 have been done for Iceberg: `AppendData`, `OverwriteByExpression`, `OverwritePartitionsDynamic`, and copy-on-write `ReplaceData` all go through Comet's split write/commit plan and the native iceberg-rust writer. That work was #4658, #5298, and #5361. It's experimental and off by default for now (`spark.comet.write.iceberg.splitOperator.enabled` and `spark.comet.iceberg.write.enabled`). The docs are at `docs/source/user-guide/latest/iceberg-writes.md`. The remaining Iceberg write work is tracked in #5649, split into phases (correctness, failure handling, coverage, enabling by default, performance). If you'd like to help, anything unassigned there is a great place to start. For example, #5318 (native `MergeRowsExec`) is the natural next step for row-level MERGE, and #5306 (reconciling the two config namespaces for native writes) and #5643 (a keep-or-lift decision for each remaining eligibility restriction) are more self-contained. Comment on the one you pick so we can assign it to you. Stage 3 (generic V2 writes) doesn't have a clear target right now, since Spark's built-in file sources write through the V1 path (#1625). So I'm going to close this issue in favor of #5649. If someone has a V2 sink that would benefit, we can open a focused issue for it. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
