RussellSpitzer commented on code in PR #17868: URL: https://github.com/apache/iceberg/pull/17868#discussion_r3884357200
########## docs/docs/spark-procedures.md: ########## @@ -415,6 +415,7 @@ Iceberg can compact data files in parallel using Spark with the `rewriteDataFile | `output-spec-id` | current partition spec id | Identifier of the output partition spec. Data will be reorganized during the rewrite to align with the output partitioning. | | `remove-dangling-deletes` | false | Remove dangling position and equality deletes after rewriting. A delete file is considered dangling if it does not apply to any live data files. Enabling this will generate an additional commit for the removal. | | `max-files-to-rewrite` | null | This option sets an upper limit on the number of eligible files that will be rewritten. If this option is not specified, all eligible files will be rewritten. | +| `executor-cache.delete-files.enabled` | false | Use the executor cache for delete files while rewriting. Enable this when the same delete file applies to many data files, which is common with equality deletes | Review Comment: I'd rather we not use a "." property if we can help. The partial progress ones are not a good example based on all the other procedure options we have use the kabob thing and I think it's because we originally defined them in "action". enable-executor-cache is also fine, but I don't think it really explains what the option is doing. It only effects deletes and specifically delete files so I'd try to keep the name tied to that functionality r -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
