andygrove commented on PR #5724: URL: https://github.com/apache/datafusion-comet/pull/5724#issuecomment-5876500709
This is a light fully automated review since there are so many PRs open. Since `fix: ignore unused bloom max byte caps`, `requireNativeSupportedBloomFilterProperties` only range-checks `write.parquet.bloom-filter-max-bytes` when a `write.parquet.bloom-filter-enabled.column.*` key is present. Without one, any parseable int stays native, and `unused bloom-filter max-bytes does not force fallback` at `CometIcebergWriteDetectionSuite.scala:293` pins that for `100`. The user guide still describes the earlier rule. The eligibility row at `docs/source/user-guide/latest/iceberg-writes.md:159` says every value other than a power of two in [32, 128 MiB] falls back to iceberg-java, and the sizing section at lines 250-253 says `CometIcebergWriteExec` is not used for a non-power-of-two or out-of-range value. So a table with `max-bytes=100` and no Bloom columns is documented as falling back but writes natively. The comment on `bloom_filter_max_bytes` at `native/proto/src/proto/operator.proto:686` makes the same claim. Could all three say the range only applies once a Bloom column is configured, and match whatever the unused-cap branch accepts after the non-positive check sunchao asked for? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
