breken-ai opened a new pull request, #4030:
URL: https://github.com/apache/iceberg-python/pull/4030

   <!--
   Thanks for opening a pull request!
   -->
   
   <!-- In the case this PR will resolve an issue, please replace 
${GITHUB_ISSUE_ID} below with the actual Github issue id. -->
   <!-- Closes #${GITHUB_ISSUE_ID} -->
   
   # Rationale for this change
   
   Appending a NaN to a table with an identity partition on a `double` or 
`float` column fails:
   
   ```python
   schema = Schema(NestedField(1, "id", IntegerType()), NestedField(2, "value", 
DoubleType()))
   spec = PartitionSpec(PartitionField(source_id=2, field_id=1000, 
transform=IdentityTransform(), name="value"))
   tbl = catalog.create_table("default.t", schema=schema, partition_spec=spec)
   tbl.append(pa.table({"id": [1, 2, 3], "value": [1.0, float("nan"), None]}))
   # ZeroDivisionError: division by zero  (bin_pack_arrow_table: tbl.nbytes / 
tbl.num_rows)
   ```
   
   `_determine_partitions` groups the rows by partition value, which yields one 
NaN group, and then selects each group's rows with `pc.field(name) == value`. 
NaN never compares equal to itself, so the NaN partition selects zero rows and 
the empty table reaches `bin_pack_arrow_table`. Nulls are already handled with 
`is_null()`. This PR selects the NaN partition with `pc.is_nan()` in the same 
way.
   
   After the fix the NaN rows are written to their own partition, and `value is 
nan` returns them.
   
   ## Are these changes tested?
   
   Yes.
   
   - `tests/io/test_pyarrow.py`: `test_determine_partitions_identity_nan` 
checks that the rows for `1.0`, `NaN` and `None` land in their partitions. On 
`main` it returns `{'nan': []}`.
   - `tests/catalog/test_catalog_behaviors.py`: 
`test_append_nan_to_identity_partitioned_table` appends to a real table (memory 
and SQL catalogs), then checks that a full scan returns all 4 rows and `value 
is nan` returns the two NaN rows. On `main` the append raises 
`ZeroDivisionError`.
   
   All 4 new tests fail on `main` and pass with the fix. 
`tests/io/test_pyarrow.py` and `tests/catalog/test_catalog_behaviors.py` pass. 
`prek run --files` (ruff, ruff-format, mypy, pydocstyle, codespell) passes. I 
also checked `float` columns by hand.
   
   ## Are there any user-facing changes?
   
   Yes. `append()`/`overwrite()` now write NaN values into identity-partitioned 
floating-point columns instead of failing.
   
   AI disclosure: this bug was found, fixed and tested by an AI coding agent 
(Claude) running under the breken-ai account; the red/green runs above are its 
local results.
   
   <!-- In the case of user-facing changes, please add the changelog label. -->
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to