nhobin219 opened a new issue, #4084:
URL: https://github.com/apache/iceberg-python/issues/4084
### Apache Iceberg version
0.12.0 (latest release)
### Please describe the bug 🐞
`DataScan.to_arrow_batch_reader()` aborts the whole Python process when the
scan has a `row_filter` and the table has a `map<string, struct<...>>` column.
`to_arrow()` on the same scan works. The abort is a failed pyarrow `DCHECK`, so
it can't be caught:
```
/arrow/cpp/src/arrow/array/array_nested.cc:912: Check failed: _s.ok()
Operation failed: ValidateChildData(data->child_data)
Bad status: Invalid: Map array keys array should have no nulls
Fatal Python error: Aborted
```
It needs all three of these:
- a `row_filter`: the same scan without one is fine;
- a map whose **values are structs**: `map<string, string>` with the same
filter is fine;
- `to_arrow_batch_reader()`: `to_arrow()` returns the right 2 rows.
The table can be written with `append` or `add_files`; both crash.
**Repro** (pyiceberg 0.12.0, pyarrow 25.0.1, Python 3.11, Linux x86_64):
```python
import tempfile
import pyarrow as pa
from pyiceberg.catalog.sql import SqlCatalog
from pyiceberg.expressions import GreaterThanOrEqual
value = pa.struct([pa.field("s", pa.string()), pa.field("i", pa.int64())])
schema = pa.schema([
pa.field("id", pa.int64(), nullable=False),
pa.field("attrs", pa.map_(pa.string(), value)),
])
data = pa.table(
{"id": [1, 2, 3], "attrs": [[("a", {"s": "x"})], [("b", {"i": 2})],
[("c", {"s": "y"})]]},
schema=schema,
)
warehouse = tempfile.mkdtemp()
catalog = SqlCatalog("t", uri=f"sqlite:///{warehouse}/c.db",
warehouse=f"file://{warehouse}")
catalog.create_namespace("ns")
table = catalog.create_table("ns.t", schema=schema)
table.append(data)
scan = table.scan(row_filter=GreaterThanOrEqual("id", 2))
print("to_arrow:", scan.to_arrow().num_rows, flush=True) # 2
for batch in scan.to_arrow_batch_reader(): # aborts the process
print("batch:", batch.num_rows, flush=True)
```
Output:
```
to_arrow: 2
/arrow/cpp/src/arrow/array/array_nested.cc:912: Check failed: _s.ok()
Operation failed: ValidateChildData(data->child_data)
Bad status: Invalid: Map array keys array should have no nulls
Fatal Python error: Aborted
```
The keys are never null in the data (Iceberg map keys are required). My
guess is that the batch path filters or projects the map's struct values in a
way that leaves the keys child misaligned or marked nullable, where
`to_arrow()` concatenates first. I haven't confirmed that.
Found in [litelink](https://github.com/nhobin219/litelink), where compaction
streams a file's rows through `to_arrow_batch_reader()`. We work around it by
using `to_arrow()` per data file.
### Willingness to contribute
- [ ] I can contribute a fix for this bug independently
- [ ] I would be willing to contribute a fix for this bug with guidance from
the Iceberg community
- [x] I cannot contribute a fix for this bug at this time
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]