nhobin219 opened a new issue, #4084:
URL: https://github.com/apache/iceberg-python/issues/4084

   ### Apache Iceberg version
   
   0.12.0 (latest release)
   
   ### Please describe the bug 🐞
   
   `DataScan.to_arrow_batch_reader()` aborts the whole Python process when the 
scan has a `row_filter` and the table has a `map<string, struct<...>>` column. 
`to_arrow()` on the same scan works. The abort is a failed pyarrow `DCHECK`, so 
it can't be caught:
   
   ```
   /arrow/cpp/src/arrow/array/array_nested.cc:912:  Check failed: _s.ok() 
Operation failed: ValidateChildData(data->child_data)
   Bad status: Invalid: Map array keys array should have no nulls
   Fatal Python error: Aborted
   ```
   
   It needs all three of these:
   - a `row_filter`: the same scan without one is fine;
   - a map whose **values are structs**: `map<string, string>` with the same 
filter is fine;
   - `to_arrow_batch_reader()`: `to_arrow()` returns the right 2 rows.
   
   The table can be written with `append` or `add_files`; both crash.
   
   **Repro** (pyiceberg 0.12.0, pyarrow 25.0.1, Python 3.11, Linux x86_64):
   
   ```python
   import tempfile
   
   import pyarrow as pa
   from pyiceberg.catalog.sql import SqlCatalog
   from pyiceberg.expressions import GreaterThanOrEqual
   
   value = pa.struct([pa.field("s", pa.string()), pa.field("i", pa.int64())])
   schema = pa.schema([
       pa.field("id", pa.int64(), nullable=False),
       pa.field("attrs", pa.map_(pa.string(), value)),
   ])
   data = pa.table(
       {"id": [1, 2, 3], "attrs": [[("a", {"s": "x"})], [("b", {"i": 2})], 
[("c", {"s": "y"})]]},
       schema=schema,
   )
   
   warehouse = tempfile.mkdtemp()
   catalog = SqlCatalog("t", uri=f"sqlite:///{warehouse}/c.db", 
warehouse=f"file://{warehouse}")
   catalog.create_namespace("ns")
   table = catalog.create_table("ns.t", schema=schema)
   table.append(data)
   
   scan = table.scan(row_filter=GreaterThanOrEqual("id", 2))
   print("to_arrow:", scan.to_arrow().num_rows, flush=True)  # 2
   for batch in scan.to_arrow_batch_reader():  # aborts the process
       print("batch:", batch.num_rows, flush=True)
   ```
   
   Output:
   
   ```
   to_arrow: 2
   /arrow/cpp/src/arrow/array/array_nested.cc:912:  Check failed: _s.ok() 
Operation failed: ValidateChildData(data->child_data)
   Bad status: Invalid: Map array keys array should have no nulls
   Fatal Python error: Aborted
   ```
   
   The keys are never null in the data (Iceberg map keys are required). My 
guess is that the batch path filters or projects the map's struct values in a 
way that leaves the keys child misaligned or marked nullable, where 
`to_arrow()` concatenates first. I haven't confirmed that.
   
   Found in [litelink](https://github.com/nhobin219/litelink), where compaction 
streams a file's rows through `to_arrow_batch_reader()`. We work around it by 
using `to_arrow()` per data file.
   
   ### Willingness to contribute
   
   - [ ] I can contribute a fix for this bug independently
   - [ ] I would be willing to contribute a fix for this bug with guidance from 
the Iceberg community
   - [x] I cannot contribute a fix for this bug at this time
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to