srujankgandla commented on issue #2407:
URL: 
https://github.com/apache/iceberg-python/issues/2407#issuecomment-5891717436

   Hi, I'd like to take this on. I reproduced it locally: reading 1.2M rows 
(5.9 MB across 12 parquet files) through to_arrow_batch_reader() peaks at ~74 
MB RSS — 12.6x the on-disk size — and instrumenting _task_to_record_batches 
shows every file is fully read and materialized even when the consumer only 
takes a single batch. That matches Tommo56700's diagnosis: batches_for_task() 
wraps the per-task iterator in list() and executor.map submits all file tasks 
eagerly, so the "streaming" reader is eager in practice.
   My planned approach: add a lazy iteration path for the batch reader that 
streams each file's batches without the per-task list() materialization or the 
eager executor fan-out, leaving to_table()/to_pandas() on the existing threaded 
path untouched. I'll put up a PR once I have it working with a regression test.
   Let me know if you have concerns with that direction. Thanks!


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to