srujankgandla commented on issue #2407: URL: https://github.com/apache/iceberg-python/issues/2407#issuecomment-5891717436
Hi, I'd like to take this on. I reproduced it locally: reading 1.2M rows (5.9 MB across 12 parquet files) through to_arrow_batch_reader() peaks at ~74 MB RSS — 12.6x the on-disk size — and instrumenting _task_to_record_batches shows every file is fully read and materialized even when the consumer only takes a single batch. That matches Tommo56700's diagnosis: batches_for_task() wraps the per-task iterator in list() and executor.map submits all file tasks eagerly, so the "streaming" reader is eager in practice. My planned approach: add a lazy iteration path for the batch reader that streams each file's batches without the per-task list() materialization or the eager executor fan-out, leaving to_table()/to_pandas() on the existing threaded path untouched. I'll put up a PR once I have it working with a regression test. Let me know if you have concerns with that direction. Thanks! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
