mneha05 opened a new issue, #2191:
URL: https://github.com/apache/libcloud/issues/2191

   ## Summary
   
   Local storage range downloads allocate memory proportional to the entire 
source file, even when only a small range is requested, and ignore 
`chunk_size`. The file-download path also buffers the complete requested range 
before writing it.
   
   ## Detailed Information
   
   Reproduced on trunk at `8d659f527f28c76db41f19c86d056a2338717b8c`, Python 
3.12.14 on Linux.
   
   ```python
   import os
   import tempfile
   
   from libcloud.storage.drivers.local import LocalStorageDriver
   
   with tempfile.TemporaryDirectory() as root:
       driver = LocalStorageDriver(root)
       container = driver.create_container("example")
       with open(os.path.join(root, "example", "object"), "wb") as fp:
           fp.write(b"0123456789" * 2000)
       obj = container.get_object("object")
       chunks = list(driver.download_object_range_as_stream(
           obj, start_bytes=5, end_bytes=12, chunk_size=3
       ))
       print(chunks)
   ```
   
   Actual: `[b'5678901']`, following an unbounded read of the entire 
20,000-byte source file.
   Expected: `[b'567', b'890', b'1']`, reading only the seven requested bytes.
   
   `download_object_range_as_stream` calls `obj_file.read()` without a size to 
determine file length, then reads the entire requested range into one chunk. 
`download_object_range` calls `exhaust_iterator` before writing anything.
   
   A local `tracemalloc` experiment requesting 1 KiB from a 64 MiB sparse file 
with `chunk_size=128` measured 67,113,884 bytes of peak traced allocation 
before the fix and 6,242 bytes afterward. This measures Python allocations for 
that experiment, not total process memory.
   
   A focused fix and regression tests are prepared: use descriptor metadata for 
file size, read only the selected range in bounded chunks, and write chunks as 
they arrive. Existing non-inclusive range boundaries are preserved.
   
   AI assistance: This report and the proposed fix were prepared with OpenAI 
Codex (GPT-6).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to