dependabot[bot] opened a new pull request, #4053: URL: https://github.com/apache/iceberg-python/pull/4053
Bumps [gcsfs](https://github.com/fsspec/gcsfs) from 2026.6.0 to 2026.8.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/fsspec/gcsfs/releases">gcsfs's releases</a>.</em></p> <blockquote> <h2>2026.8.1</h2> <h2>What's Changed</h2> <p><strong>Bug Fixes</strong> Fixes <a href="https://redirect.github.com/fsspec/gcsfs/issues/1048">fsspec/gcsfs#1048</a>: In release 2026.8.0, cat_file began using a default concurrency of 4 instead of 1. This caused a performance regression for small cat_file operations with unknown sizes, as an unnecessary _info() request was being triggered to determine the size for concurrent reads.</p> <p><strong>Full Changelog</strong>: <a href="https://github.com/fsspec/gcsfs/compare/2026.8.0...2026.8.1">https://github.com/fsspec/gcsfs/compare/2026.8.0...2026.8.1</a></p> <h2>2026.8.0</h2> <h2>What's Changed</h2> <p><strong>Adaptive Concurrent Prefetching is now the default read path</strong></p> <p><strong>Enhanced read path via adaptive concurrent prefetching is now the default in GCSFS.</strong> Starting with this version, GCSFS predicts the next byte range an application will read, fetches it in the background across several concurrent HTTP requests, and keeps the bytes in memory before next read() is called. Network round-trips overlap with application compute instead of blocking calls where compute has to wait for data to be fetched. We are also enabling read concurrency, so a single reader is no longer limited by the bandwidth of a single HTTP connection.</p> <p>GCSFS prefetcher adapts to workload read IO patterns. It tracks the rolling average of recent read sizes and scales the prefetch window linearly with the detected sequential streak, rather than using a fixed block size or exponential doubling. This is inline to what modern Linux kernels will do to balance prefetch and memory footprint. When the pattern turns random read i.e. we are not able to leverage the prefetched buffer to answer the next read() call, it drains the buffer to zero, so that random-access workloads pay no bandwidth or memory penalty.</p> <p><strong>Why this matters for AI/ML workloads</strong></p> <ul> <li><strong>Model loading and checkpoint restore are typically large sequential reads - and prefetcher shines there.</strong> In our benchmarking, a single-stream sequential throughput improved from 23.69 MB/s to 658.71 MB/s for 1 MB I/O, and from 156 MB/s to 736 MB/s for 16 MB I/O.</li> <li><strong>Training data pipelines stay fed.</strong> Parquet and sharded dataset reads issue small-to-medium sequential ranges that previously suffered from low throughput, but now achieve significantly more; at 16 MB I/O throughput rises from 150 MB/s to 730 MB/s. Reducing the wait time for data loading improves accelerator goodput(amount of time accelerator is utilised for training than waiting).</li> <li><strong>Multi-worker dataloader scaling.</strong> The prefetcher manufactures its own parallelism per worker instead of relying on process count alone.</li> <li><strong>Accelerate the throughput even further with Rapid Buckets.</strong> With Rapid Buckets single node throughput reaches 21 GiB/s with 16-process sequentially reading at 16 MiB I/O compared to standard buckets with 48processes.</li> </ul> <p><strong>Adaptive prefetcher is enabled by default</strong> when cache_type is not explicitly set and concurrency value is set at 4(DEFAULT_GCSFS_CONCURRENCY=4) for both Standard and Rapid buckets. You can disable adaptive prefetcher by setting an explicit cache_type, or by setting USE_EXPERIMENTAL_ADAPTIVE_PREFETCHING='false', or by passing use_experimental_adaptive_prefetching=False to open() call.</p> <p><strong>(Warning) Impact on memory:</strong> Prefetching trades memory for throughput. Peak memory rises from ~170 MB to 600 MB on single-stream reads for 16 MB IO size and varies with requested IO sizes, and would be materially more under high process counts. Please ensure that application memory limits accordingly to use prefetcher without any Out of Memory(OOM) issues. To put hard limit, you can also use <a href="https://github.com/fsspec/gcsfs/blob/main/gcsfs/prefetcher.py#L154">user_max_prefetch_size</a></p> <p>For details on architecture, tuning, full benchmark tables, along with known limitations please refer to : <a href="https://github.com/fsspec/gcsfs/blob/main/docs/source/prefetcher.rst">https://github.com/fsspec/gcsfs/blob/main/docs/source/prefetcher.rst</a></p> <p>(<a href="https://redirect.github.com/fsspec/gcsfs/issues/795">#795</a>, <a href="https://redirect.github.com/fsspec/gcsfs/issues/805">#805</a>, <a href="https://redirect.github.com/fsspec/gcsfs/issues/818">#818</a>, <a href="https://redirect.github.com/fsspec/gcsfs/issues/877">#877</a>)</p> <p><strong>Bug Fixes & Improvements</strong></p> <ul> <li>Zero-cost local backward seeks in PrefetchConsumer - Parquet footer and ZIP directory reads are served from the existing buffer instead of re-issuing a network request. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/930">#930</a>)</li> <li>Concurrent downloads cap task count against a minimum chunk size, removing per-task overhead on small ranges. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/926">#926</a>)</li> <li>Generation consistency across parallel fetches, guaranteeing every chunk comes from the same object version. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/921">#921</a>)</li> <li>Fixed silent truncation on short reads in zonal bucket downloads. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/920">#920</a>)</li> <li>Zero-copy read and write paths via memoryview, cutting CPU and transient memory in the hot path. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/840">#840</a>, <a href="https://redirect.github.com/fsspec/gcsfs/issues/907">#907</a>, <a href="https://redirect.github.com/fsspec/gcsfs/issues/928">#928</a>)</li> <li>Graceful fallback where ctypes.pythonapi is unavailable. (<a href="https://redirect.github.com/fsspec/gcsfs/issues/938">#938</a>)</li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/ccbcqbz"><code>@ccbcqbz</code></a> made their first contribution in <a href="https://redirect.github.com/fsspec/gcsfs/pull/960">fsspec/gcsfs#960</a></li> <li><a href="https://github.com/raj-prince"><code>@raj-prince</code></a> made their first contribution in <a href="https://redirect.github.com/fsspec/gcsfs/pull/989">fsspec/gcsfs#989</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/fsspec/gcsfs/compare/2026.7.0...2026.8.0">https://github.com/fsspec/gcsfs/compare/2026.7.0...2026.8.0</a></p> <h2>2026.7.0</h2> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/fsspec/gcsfs/commit/73929a1b18ae3ef7b56b586a5403240cd827e0cf"><code>73929a1</code></a> fix: reverting default cat_file concurrency back to 1 (<a href="https://redirect.github.com/fsspec/gcsfs/issues/1051">#1051</a>) (<a href="https://redirect.github.com/fsspec/gcsfs/issues/1055">#1055</a>)</li> <li><a href="https://github.com/fsspec/gcsfs/commit/dc33f2366db1722dad3095d0b6497bb6befe61fc"><code>dc33f23</code></a> Merge branch 'main' into release-2026.8.0</li> <li><a href="https://github.com/fsspec/gcsfs/commit/e88cd8b2c05424cf9dd54104aef50abfcc3b5630"><code>e88cd8b</code></a> lint fixes</li> <li><a href="https://github.com/fsspec/gcsfs/commit/d98d04b3068264d30c8e0ac5fefefe92e0820ae3"><code>d98d04b</code></a> Update changelog.rst</li> <li><a href="https://github.com/fsspec/gcsfs/commit/d5c45dda9db940aa88675e77ac0d3de2429eedef"><code>d5c45dd</code></a> Update changelog.rst</li> <li><a href="https://github.com/fsspec/gcsfs/commit/158badb54be82990bdd786be77777e80c5bf2b4b"><code>158badb</code></a> feat(subsystembenchmarks): add webdataset image dataloading read benchmark (#...</li> <li><a href="https://github.com/fsspec/gcsfs/commit/2089c10349c6e044626859b94d5a4bb10f96ca19"><code>2089c10</code></a> feat(subsystembenchmarks): generalize dataloading seam for multiple loaders (...</li> <li><a href="https://github.com/fsspec/gcsfs/commit/158c361c0468cd6a18da8f900824c413c31170eb"><code>158c361</code></a> Update pyproject.toml</li> <li><a href="https://github.com/fsspec/gcsfs/commit/c99137ddff39cd1694e1214dd3ffd9c44c73a2d1"><code>c99137d</code></a> subsystembenchmarks: support model id substitution in checkpointing save (<a href="https://redirect.github.com/fsspec/gcsfs/issues/1003">#1003</a>)</li> <li><a href="https://github.com/fsspec/gcsfs/commit/9d3887527bff8208202bf3cfb8aaadd196442452"><code>9d38875</code></a> Track whether cache_type is explicitly set or defaulting in User-Agent header...</li> <li>Additional commits viewable in <a href="https://github.com/fsspec/gcsfs/compare/2026.6.0...2026.8.1">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
