[
https://issues.apache.org/jira/browse/HDFS-17975?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HDFS-17975:
----------------------------------
Labels: pull-request-available (was: )
> HDFS Client-Side Block Prefetch to Improve Large Sequential Read Throughput
> ---------------------------------------------------------------------------
>
> Key: HDFS-17975
> URL: https://issues.apache.org/jira/browse/HDFS-17975
> Project: Hadoop HDFS
> Issue Type: Improvement
> Components: hdfs-client
> Reporter: Rajan Dhabalia
> Priority: Major
> Labels: pull-request-available
>
> h2. Summary
> Add client-side block prefetching (parallel read-ahead) to `DFSInputStream`
> to improve throughput for large and sequential read workloads.
> h2. Issue Type
> Improvement
> h2. Component/s
> hdfs-client
> h2. Description
> h3. Motivation
> Large scan and sequential read workloads on `DFSInputStream` can be limited
> by synchronous, one-block-at-a-time reads. The latency of each remote block
> read can stall the consumer and prevent full utilization of available network
> and DataNode capacity.
> Client-side prefetching hides this latency by fetching upcoming blocks in
> parallel while the consumer processes previously read data.
> h3. Approach
> Introduce a config-gated, default-off `BlockPrefetcher` in `DFSInputStream`.
> When enabled:
> * Asynchronously prefetch upcoming data using a bounded thread pool.
> * Store prefetched data in a bounded in-memory cache.
> * Serve reads directly from the cache when available.
> * Fall back to the existing synchronous read path on cache misses.
> * Support configurable prefetch window, chunk size, cache size, worker
> threads, and TTL.
> * Provide optional periodic metrics logging for cache hit ratio and
> prefetched bytes.
> * Add `PrefetchReadExample` and unit tests.
> The existing read path remains unchanged when prefetching is disabled.
> h3. Configuration
> ||Property||Default||Description||
> |`dfs.client.prefetch.enabled`|false|Master switch for client-side
> prefetching|
> |`dfs.client.prefetch.size`|See docs|Prefetch window size|
> |`dfs.client.prefetch.max.bytes`|See docs|Maximum bytes held in the prefetch
> cache|
> |`dfs.client.prefetch.chunk.size`|See docs|Size of each prefetched chunk|
> |`dfs.client.prefetch.threads`|See docs|Number of prefetch worker threads|
> |`dfs.client.prefetch.threadpool.size`|See docs|Prefetch thread-pool size|
> |`dfs.client.prefetch.ttl.ms`|See docs|TTL for cached prefetched data|
> |`dfs.client.prefetch.metrics.log.enabled`|false|Enable periodic prefetch
> metrics logging|
> |`dfs.client.prefetch.metrics.log.interval.ms`|See docs|Metrics logging
> interval|
> h3. Results
> Testing on large sequential read workloads showed:
> * {*}Up to 3.49x higher average read throughput{*}, approximately {*}249%
> improvement{*}.
> * Peak throughput of approximately {*}1.58 GB/s{*}.
> * Approximately {*}81% prefetch cache hit ratio{*}.
> * Approximately *4 out of 5 reads* served from the prefetch cache.
> * No change to the existing read path when the feature is disabled.
> h3. Estimated Improvement
> The benefit is workload-dependent and increases with read sequentiality and
> the amount of latency that can be hidden through parallel prefetching.
> ||Workload||Expected Benefit||Rationale||
> |Large sequential scan|~3x to 3.5x|Upcoming blocks can be fetched in parallel
> while earlier data is processed|
> |Mixed sequential + occasional seek|~1.8x to 2.5x|Sequential portions
> benefit, while seeks reduce cache effectiveness|
> |Random / small reads|~1x|Limited sequentiality provides little opportunity
> for prefetching|
> h3. Resource Protection
> Prefetching is bounded to prevent uncontrolled memory or CPU consumption:
> * Maximum prefetch cache size
> * Configurable prefetch window and chunk size
> * Bounded worker threads and thread-pool size
> * Entry TTL
> * Optional metrics logging
> For workloads with limited sequentiality, these limits reduce unnecessary
> memory usage and background work.
> h3. Expected Impact
> * Improve throughput for large sequential and scan workloads.
> * Hide remote block-read latency through parallelism.
> * Improve utilization of network and DataNode capacity.
> * Reduce the impact of per-block RPC latency on a single reader.
> * Preserve existing behavior when disabled.
> * Provide tunable resource limits for different workloads.
> h2. Backward Compatibility
> The feature is {*}disabled by default{*}. Existing `DFSInputStream` behavior
> and read paths remain unchanged unless client-side prefetching is explicitly
> enabled.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]