[
https://issues.apache.org/jira/browse/HDFS-17975?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18111870#comment-18111870
]
ASF GitHub Bot commented on HDFS-17975:
---------------------------------------
hadoop-yetus commented on PR #8716:
URL: https://github.com/apache/hadoop/pull/8716#issuecomment-5551838526
:broken_heart: **-1 overall**
| Vote | Subsystem | Runtime | Logfile | Comment |
|:----:|----------:|--------:|:--------:|:-------:|
| +0 :ok: | reexec | 1m 10s | | Docker mode activated. |
|||| _ Prechecks _ |
| +1 :green_heart: | dupname | 0m 0s | | No case conflicting files
found. |
| +0 :ok: | codespell | 0m 0s | | codespell was not available. |
| +0 :ok: | detsecrets | 0m 0s | | detect-secrets was not available.
|
| +0 :ok: | xmllint | 0m 0s | | xmllint was not available. |
| +1 :green_heart: | @author | 0m 0s | | The patch does not contain
any @author tags. |
| +1 :green_heart: | test4tests | 0m 0s | | The patch appears to
include 2 new or modified test files. |
|||| _ trunk Compile Tests _ |
| +0 :ok: | mvndep | 2m 10s | | Maven dependency ordering for branch |
| +1 :green_heart: | mvninstall | 46m 14s | | trunk passed |
| +1 :green_heart: | compile | 5m 50s | | trunk passed with JDK
Ubuntu-21.0.12+8-1-24.04-Ubuntu |
| +1 :green_heart: | compile | 6m 13s | | trunk passed with JDK
Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| +1 :green_heart: | checkstyle | 2m 10s | | trunk passed |
| +1 :green_heart: | mvnsite | 3m 12s | | trunk passed |
| +1 :green_heart: | javadoc | 2m 34s | | trunk passed with JDK
Ubuntu-21.0.12+8-1-24.04-Ubuntu |
| +1 :green_heart: | javadoc | 2m 32s | | trunk passed with JDK
Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| +1 :green_heart: | spotbugs | 7m 37s | | trunk passed |
| +1 :green_heart: | shadedclient | 33m 25s | | branch has no errors
when building and testing our client artifacts. |
|||| _ Patch Compile Tests _ |
| +0 :ok: | mvndep | 0m 29s | | Maven dependency ordering for patch |
| +1 :green_heart: | mvninstall | 2m 30s | | the patch passed |
| +1 :green_heart: | compile | 5m 50s | | the patch passed with JDK
Ubuntu-21.0.12+8-1-24.04-Ubuntu |
| +1 :green_heart: | javac | 5m 50s | | the patch passed |
| +1 :green_heart: | compile | 6m 17s | | the patch passed with JDK
Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| +1 :green_heart: | javac | 6m 17s | | the patch passed |
| +1 :green_heart: | blanks | 0m 0s | | The patch has no blanks
issues. |
| -0 :warning: | checkstyle | 1m 48s |
[/results-checkstyle-hadoop-hdfs-project.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/3/artifact/out/results-checkstyle-hadoop-hdfs-project.txt)
| hadoop-hdfs-project: The patch generated 40 new + 109 unchanged - 0 fixed =
149 total (was 109) |
| +1 :green_heart: | mvnsite | 2m 46s | | the patch passed |
| +1 :green_heart: | javadoc | 1m 44s | | the patch passed with JDK
Ubuntu-21.0.12+8-1-24.04-Ubuntu |
| +1 :green_heart: | javadoc | 1m 49s | | the patch passed with JDK
Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| +1 :green_heart: | spotbugs | 8m 15s | | the patch passed |
| +1 :green_heart: | shadedclient | 37m 13s | | patch has no errors
when building and testing our client artifacts. |
|||| _ Other Tests _ |
| +1 :green_heart: | unit | 2m 47s | | hadoop-hdfs-client in the patch
passed. |
| -1 :x: | unit | 300m 40s |
[/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/3/artifact/out/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt)
| hadoop-hdfs in the patch passed. |
| +1 :green_heart: | asflicense | 0m 47s | | The patch does not
generate ASF License warnings. |
| | | 484m 55s | | |
| Reason | Tests |
|-------:|:------|
| Failed junit tests | hadoop.hdfs.server.namenode.ha.TestBootstrapStandby |
| | hadoop.tools.TestHdfsConfigFields |
| | hadoop.hdfs.server.namenode.ha.TestStandbyCheckpoints |
| Subsystem | Report/Notes |
|----------:|:-------------|
| Docker | ClientAPI=1.56 ServerAPI=1.56 base:
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/3/artifact/out/Dockerfile
|
| GITHUB PR | https://github.com/apache/hadoop/pull/8716 |
| Optional Tests | dupname asflicense compile javac javadoc mvninstall
mvnsite unit shadedclient spotbugs checkstyle codespell detsecrets xmllint |
| uname | Linux 3384f607c02f 5.15.0-190-generic #200-Ubuntu SMP Fri Aug 7
15:06:04 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux |
| Build tool | maven |
| Personality | dev-support/bin/hadoop.sh |
| git revision | trunk / 9b42d55e34eea1da175f4780dbe2b870c19c768d |
| Default Java | Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| Multi-JDK versions |
/usr/lib/jvm/java-21-openjdk-amd64:Ubuntu-21.0.12+8-1-24.04-Ubuntu
/usr/lib/jvm/java-17-openjdk-amd64:Ubuntu-17.0.20+8-1-24.04-Ubuntu |
| Test Results |
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/3/testReport/ |
| Max. process+thread count | 3119 (vs. ulimit of 10000) |
| modules | C: hadoop-hdfs-project/hadoop-hdfs-client
hadoop-hdfs-project/hadoop-hdfs U: hadoop-hdfs-project |
| Console output |
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/3/console |
| versions | git=2.43.0 maven=3.9.15 spotbugs=4.9.7 |
| Powered by | Apache Yetus 0.14.1 https://yetus.apache.org |
This message was automatically generated.
> HDFS Client-Side Block Prefetch to Improve Large Sequential Read Throughput
> ---------------------------------------------------------------------------
>
> Key: HDFS-17975
> URL: https://issues.apache.org/jira/browse/HDFS-17975
> Project: Hadoop HDFS
> Issue Type: Improvement
> Components: hdfs-client
> Reporter: Rajan Dhabalia
> Priority: Major
> Labels: pull-request-available
>
> h2. Summary
> Add client-side block prefetching (parallel read-ahead) to `DFSInputStream`
> to improve throughput for large and sequential read workloads.
> h2. Issue Type
> Improvement
> h2. Component/s
> hdfs-client
> h2. Description
> h3. Motivation
> Large scan and sequential read workloads on `DFSInputStream` can be limited
> by synchronous, one-block-at-a-time reads. The latency of each remote block
> read can stall the consumer and prevent full utilization of available network
> and DataNode capacity.
> Client-side prefetching hides this latency by fetching upcoming blocks in
> parallel while the consumer processes previously read data.
> h3. Approach
> Introduce a config-gated, default-off `BlockPrefetcher` in `DFSInputStream`.
> When enabled:
> * Asynchronously prefetch upcoming data using a bounded thread pool.
> * Store prefetched data in a bounded in-memory cache.
> * Serve reads directly from the cache when available.
> * Fall back to the existing synchronous read path on cache misses.
> * Support configurable prefetch window, chunk size, cache size, worker
> threads, and TTL.
> * Provide optional periodic metrics logging for cache hit ratio and
> prefetched bytes.
> * Add `PrefetchReadExample` and unit tests.
> The existing read path remains unchanged when prefetching is disabled.
> h3. Configuration
> ||Property||Default||Description||
> |`dfs.client.prefetch.enabled`|false|Master switch for client-side
> prefetching|
> |`dfs.client.prefetch.size`|See docs|Prefetch window size|
> |`dfs.client.prefetch.max.bytes`|See docs|Maximum bytes held in the prefetch
> cache|
> |`dfs.client.prefetch.chunk.size`|See docs|Size of each prefetched chunk|
> |`dfs.client.prefetch.threads`|See docs|Number of prefetch worker threads|
> |`dfs.client.prefetch.threadpool.size`|See docs|Prefetch thread-pool size|
> |`dfs.client.prefetch.ttl.ms`|See docs|TTL for cached prefetched data|
> |`dfs.client.prefetch.metrics.log.enabled`|false|Enable periodic prefetch
> metrics logging|
> |`dfs.client.prefetch.metrics.log.interval.ms`|See docs|Metrics logging
> interval|
> h3. Results
> Testing on large sequential read workloads showed:
> * {*}Up to 3.49x higher average read throughput{*}, approximately {*}249%
> improvement{*}.
> * Peak throughput of approximately {*}1.58 GB/s{*}.
> * Approximately {*}81% prefetch cache hit ratio{*}.
> * Approximately *4 out of 5 reads* served from the prefetch cache.
> * No change to the existing read path when the feature is disabled.
> h3. Estimated Improvement
> The benefit is workload-dependent and increases with read sequentiality and
> the amount of latency that can be hidden through parallel prefetching.
> ||Workload||Expected Benefit||Rationale||
> |Large sequential scan|~3x to 3.5x|Upcoming blocks can be fetched in parallel
> while earlier data is processed|
> |Mixed sequential + occasional seek|~1.8x to 2.5x|Sequential portions
> benefit, while seeks reduce cache effectiveness|
> |Random / small reads|~1x|Limited sequentiality provides little opportunity
> for prefetching|
> h3. Resource Protection
> Prefetching is bounded to prevent uncontrolled memory or CPU consumption:
> * Maximum prefetch cache size
> * Configurable prefetch window and chunk size
> * Bounded worker threads and thread-pool size
> * Entry TTL
> * Optional metrics logging
> For workloads with limited sequentiality, these limits reduce unnecessary
> memory usage and background work.
> h3. Expected Impact
> * Improve throughput for large sequential and scan workloads.
> * Hide remote block-read latency through parallelism.
> * Improve utilization of network and DataNode capacity.
> * Reduce the impact of per-block RPC latency on a single reader.
> * Preserve existing behavior when disabled.
> * Provide tunable resource limits for different workloads.
> h2. Backward Compatibility
> The feature is {*}disabled by default{*}. Existing `DFSInputStream` behavior
> and read paths remain unchanged unless client-side prefetching is explicitly
> enabled.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]