[ 
https://issues.apache.org/jira/browse/HDFS-17975?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112083#comment-18112083
 ] 

ASF GitHub Bot commented on HDFS-17975:
---------------------------------------

hadoop-yetus commented on PR #8716:
URL: https://github.com/apache/hadoop/pull/8716#issuecomment-5560065914

   :broken_heart: **-1 overall**
   
   
   
   
   
   
   | Vote | Subsystem | Runtime |  Logfile | Comment |
   |:----:|----------:|--------:|:--------:|:-------:|
   | +0 :ok: |  reexec  |  19m 34s |  |  Docker mode activated.  |
   |||| _ Prechecks _ |
   | +1 :green_heart: |  dupname  |   0m  0s |  |  No case conflicting files 
found.  |
   | +0 :ok: |  codespell  |   0m  0s |  |  codespell was not available.  |
   | +0 :ok: |  detsecrets  |   0m  0s |  |  detect-secrets was not available.  
|
   | +0 :ok: |  xmllint  |   0m  0s |  |  xmllint was not available.  |
   | +1 :green_heart: |  @author  |   0m  0s |  |  The patch does not contain 
any @author tags.  |
   | +1 :green_heart: |  test4tests  |   0m  0s |  |  The patch appears to 
include 2 new or modified test files.  |
   |||| _ trunk Compile Tests _ |
   | +0 :ok: |  mvndep  |   1m 52s |  |  Maven dependency ordering for branch  |
   | +1 :green_heart: |  mvninstall  |  42m 44s |  |  trunk passed  |
   | +1 :green_heart: |  compile  |   5m 26s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  compile  |   5m 49s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  checkstyle  |   2m 11s |  |  trunk passed  |
   | +1 :green_heart: |  mvnsite  |   3m 23s |  |  trunk passed  |
   | +1 :green_heart: |  javadoc  |   2m 42s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   2m 40s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   7m 48s |  |  trunk passed  |
   | +1 :green_heart: |  shadedclient  |  31m 41s |  |  branch has no errors 
when building and testing our client artifacts.  |
   |||| _ Patch Compile Tests _ |
   | +0 :ok: |  mvndep  |   0m 28s |  |  Maven dependency ordering for patch  |
   | +1 :green_heart: |  mvninstall  |   2m 13s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   4m 53s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   4m 53s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   5m 23s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   5m 23s |  |  the patch passed  |
   | +1 :green_heart: |  blanks  |   0m  0s |  |  The patch has no blanks 
issues.  |
   | -0 :warning: |  checkstyle  |   1m 38s | 
[/results-checkstyle-hadoop-hdfs-project.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/4/artifact/out/results-checkstyle-hadoop-hdfs-project.txt)
 |  hadoop-hdfs-project: The patch generated 32 new + 109 unchanged - 0 fixed = 
141 total (was 109)  |
   | +1 :green_heart: |  mvnsite  |   2m 23s |  |  the patch passed  |
   | +1 :green_heart: |  javadoc  |   1m 37s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   1m 44s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   7m 17s |  |  the patch passed  |
   | +1 :green_heart: |  shadedclient  |  30m 51s |  |  patch has no errors 
when building and testing our client artifacts.  |
   |||| _ Other Tests _ |
   | +1 :green_heart: |  unit  |   2m 42s |  |  hadoop-hdfs-client in the patch 
passed.  |
   | -1 :x: |  unit  | 285m 47s | 
[/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/4/artifact/out/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch passed.  |
   | +1 :green_heart: |  asflicense  |   0m 48s |  |  The patch does not 
generate ASF License warnings.  |
   |  |   | 472m 36s |  |  |
   
   
   | Reason | Tests |
   |-------:|:------|
   | Failed junit tests | hadoop.tools.TestHdfsConfigFields |
   
   
   | Subsystem | Report/Notes |
   |----------:|:-------------|
   | Docker | ClientAPI=1.56 ServerAPI=1.56 base: 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/4/artifact/out/Dockerfile
 |
   | GITHUB PR | https://github.com/apache/hadoop/pull/8716 |
   | Optional Tests | dupname asflicense compile javac javadoc mvninstall 
mvnsite unit shadedclient spotbugs checkstyle codespell detsecrets xmllint |
   | uname | Linux 695bbc6afa35 5.15.0-190-generic #200-Ubuntu SMP Fri Aug 7 
15:06:04 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux |
   | Build tool | maven |
   | Personality | dev-support/bin/hadoop.sh |
   | git revision | trunk / 6081187d44c9d13c0690bd975716c99272b4dbd3 |
   | Default Java | Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   | Multi-JDK versions | 
/usr/lib/jvm/java-21-openjdk-amd64:Ubuntu-21.0.12+8-1-24.04-Ubuntu 
/usr/lib/jvm/java-17-openjdk-amd64:Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   |  Test Results | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/4/testReport/ |
   | Max. process+thread count | 3019 (vs. ulimit of 10000) |
   | modules | C: hadoop-hdfs-project/hadoop-hdfs-client 
hadoop-hdfs-project/hadoop-hdfs U: hadoop-hdfs-project |
   | Console output | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8716/4/console |
   | versions | git=2.43.0 maven=3.9.15 spotbugs=4.9.7 |
   | Powered by | Apache Yetus 0.14.1 https://yetus.apache.org |
   
   
   This message was automatically generated.
   
   




> HDFS Client-Side Block Prefetch to Improve Large Sequential Read Throughput
> ---------------------------------------------------------------------------
>
>                 Key: HDFS-17975
>                 URL: https://issues.apache.org/jira/browse/HDFS-17975
>             Project: Hadoop HDFS
>          Issue Type: Improvement
>          Components: hdfs-client
>            Reporter: Rajan Dhabalia
>            Priority: Major
>              Labels: pull-request-available
>
> h2. Summary
> Add client-side block prefetching (parallel read-ahead) to `DFSInputStream` 
> to improve throughput for large and sequential read workloads.
> h2. Issue Type
> Improvement
> h2. Component/s
> hdfs-client
> h2. Description
> h3. Motivation
> Large scan and sequential read workloads on `DFSInputStream` can be limited 
> by synchronous, one-block-at-a-time reads. The latency of each remote block 
> read can stall the consumer and prevent full utilization of available network 
> and DataNode capacity.
> Client-side prefetching hides this latency by fetching upcoming blocks in 
> parallel while the consumer processes previously read data.
> h3. Approach
> Introduce a config-gated, default-off `BlockPrefetcher` in `DFSInputStream`.
> When enabled:
>  * Asynchronously prefetch upcoming data using a bounded thread pool.
>  * Store prefetched data in a bounded in-memory cache.
>  * Serve reads directly from the cache when available.
>  * Fall back to the existing synchronous read path on cache misses.
>  * Support configurable prefetch window, chunk size, cache size, worker 
> threads, and TTL.
>  * Provide optional periodic metrics logging for cache hit ratio and 
> prefetched bytes.
>  * Add `PrefetchReadExample` and unit tests.
> The existing read path remains unchanged when prefetching is disabled.
> h3. Configuration
> ||Property||Default||Description||
> |`dfs.client.prefetch.enabled`|false|Master switch for client-side 
> prefetching|
> |`dfs.client.prefetch.size`|See docs|Prefetch window size|
> |`dfs.client.prefetch.max.bytes`|See docs|Maximum bytes held in the prefetch 
> cache|
> |`dfs.client.prefetch.chunk.size`|See docs|Size of each prefetched chunk|
> |`dfs.client.prefetch.threads`|See docs|Number of prefetch worker threads|
> |`dfs.client.prefetch.threadpool.size`|See docs|Prefetch thread-pool size|
> |`dfs.client.prefetch.ttl.ms`|See docs|TTL for cached prefetched data|
> |`dfs.client.prefetch.metrics.log.enabled`|false|Enable periodic prefetch 
> metrics logging|
> |`dfs.client.prefetch.metrics.log.interval.ms`|See docs|Metrics logging 
> interval|
> h3. Results
> Testing on large sequential read workloads showed:
>  * {*}Up to 3.49x higher average read throughput{*}, approximately {*}249% 
> improvement{*}.
>  * Peak throughput of approximately {*}1.58 GB/s{*}.
>  * Approximately {*}81% prefetch cache hit ratio{*}.
>  * Approximately *4 out of 5 reads* served from the prefetch cache.
>  * No change to the existing read path when the feature is disabled.
> h3. Estimated Improvement
> The benefit is workload-dependent and increases with read sequentiality and 
> the amount of latency that can be hidden through parallel prefetching.
> ||Workload||Expected Benefit||Rationale||
> |Large sequential scan|~3x to 3.5x|Upcoming blocks can be fetched in parallel 
> while earlier data is processed|
> |Mixed sequential + occasional seek|~1.8x to 2.5x|Sequential portions 
> benefit, while seeks reduce cache effectiveness|
> |Random / small reads|~1x|Limited sequentiality provides little opportunity 
> for prefetching|
> h3. Resource Protection
> Prefetching is bounded to prevent uncontrolled memory or CPU consumption:
>  * Maximum prefetch cache size
>  * Configurable prefetch window and chunk size
>  * Bounded worker threads and thread-pool size
>  * Entry TTL
>  * Optional metrics logging
> For workloads with limited sequentiality, these limits reduce unnecessary 
> memory usage and background work.
> h3. Expected Impact
>  * Improve throughput for large sequential and scan workloads.
>  * Hide remote block-read latency through parallelism.
>  * Improve utilization of network and DataNode capacity.
>  * Reduce the impact of per-block RPC latency on a single reader.
>  * Preserve existing behavior when disabled.
>  * Provide tunable resource limits for different workloads.
> h2. Backward Compatibility
> The feature is {*}disabled by default{*}. Existing `DFSInputStream` behavior 
> and read paths remain unchanged unless client-side prefetching is explicitly 
> enabled.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to