[ 
https://issues.apache.org/jira/browse/HDFS-17977?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18111780#comment-18111780
 ] 

ASF GitHub Bot commented on HDFS-17977:
---------------------------------------

hadoop-yetus commented on PR #8719:
URL: https://github.com/apache/hadoop/pull/8719#issuecomment-5548716033

   :broken_heart: **-1 overall**
   
   
   
   
   
   
   | Vote | Subsystem | Runtime |  Logfile | Comment |
   |:----:|----------:|--------:|:--------:|:-------:|
   | +0 :ok: |  reexec  |   0m 21s |  |  Docker mode activated.  |
   |||| _ Prechecks _ |
   | +1 :green_heart: |  dupname  |   0m  0s |  |  No case conflicting files 
found.  |
   | +0 :ok: |  codespell  |   0m  0s |  |  codespell was not available.  |
   | +0 :ok: |  detsecrets  |   0m  0s |  |  detect-secrets was not available.  
|
   | +0 :ok: |  markdownlint  |   0m  0s |  |  markdownlint was not available.  
|
   | +0 :ok: |  xmllint  |   0m  1s |  |  xmllint was not available.  |
   | +1 :green_heart: |  @author  |   0m  0s |  |  The patch does not contain 
any @author tags.  |
   | +1 :green_heart: |  test4tests  |   0m  0s |  |  The patch appears to 
include 3 new or modified test files.  |
   |||| _ trunk Compile Tests _ |
   | +0 :ok: |  mvndep  |   2m 25s |  |  Maven dependency ordering for branch  |
   | +1 :green_heart: |  mvninstall  |  29m 12s |  |  trunk passed  |
   | +1 :green_heart: |  compile  |   9m 36s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  compile  |   9m 44s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  checkstyle  |   3m 35s |  |  trunk passed  |
   | +1 :green_heart: |  mvnsite  |   1m 29s |  |  trunk passed  |
   | +1 :green_heart: |  javadoc  |   1m 16s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   1m 11s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +0 :ok: |  spotbugs  |   0m 25s |  |  branch/hadoop-project no spotbugs 
output file (spotbugsXml.xml)  |
   | +1 :green_heart: |  shadedclient  |  19m 14s |  |  branch has no errors 
when building and testing our client artifacts.  |
   |||| _ Patch Compile Tests _ |
   | +0 :ok: |  mvndep  |   0m 18s |  |  Maven dependency ordering for patch  |
   | -1 :x: |  mvninstall  |   0m 20s | 
[/patch-mvninstall-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-mvninstall-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch failed.  |
   | -1 :x: |  compile  |   1m 10s | 
[/patch-compile-root-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-compile-root-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt)
 |  root in the patch failed with JDK Ubuntu-21.0.12+8-1-24.04-Ubuntu.  |
   | -1 :x: |  javac  |   1m 10s | 
[/patch-compile-root-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-compile-root-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt)
 |  root in the patch failed with JDK Ubuntu-21.0.12+8-1-24.04-Ubuntu.  |
   | -1 :x: |  compile  |   1m 17s | 
[/patch-compile-root-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-compile-root-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt)
 |  root in the patch failed with JDK Ubuntu-17.0.20+8-1-24.04-Ubuntu.  |
   | -1 :x: |  javac  |   1m 17s | 
[/patch-compile-root-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-compile-root-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt)
 |  root in the patch failed with JDK Ubuntu-17.0.20+8-1-24.04-Ubuntu.  |
   | +1 :green_heart: |  blanks  |   0m  0s |  |  The patch has no blanks 
issues.  |
   | -0 :warning: |  checkstyle  |   3m 11s | 
[/results-checkstyle-root.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/results-checkstyle-root.txt)
 |  root: The patch generated 2 new + 0 unchanged - 0 fixed = 2 total (was 0)  |
   | -1 :x: |  mvnsite  |   0m 21s | 
[/patch-mvnsite-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-mvnsite-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch failed.  |
   | -1 :x: |  javadoc  |   0m 18s | 
[/patch-javadoc-hadoop-hdfs-project_hadoop-hdfs-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-javadoc-hadoop-hdfs-project_hadoop-hdfs-jdkUbuntu-21.0.12+8-1-24.04-Ubuntu.txt)
 |  hadoop-hdfs in the patch failed with JDK Ubuntu-21.0.12+8-1-24.04-Ubuntu.  |
   | -1 :x: |  javadoc  |   0m 21s | 
[/patch-javadoc-hadoop-hdfs-project_hadoop-hdfs-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-javadoc-hadoop-hdfs-project_hadoop-hdfs-jdkUbuntu-17.0.20+8-1-24.04-Ubuntu.txt)
 |  hadoop-hdfs in the patch failed with JDK Ubuntu-17.0.20+8-1-24.04-Ubuntu.  |
   | +0 :ok: |  spotbugs  |   0m 11s |  |  hadoop-project has no data from 
spotbugs  |
   | -1 :x: |  spotbugs  |   0m 22s | 
[/patch-spotbugs-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-spotbugs-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch failed.  |
   | -1 :x: |  shadedclient  |   5m  1s |  |  patch has errors when building 
and testing our client artifacts.  |
   |||| _ Other Tests _ |
   | +1 :green_heart: |  unit  |   0m 10s |  |  hadoop-project in the patch 
passed.  |
   | -1 :x: |  unit  |   0m 18s | 
[/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch failed.  |
   | +1 :green_heart: |  asflicense  |   0m 21s |  |  The patch does not 
generate ASF License warnings.  |
   |  |   |  99m  2s |  |  |
   
   
   | Subsystem | Report/Notes |
   |----------:|:-------------|
   | Docker | ClientAPI=1.56 ServerAPI=1.56 base: 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/artifact/out/Dockerfile
 |
   | GITHUB PR | https://github.com/apache/hadoop/pull/8719 |
   | Optional Tests | dupname asflicense mvnsite codespell detsecrets 
markdownlint compile javac javadoc mvninstall unit shadedclient spotbugs 
checkstyle xmllint |
   | uname | Linux 9a868832f9ce 5.15.0-190-generic #200-Ubuntu SMP Fri Aug 7 
15:06:04 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux |
   | Build tool | maven |
   | Personality | dev-support/bin/hadoop.sh |
   | git revision | trunk / 5935cb9bcc7ee2011891f373076b525d4287a166 |
   | Default Java | Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   | Multi-JDK versions | 
/usr/lib/jvm/java-21-openjdk-amd64:Ubuntu-21.0.12+8-1-24.04-Ubuntu 
/usr/lib/jvm/java-17-openjdk-amd64:Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   |  Test Results | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/testReport/ |
   | Max. process+thread count | 610 (vs. ulimit of 10000) |
   | modules | C: hadoop-project hadoop-hdfs-project/hadoop-hdfs U: . |
   | Console output | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8719/1/console |
   | versions | git=2.43.0 maven=3.9.15 spotbugs=4.9.7 |
   | Powered by | Apache Yetus 0.14.1 https://yetus.apache.org |
   
   
   This message was automatically generated.
   
   




> HDFS-Test Add a compact standalone HDFS DataNode read/write stress tester
> -------------------------------------------------------------------------
>
>                 Key: HDFS-17977
>                 URL: https://issues.apache.org/jira/browse/HDFS-17977
>             Project: Hadoop HDFS
>          Issue Type: Improvement
>          Components: benchmarks, hdfs, test
>            Reporter: Rajan Dhabalia
>            Priority: Major
>              Labels: pull-request-available
>
> h2. Summary
> Add a compact, single-process HDFS read/write stress tester that generates 
> controlled throughput against targeted DataNodes and reports client-side 
> latency distributions.
> h2. Description
> h3. Motivation
> `TestDFSIO` is the standard HDFS I/O benchmark but is not well suited for 
> targeted DataNode stress testing.
> Key limitations:
>  * Requires MapReduce/YARN.
>  * Load is distributed through MapReduce scheduling rather than directly 
> targeting DataNodes.
>  * Limited control over steady throughput/QPS against selected DataNodes.
>  * Primarily reports aggregate throughput rather than client-side p50/p95/p99 
> latency.
>  * Does not provide a reliable mechanism for cold-read workloads targeting 
> disk I/O.
> `HdfsStressTest` provides a lightweight standalone tool for controlled 
> DataNode performance and stress testing.
> h3. Approach
> Introduce `HdfsStressTest`, a single-process load generator that depends only 
> on the HDFS client.
> Key capabilities:
>  # *Controlled throughput*
>  ## Global token-bucket rate limiter.
>  ## Configurable read/write throughput.
>  ## Optional linear throughput ramp.
>  # *Targeted DataNodes*
>  ## Use HDFS favored-nodes hints to direct writes to selected DataNodes.
>  # *Configurable I/O size*
>  ## Configurable block/file size.
>  ## Block-sized operations for consistent latency measurements.
>  # *Cold-read workload*
>  ## Pre-create a read corpus before testing.
>  ## Corpus can exceed DataNode page cache to reduce cache effects.
>  # *Latency and throughput metrics*
>  ## p50, p75, p95, p99, min, max, mean, and standard deviation.
>  ## Effective QPS and throughput for reads/writes.
>  # *Client-side scale-out*
>  ## Multiple clients can run concurrently.
>  ## Aggregate load is the sum of configured load across clients.
> Configuration is provided through a Java properties file, with individual 
> properties optionally overridden using `-D` through `ToolRunner`.
> h3. Configuration
> ||Property||Default||Description||
> |`favoredDataNodes`|None|Comma-separated DataNode host:port list used as 
> favored nodes for writes|
> |`replication`|3|Replication factor for generated files|
> |`blockSizeMB`|128|Block/file size for read/write operations|
> |`testWriteDirectory`|None|HDFS directory for write workload; omit to disable 
> writes|
> |`writeThroughputMB`|0|Target write throughput; `0` disables writes|
> |`endWriteThroughputMB`|0|Optional end value for linear write-throughput ramp|
> |`writeThreads`|-1|Number of writer threads; `-1` uses automatic 
> configuration|
> |`testReadDirectories`|None|HDFS directories for read workload/cold-read 
> corpus|
> |`readThroughputMB`|0|Target read throughput; `0` disables reads|
> |`endReadThroughputMB`|0|Optional end value for linear read-throughput ramp|
> |`readThreads`|-1|Number of reader threads; `-1` uses automatic configuration|
> |`testReadFileSizeGB`|0|Size of pre-created cold-read corpus; `0` disables 
> corpus generation|
> |`preTestWriteThroughputMB`|0|Optional throughput limit for corpus generation|
> |`preTestWriteDurationSeconds`|0|Optional time limit for corpus generation|
> |`testDurationSeconds`|60|Duration of measured workload|
> h3. Example
> {code:java}
> hadoop jar hadoop-hdfs-<version>-tests.jar \
>   org.apache.hadoop.hdfs.HdfsStressTest \
>   /path/to/stress.properties
> {code}
> h3. Results
> The stress tester provides:
>  * Controlled read/write load against targeted DataNodes.
>  * Reproducible workloads against specific replica sets.
>  * Cold-read workloads for storage-path testing.
>  * Client-side latency distributions and tail latency.
>  * Effective throughput and QPS measurements.
>  * Execution without MapReduce/YARN.
>  * Scale-out through multiple client processes.
> The tool complements rather than replaces `TestDFSIO`.
> h3. Comparison with TestDFSIO
> ||Dimension||TestDFSIO||HdfsStressTest||
> |Target specific DataNodes|Limited|Supported through favored-nodes hints|
> |Controlled throughput|Limited by MapReduce|Explicit control|
> |Throughput ramp|No|Supported|
> |Cold-read workload|Not specifically supported|Supported|
> |Client-side latency|Aggregate metrics|p50/p75/p95/p99 and other statistics|
> |External dependencies|MapReduce/YARN|HDFS client only|
> |Targeted DataNode stress|Difficult|Primary use case|
> |Scale-out|MapReduce-based|Multiple standalone clients|
> |Single-process execution|No|Yes|
> h3. Expected Impact
> Improves reproducibility and control for DataNode performance testing through:
>  * Precise offered-load control.
>  * Targeted DataNode/replica-set testing.
>  * Repeatable cold-read workloads.
>  * Client-side tail-latency measurements.
>  * Lightweight execution without MapReduce/YARN.
>  * Easy scale-out using multiple client processes.
> The primary benefit is {*}better observability and control during DataNode 
> performance and stress testing{*}, rather than production throughput 
> improvement.
> h2. Robustness and Input Validation
> The tool validates configuration at startup to prevent misleading benchmark 
> results:
>  * `blockSizeMB` must be positive.
>  * Read workloads (`readThroughputMB > 0`) require a positive 
> `testReadFileSizeGB`.
>  * If `preTestWriteDurationSeconds` limits corpus generation before the 
> requested `testReadFileSizeGB` is reached, the tool prints a *WARNING* that 
> reads may be served from OS page cache and recommends increasing or removing 
> the time limit.
> h2. Testing
> `TestHdfsStressTest` (MiniDFSCluster) covers:
>  * Block-sized file creation with configured replication and payload.
>  * Full-file reads to EOF.
>  * Pre-test cold-read corpus generation across multiple directories.
>  * Corpus generation duration limits and cache warnings.
>  * Write-only execution through `ToolRunner`.
> `TestHdfsStressTestHelpers` provides cluster-free unit tests for:
>  * Token-bucket rate limiting and pacing.
>  * Runtime rate changes.
>  * Latency conversion and sorting.
>  * Bounded-memory reservoir sampling.
>  * Configuration validation.
>  * Valid read-only and write-only configurations.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to