Rajan Dhabalia created HDFS-17977:
-------------------------------------

             Summary: HDFS-Test Add a compact standalone HDFS DataNode 
read/write stress tester
                 Key: HDFS-17977
                 URL: https://issues.apache.org/jira/browse/HDFS-17977
             Project: Hadoop HDFS
          Issue Type: Improvement
          Components: benchmarks, hdfs, test
            Reporter: Rajan Dhabalia


h2. Summary

Add a compact, single-process HDFS read/write stress tester that generates 
controlled throughput against targeted DataNodes and reports client-side 
latency distributions.
h2. Description
h3. Motivation

`TestDFSIO` is the standard HDFS I/O benchmark but is not well suited for 
targeted DataNode stress testing.

Key limitations:
 * Requires MapReduce/YARN.
 * Load is distributed through MapReduce scheduling rather than directly 
targeting DataNodes.
 * Limited control over steady throughput/QPS against selected DataNodes.
 * Primarily reports aggregate throughput rather than client-side p50/p95/p99 
latency.
 * Does not provide a reliable mechanism for cold-read workloads targeting disk 
I/O.

`HdfsStressTest` provides a lightweight standalone tool for controlled DataNode 
performance and stress testing.
h3. Approach

Introduce `HdfsStressTest`, a single-process load generator that depends only 
on the HDFS client.

Key capabilities:
 # *Controlled throughput*
 ## Global token-bucket rate limiter.
 ## Configurable read/write throughput.
 ## Optional linear throughput ramp.
 # *Targeted DataNodes*
 ## Use HDFS favored-nodes hints to direct writes to selected DataNodes.
 # *Configurable I/O size*
 ## Configurable block/file size.
 ## Block-sized operations for consistent latency measurements.
 # *Cold-read workload*
 ## Pre-create a read corpus before testing.
 ## Corpus can exceed DataNode page cache to reduce cache effects.
 # *Latency and throughput metrics*
 ## p50, p75, p95, p99, min, max, mean, and standard deviation.
 ## Effective QPS and throughput for reads/writes.
 # *Client-side scale-out*
 ## Multiple clients can run concurrently.
 ## Aggregate load is the sum of configured load across clients.

Configuration is provided through a Java properties file, with individual 
properties optionally overridden using `-D` through `ToolRunner`.
h3. Configuration
||Property||Default||Description||
|`favoredDataNodes`|None|Comma-separated DataNode host:port list used as 
favored nodes for writes|
|`replication`|3|Replication factor for generated files|
|`blockSizeMB`|128|Block/file size for read/write operations|
|`testWriteDirectory`|None|HDFS directory for write workload; omit to disable 
writes|
|`writeThroughputMB`|0|Target write throughput; `0` disables writes|
|`endWriteThroughputMB`|0|Optional end value for linear write-throughput ramp|
|`writeThreads`|-1|Number of writer threads; `-1` uses automatic configuration|
|`testReadDirectories`|None|HDFS directories for read workload/cold-read corpus|
|`readThroughputMB`|0|Target read throughput; `0` disables reads|
|`endReadThroughputMB`|0|Optional end value for linear read-throughput ramp|
|`readThreads`|-1|Number of reader threads; `-1` uses automatic configuration|
|`testReadFileSizeGB`|0|Size of pre-created cold-read corpus; `0` disables 
corpus generation|
|`preTestWriteThroughputMB`|0|Optional throughput limit for corpus generation|
|`preTestWriteDurationSeconds`|0|Optional time limit for corpus generation|
|`testDurationSeconds`|60|Duration of measured workload|
h3. Example
{code:java}
hadoop jar hadoop-hdfs-<version>-tests.jar \
  org.apache.hadoop.hdfs.HdfsStressTest \
  /path/to/stress.properties
{code}
h3. Results

The stress tester provides:
 * Controlled read/write load against targeted DataNodes.
 * Reproducible workloads against specific replica sets.
 * Cold-read workloads for storage-path testing.
 * Client-side latency distributions and tail latency.
 * Effective throughput and QPS measurements.
 * Execution without MapReduce/YARN.
 * Scale-out through multiple client processes.

The tool complements rather than replaces `TestDFSIO`.
h3. Comparison with TestDFSIO
||Dimension||TestDFSIO||HdfsStressTest||
|Target specific DataNodes|Limited|Supported through favored-nodes hints|
|Controlled throughput|Limited by MapReduce|Explicit control|
|Throughput ramp|No|Supported|
|Cold-read workload|Not specifically supported|Supported|
|Client-side latency|Aggregate metrics|p50/p75/p95/p99 and other statistics|
|External dependencies|MapReduce/YARN|HDFS client only|
|Targeted DataNode stress|Difficult|Primary use case|
|Scale-out|MapReduce-based|Multiple standalone clients|
|Single-process execution|No|Yes|
h3. Expected Impact

Improves reproducibility and control for DataNode performance testing through:
 * Precise offered-load control.
 * Targeted DataNode/replica-set testing.
 * Repeatable cold-read workloads.
 * Client-side tail-latency measurements.
 * Lightweight execution without MapReduce/YARN.
 * Easy scale-out using multiple client processes.

The primary benefit is {*}better observability and control during DataNode 
performance and stress testing{*}, rather than production throughput 
improvement.
h2. Robustness and Input Validation

The tool validates configuration at startup to prevent misleading benchmark 
results:
 * `blockSizeMB` must be positive.
 * Read workloads (`readThroughputMB > 0`) require a positive 
`testReadFileSizeGB`.
 * If `preTestWriteDurationSeconds` limits corpus generation before the 
requested `testReadFileSizeGB` is reached, the tool prints a *WARNING* that 
reads may be served from OS page cache and recommends increasing or removing 
the time limit.

h2. Testing

`TestHdfsStressTest` (MiniDFSCluster) covers:
 * Block-sized file creation with configured replication and payload.
 * Full-file reads to EOF.
 * Pre-test cold-read corpus generation across multiple directories.
 * Corpus generation duration limits and cache warnings.
 * Write-only execution through `ToolRunner`.

`TestHdfsStressTestHelpers` provides cluster-free unit tests for:
 * Token-bucket rate limiting and pacing.
 * Runtime rate changes.
 * Latency conversion and sorting.
 * Bounded-memory reservoir sampling.
 * Configuration validation.
 * Valid read-only and write-only configurations.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to