[ 
https://issues.apache.org/jira/browse/HDFS-17976?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18111814#comment-18111814
 ] 

ASF GitHub Bot commented on HDFS-17976:
---------------------------------------

hadoop-yetus commented on PR #8718:
URL: https://github.com/apache/hadoop/pull/8718#issuecomment-5550348395

   :broken_heart: **-1 overall**
   
   
   
   
   
   
   | Vote | Subsystem | Runtime |  Logfile | Comment |
   |:----:|----------:|--------:|:--------:|:-------:|
   | +0 :ok: |  reexec  |   0m 22s |  |  Docker mode activated.  |
   |||| _ Prechecks _ |
   | +1 :green_heart: |  dupname  |   0m  0s |  |  No case conflicting files 
found.  |
   | +0 :ok: |  codespell  |   0m  0s |  |  codespell was not available.  |
   | +0 :ok: |  detsecrets  |   0m  0s |  |  detect-secrets was not available.  
|
   | +0 :ok: |  xmllint  |   0m  0s |  |  xmllint was not available.  |
   | +1 :green_heart: |  @author  |   0m  0s |  |  The patch does not contain 
any @author tags.  |
   | +1 :green_heart: |  test4tests  |   0m  0s |  |  The patch appears to 
include 2 new or modified test files.  |
   |||| _ trunk Compile Tests _ |
   | +1 :green_heart: |  mvninstall  |  26m 50s |  |  trunk passed  |
   | +1 :green_heart: |  compile  |   0m 59s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  compile  |   0m 59s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  checkstyle  |   1m  8s |  |  trunk passed  |
   | +1 :green_heart: |  mvnsite  |   1m  7s |  |  trunk passed  |
   | +1 :green_heart: |  javadoc  |   0m 57s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   0m 57s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   2m 19s |  |  trunk passed  |
   | +1 :green_heart: |  shadedclient  |  18m 22s |  |  branch has no errors 
when building and testing our client artifacts.  |
   | -0 :warning: |  patch  |  18m 41s |  |  Used diff version of patch file. 
Binary files and potentially other changes not applied. Please rebase and 
squash commits if necessary.  |
   |||| _ Patch Compile Tests _ |
   | +1 :green_heart: |  mvninstall  |   0m 50s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   0m 42s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   0m 42s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   0m 47s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   0m 47s |  |  the patch passed  |
   | +1 :green_heart: |  blanks  |   0m  0s |  |  The patch has no blanks 
issues.  |
   | -0 :warning: |  checkstyle  |   0m 42s | 
[/results-checkstyle-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8718/2/artifact/out/results-checkstyle-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs-project/hadoop-hdfs: The patch generated 5 new + 309 unchanged 
- 0 fixed = 314 total (was 309)  |
   | +1 :green_heart: |  mvnsite  |   0m 50s |  |  the patch passed  |
   | +1 :green_heart: |  javadoc  |   0m 35s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   0m 37s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   2m  5s |  |  the patch passed  |
   | +1 :green_heart: |  shadedclient  |  16m 53s |  |  patch has no errors 
when building and testing our client artifacts.  |
   |||| _ Other Tests _ |
   | -1 :x: |  unit  | 185m 55s | 
[/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8718/2/artifact/out/patch-unit-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs in the patch passed.  |
   | +1 :green_heart: |  asflicense  |   0m 30s |  |  The patch does not 
generate ASF License warnings.  |
   |  |   | 263m 40s |  |  |
   
   
   | Reason | Tests |
   |-------:|:------|
   | Failed junit tests | 
hadoop.hdfs.server.balancer.TestBalancerWithHANameNodes |
   
   
   | Subsystem | Report/Notes |
   |----------:|:-------------|
   | Docker | ClientAPI=1.56 ServerAPI=1.56 base: 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8718/2/artifact/out/Dockerfile
 |
   | Optional Tests | dupname asflicense compile javac javadoc mvninstall 
mvnsite unit shadedclient spotbugs checkstyle codespell detsecrets xmllint |
   | uname | Linux 1601862756d9 5.15.0-190-generic #200-Ubuntu SMP Fri Aug 7 
15:06:04 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux |
   | Build tool | maven |
   | Personality | dev-support/bin/hadoop.sh |
   | git revision | trunk / 706d4dcb3266f35fcd5df77add1352a5541d28c4 |
   | Default Java | Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   | Multi-JDK versions | 
/usr/lib/jvm/java-21-openjdk-amd64:Ubuntu-21.0.12+8-1-24.04-Ubuntu 
/usr/lib/jvm/java-17-openjdk-amd64:Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   |  Test Results | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8718/2/testReport/ |
   | Max. process+thread count | 4375 (vs. ulimit of 10000) |
   | modules | C: hadoop-hdfs-project/hadoop-hdfs U: 
hadoop-hdfs-project/hadoop-hdfs |
   | Console output | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8718/2/console |
   | versions | git=2.43.0 maven=3.9.15 spotbugs=4.9.7 |
   | Powered by | Apache Yetus 0.14.1 https://yetus.apache.org |
   
   
   This message was automatically generated.
   
   




> HDFS DataNode add configurable inactivity-based timeout for DataNode block 
> transfers
> ------------------------------------------------------------------------------------
>
>                 Key: HDFS-17976
>                 URL: https://issues.apache.org/jira/browse/HDFS-17976
>             Project: Hadoop HDFS
>          Issue Type: Improvement
>          Components: hdfs
>            Reporter: Rajan Dhabalia
>            Priority: Major
>              Labels: pull-request-available
>
> h2. Summary
> Add a configurable inactivity-based timeout to safely terminate stalled 
> DataNode block transfers, reclaim resources, and preserve partially written 
> replicas for recovery.
> h2. Issue Type
> Improvement
> h2. Component/s
> datanode
> h2. Description
> h3. Motivation
> A block write remains active until the client sends the final packet and the 
> block is finalized. If a client crashes, loses connectivity, or becomes 
> unresponsive, the DataNode transfer thread can remain blocked waiting for 
> additional data.
> Stalled transfers can accumulate over time, consuming transfer threads and 
> other resources and potentially affecting healthy client operations.
> h3. Approach
> Introduce a configurable inactivity timeout for DataNode block transfers.
> When enabled:
>  * Track packet-arrival activity for each ongoing block write.
>  * Schedule an inactivity check for each active transfer.
>  * Terminate the transfer if no packet is received within the configured 
> timeout.
>  * Flush buffered data before termination.
>  * Close the input stream to interrupt blocked reads.
>  * Keep the replica in *RBW (Replica Being Written)* state.
>  * Allow existing NameNode lease recovery and block synchronization to 
> recover and finalize the replica.
> The timeout uses a shared `ScheduledExecutorService` and is created lazily 
> only when enabled. When disabled, there are no additional scheduler threads 
> or timeout processing.
> h3. Configuration
> ||Property||Default||Description||
> |`dfs.datanode.last.packet.receive.timeout.ms`|0|Inactivity timeout for an 
> ongoing block transfer. If no packet is received within the configured 
> window, the transfer is considered stalled and terminated. `0` or negative 
> disables the feature.|
> h3. Transfer Lifecycle
> {code:java}
> Client starts block write
>         |
>         v
> DataNode receives packets
>         |
>         v
> Track packet activity
>         |
>         v
> Packet received?
>    /            \
>  Yes             No
>   |               |
>   v               v
> Reset timer    Timeout reached
>                   |
>                   v
>           Flush buffered data
>                   |
>                   v
>           Close input stream
>                   |
>                   v
>           Replica remains RBW
>                   |
>                   v
>        Existing lease recovery /
>        block synchronization
>                   |
>                   v
>             Block finalized
> {code}
> h3. Data Safety
> When a stalled transfer is terminated:
>  # Buffered data is flushed.
>  # The input stream is closed.
>  # The replica remains in RBW state.
>  # Existing HDFS recovery mechanisms handle subsequent recovery/finalization.
> This preserves data already received by the DataNode while reclaiming stalled 
> transfer resources.
> h3. Results
> The inactivity timeout provides:
>  * Bounded resource usage for stalled transfers.
>  * Prevention of indefinite transfer-thread retention after client failures.
>  * Reduced risk of transfer-thread exhaustion during client failure bursts.
>  * Preservation of partially written replica data.
>  * Compatibility with existing HDFS recovery mechanisms.
>  * No additional overhead when disabled.
> h3. Estimated Impact
> This is primarily a reliability and resource-reclamation improvement rather 
> than a throughput optimization.
> ||Dimension||Without Timeout||With Inactivity Timeout||
> |Stalled transfer lifetime|Potentially unbounded|Bounded by configured 
> timeout|
> |Transfer threads|Can accumulate|Reclaimed after inactivity|
> |Client failure bursts|Risk of resource exhaustion|Resource usage remains 
> bounded|
> |Partially written data|Preserved through existing recovery|Preserved; 
> replica remains RBW|
> |Normal transfers|Existing behavior|Unchanged|
> |Disabled overhead|Existing behavior|No additional scheduler/timeout 
> processing|
> h3. Operational Considerations
> The timeout should be configured based on workload characteristics. An overly 
> aggressive timeout could terminate legitimately slow transfers, while a 
> sufficiently large timeout allows temporary network or client stalls while 
> still reclaiming resources from genuinely stalled transfers.
> The feature is therefore *disabled by default* and can be enabled explicitly 
> by operators.
> h2. Expected Impact
>  * Improve DataNode resilience to crashed, disconnected, or hung clients.
>  * Prevent stalled transfers from holding resources indefinitely.
>  * Bound transfer-resource usage during client failure bursts.
>  * Preserve partially written data for existing recovery mechanisms.
>  * Keep normal block-transfer behavior unchanged when disabled.
> h2. Backward Compatibility
> The feature is {*}disabled by default{*}. Existing DataNode block-transfer 
> behavior remains unchanged unless 
> `dfs.datanode.last.packet.receive.timeout.ms` is explicitly configured.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to