[ 
https://issues.apache.org/jira/browse/HDFS-17974?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112107#comment-18112107
 ] 

ASF GitHub Bot commented on HDFS-17974:
---------------------------------------

hadoop-yetus commented on PR #8715:
URL: https://github.com/apache/hadoop/pull/8715#issuecomment-5562109001

   :confetti_ball: **+1 overall**
   
   
   
   
   
   
   | Vote | Subsystem | Runtime |  Logfile | Comment |
   |:----:|----------:|--------:|:--------:|:-------:|
   | +0 :ok: |  reexec  |   0m 57s |  |  Docker mode activated.  |
   |||| _ Prechecks _ |
   | +1 :green_heart: |  dupname  |   0m  0s |  |  No case conflicting files 
found.  |
   | +0 :ok: |  codespell  |   0m  1s |  |  codespell was not available.  |
   | +0 :ok: |  detsecrets  |   0m  1s |  |  detect-secrets was not available.  
|
   | +0 :ok: |  xmllint  |   0m  1s |  |  xmllint was not available.  |
   | +1 :green_heart: |  @author  |   0m  0s |  |  The patch does not contain 
any @author tags.  |
   | +1 :green_heart: |  test4tests  |   0m  0s |  |  The patch appears to 
include 1 new or modified test files.  |
   |||| _ trunk Compile Tests _ |
   | +1 :green_heart: |  mvninstall  |  47m 14s |  |  trunk passed  |
   | +1 :green_heart: |  compile  |   1m 43s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  compile  |   1m 47s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  checkstyle  |   1m 53s |  |  trunk passed  |
   | +1 :green_heart: |  mvnsite  |   1m 57s |  |  trunk passed  |
   | +1 :green_heart: |  javadoc  |   1m 30s |  |  trunk passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   1m 30s |  |  trunk passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   4m 22s |  |  trunk passed  |
   | +1 :green_heart: |  shadedclient  |  37m 34s |  |  branch has no errors 
when building and testing our client artifacts.  |
   |||| _ Patch Compile Tests _ |
   | +1 :green_heart: |  mvninstall  |   1m 24s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   1m 17s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   1m 17s |  |  the patch passed  |
   | +1 :green_heart: |  compile  |   1m 23s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javac  |   1m 23s |  |  the patch passed  |
   | +1 :green_heart: |  blanks  |   0m  0s |  |  The patch has no blanks 
issues.  |
   | -0 :warning: |  checkstyle  |   1m 22s | 
[/results-checkstyle-hadoop-hdfs-project_hadoop-hdfs.txt](https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8715/6/artifact/out/results-checkstyle-hadoop-hdfs-project_hadoop-hdfs.txt)
 |  hadoop-hdfs-project/hadoop-hdfs: The patch generated 27 new + 382 unchanged 
- 0 fixed = 409 total (was 382)  |
   | +1 :green_heart: |  mvnsite  |   1m 28s |  |  the patch passed  |
   | +1 :green_heart: |  javadoc  |   0m 58s |  |  the patch passed with JDK 
Ubuntu-21.0.12+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  javadoc  |   1m  3s |  |  the patch passed with JDK 
Ubuntu-17.0.20+8-1-24.04-Ubuntu  |
   | +1 :green_heart: |  spotbugs  |   4m  4s |  |  the patch passed  |
   | +1 :green_heart: |  shadedclient  |  36m 45s |  |  patch has no errors 
when building and testing our client artifacts.  |
   |||| _ Other Tests _ |
   | +1 :green_heart: |  unit  | 267m 39s |  |  hadoop-hdfs in the patch 
passed.  |
   | +1 :green_heart: |  asflicense  |   0m 53s |  |  The patch does not 
generate ASF License warnings.  |
   |  |   | 416m 47s |  |  |
   
   
   | Subsystem | Report/Notes |
   |----------:|:-------------|
   | Docker | ClientAPI=1.56 ServerAPI=1.56 base: 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8715/6/artifact/out/Dockerfile
 |
   | GITHUB PR | https://github.com/apache/hadoop/pull/8715 |
   | Optional Tests | dupname asflicense compile javac javadoc mvninstall 
mvnsite unit shadedclient spotbugs checkstyle codespell detsecrets xmllint |
   | uname | Linux fd090cc4226b 5.15.0-185-generic #195-Ubuntu SMP Fri Jun 19 
17:11:50 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux |
   | Build tool | maven |
   | Personality | dev-support/bin/hadoop.sh |
   | git revision | trunk / 095387b0daafe844d8744b5ad3ab247a06e15906 |
   | Default Java | Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   | Multi-JDK versions | 
/usr/lib/jvm/java-21-openjdk-amd64:Ubuntu-21.0.12+8-1-24.04-Ubuntu 
/usr/lib/jvm/java-17-openjdk-amd64:Ubuntu-17.0.20+8-1-24.04-Ubuntu |
   |  Test Results | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8715/6/testReport/ |
   | Max. process+thread count | 3343 (vs. ulimit of 10000) |
   | modules | C: hadoop-hdfs-project/hadoop-hdfs U: 
hadoop-hdfs-project/hadoop-hdfs |
   | Console output | 
https://ci-hadoop.apache.org/job/hadoop-multibranch/job/PR-8715/6/console |
   | versions | git=2.43.0 maven=3.9.15 spotbugs=4.9.7 |
   | Powered by | Apache Yetus 0.14.1 https://yetus.apache.org |
   
   
   This message was automatically generated.
   
   




> HDFS DataNode Affinity for Tenant Isolation
> -------------------------------------------
>
>                 Key: HDFS-17974
>                 URL: https://issues.apache.org/jira/browse/HDFS-17974
>             Project: Hadoop HDFS
>          Issue Type: Improvement
>          Components: hdfs
>            Reporter: Rajan Dhabalia
>            Priority: Major
>              Labels: pull-request-available
>
> h2. Summary
> Introduce a pluggable, regex-based DataNode affinity mechanism that maps HDFS 
> paths to dedicated DataNode pools.
> This enables *tenant/dataset-level storage and I/O isolation* without 
> requiring separate HDFS clusters or rack-based workarounds.
> h2. Motivation
> The current HDFS block placement policy has no native path-based mechanism to 
> restrict data to a specific DataNode pool. This can cause noisy-neighbor 
> interference for isolation-sensitive workloads.
> The feature provides:
>  * Path-to-DataNode-pool mapping
>  * Tenant/dataset I/O isolation
>  * More predictable performance
>  * Independent capacity planning
>  * Runtime configuration updates without NameNode restart
> h2. Design
> Introduce a pluggable `DatanodeAffinityManager` configured through:
> {code:java}
> dfs.datanode.affinity.manager.classname
> {code}
> An affinity rule maps:
> {code:java}
> HDFS path regex -> DataNode hostname regex
> {code}
> Example:
> {code:java}
> /data/tenantA/.* -> dn-tenantA-.*
> /data/tenantB/.* -> dn-tenantB-.*
> {code}
> The manager resolves the hostname regex against registered DataNodes and 
> builds a restricted `NetworkTopology` containing only eligible DataNodes for 
> each affinity group.
> h2. NameNode Integration
> *DatanodeManager*
>  * Identifies DataNodes belonging to affinity pools during registration.
>  * Removes affinity-only DataNodes from the default placement topology.
>  * Prevents non-affinity workloads from using dedicated DataNodes.
>  * Refreshes affinity state through `hdfs dfsadmin -refreshNodes`.
> *BlockManager*
> For each block placement request:
>  # Match the source path against configured affinity groups.
>  # Use the group's `BlockPlacementPolicy` if matched.
>  # Select targets from the group's restricted topology.
>  # Fall back to the default placement policy when no group matches.
> This avoids large exclusion lists on the placement hot path.
> h2. Pluggable Implementation
> `DatanodeAffinityManager` is an abstraction that allows different affinity 
> sources without changing block-placement logic.
> Built-in implementation:
> {code:java}
> FileDatanodeAffinityManager
> {code}
> It loads affinity rules from a JSON configuration file and reloads them 
> through `dfsadmin -refreshNodes`.
> h2. Configuration
> ||Property||Default||Description||
> |`dfs.datanode.affinity.manager.classname`|Empty|`DatanodeAffinityManager` 
> implementation. Empty disables the feature.|
> |`dfs.datanode.affinity.file.path`|Empty|JSON affinity configuration used by 
> `FileDatanodeAffinityManager`.|
> Example:
> {code:java}
> /data/tenantA/.* -> dn-tenantA-.*
> /data/tenantB/.* -> dn-tenantB-.*
> {code}
> h2. Operational Visibility
> Add:
> {code:java}
> hdfs fsck <path> -favored-nodes
> {code}
> to display the DataNodes resolved for a path.
> Example:
> {code:java}
> hdfs fsck /data/tenantA -favored-nodes
> {code}
> This provides a dry-run mechanism to validate affinity configuration before 
> enabling it.
> h2. Runtime Refresh
> Affinity configuration and DataNode membership can be updated using:
> {code:java}
> hdfs dfsadmin -refreshNodes
> {code}
> No NameNode restart is required.
> Supported changes include:
>  * Adding/removing DataNodes from an affinity pool
>  * Updating path-to-pool mappings
>  * Changing hostname matching rules
> h2. Benefits
>  * *Tenant Isolation:* Affinity-enabled data is placed only on its dedicated 
> DataNode pool.
>  * *Predictable Performance:* Reduces noisy-neighbor impact and isolates I/O 
> capacity.
>  * *Default Pool Protection:* Dedicated DataNodes are excluded from normal 
> placement.
>  * *Efficient Placement:* Restricted topology limits placement to the 
> relevant pool.
>  * *Runtime Configuration:* Changes take effect without NameNode restart.
> h2. Expected Impact
> ||Dimension||Expected Impact||
> |Cross-tenant interference|Eliminated within the isolated DataNode pool|
> |Tail latency|Reduced when contention exists|
> |Placement scope|Limited to affinity pool|
> |Default-pool throughput|No expected regression|
> |Configuration changes|Runtime refresh|
> The primary goal is {*}isolation and performance predictability{*}, not raw 
> throughput improvement.
> h2. Implementation
> Key changes:
>  * `DatanodeAffinityManager` abstraction
>  * `FileDatanodeAffinityManager`
>  * Path-regex to DataNode-hostname-regex mapping
>  * Per-group restricted `NetworkTopology`
>  * `DatanodeManager` topology integration
>  * Per-group `BlockPlacementPolicy` in `BlockManager`
>  * Affinity-aware `chooseTarget4NewBlock`
>  * Runtime refresh via `dfsadmin -refreshNodes`
>  * `hdfs fsck -favored-nodes` validation
> h2. Backward Compatibility
>  * Disabled by default.
>  * Existing block placement behavior is unchanged when affinity is not 
> configured.
>  * Non-matching paths continue to use the default placement policy.
>  * No HDFS client changes are required.
>  * Can be enabled selectively for specific tenants or datasets.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to