[
https://issues.apache.org/jira/browse/HDFS-17504?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17845841#comment-17845841
]
ASF GitHub Bot commented on HDFS-17504:
---------------------------------------
Hexiaoqiao commented on PR #6792:
URL: https://github.com/apache/hadoop/pull/6792#issuecomment-2107018371
@zhuzilong2013 Thanks for your report and contribution! IMO, they are
independent between different BPServiceActor, if exit DN process due to one
BPServiceActor issue, it will increase number of Dead DataNode from the whole
cluster view, where I don't think it is proper in Federation Arch. Another
side, maybe we could add some BPServiceActor count metric to monitor if
BPServiceActor works fine? Thanks again.
> DN process should exit when BPServiceActor exit
> -----------------------------------------------
>
> Key: HDFS-17504
> URL: https://issues.apache.org/jira/browse/HDFS-17504
> Project: Hadoop HDFS
> Issue Type: Bug
> Reporter: Zilong Zhu
> Assignee: Zilong Zhu
> Priority: Major
> Labels: pull-request-available
>
> BPServiceActor is a very important thread. In a non-HA cluster, the exit of
> the BPServiceActor thread will cause the DN process to exit. However, in a HA
> cluster, this is not the case.
> I found HDFS-15651 causes BPServiceActor thread to exit and sets the
> "runningState" from "RunningState.FAILED" to "RunningState.EXITED", it can
> be confusing during troubleshooting.
> I believe that the DN process should exit when the flag of the BPServiceActor
> is set to RunningState.FAILED because at this point, the DN is unable to
> recover and establish a heartbeat connection with the ANN on its own.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]