[ 
https://issues.apache.org/jira/browse/HDFS-17969?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

lei w updated HDFS-17969:
-------------------------
    Description: 
Our cluster upgrade follows this procedure:
1. The active NameNode is switched to standby (failover).
2. The (now standby) node is upgraded while the other node serves as active.
3. After the upgrade completes, the upgraded node is switched back to active.
Block missing was observed under the following race condition, which occurs 
when the upgraded node is switched back to active.

Root cause analysis:

1. NN restart/failover completes; the in-memory datanodemap is reset.
2. DN Actor thread sends a heartbeat; since the NN no longer recognizes this 
DN, the heartbeat response carries a registration command. Every DN receives 
this registration command after the failover.
3. The Actor thread enqueues the command into the command queue, which is 
consumed by the cmdThread.
4. The cmdThread starts registration. Since the NN is busy at this point 
(processing block reports from all DNs), the retrieveNamespaceInfo call 
preceding registration takes a long time and may even time out.
5. Because step 4 can take longer than the heartbeat interval, the Actor thread 
sends the next heartbeat before registration finishes. Upon success, the NN 
returns another registration command, which is enqueued again.
6. The first registration succeeds; the cmdThread schedules a full block report.
7. The Actor thread notices the scheduled full block report and triggers full 
block report (covering both incremental and stored blocks). This takes a long 
time.
8. Meanwhile, the cmdThread picks up the second registration command from the 
queue and starts registration again (retrieveNamespaceInfo, clear ibr, 
register). After it completes, another full block report is scheduled.
9. The ActorThread finishes the full block report and reschedules the next full 
block report to a random time within the configured FBR window .
Any blocks written between the moment the incremental block report (IBR) is 
sent in step 7 and the moment clear ibr executes in step 8 will not be reported 
to the NN until the next full block report.  If all DataNodes holding replicas 
of a given block are in this state, the NN never receives any report for that 
block within the gap window.

  was:
Our cluster upgrade follows this procedure:
1. The active NameNode is switched to standby (failover).
2. The (now standby) node is upgraded while the other node serves as active.
3. After the upgrade completes, the upgraded node is switched back to active.
Block missing was observed under the following race condition, which occurs 
when the upgraded node is switched back to active.

Root cause analysis:

1. NN restart/failover completes; the in-memory datanodemap is reset.
2. DN Actor thread sends a heartbeat; since the NN no longer recognizes this 
DN, the heartbeat response carries a registration command. Every DN receives 
this registration command after the failover.
3. The Actor thread enqueues the command into the command queue, which is 
consumed by the cmdThread.
4. The cmdThread starts registration. Since the NN is busy at this point 
(processing block reports from all DNs), the retrieveNamespaceInfo call 
preceding registration takes a long time and may even time out.
5. Because step 4 can take longer than the heartbeat interval, the Actor thread 
sends the next heartbeat before registration finishes. Upon success, the NN 
returns another registration command, which is enqueued again.
6. The first registration succeeds; the cmdThread schedules a full block report.
7. The Actor thread notices the scheduled full block report and triggers the 
BRThread.
8. The BRThread performs the full block report (covering both incremental and 
stored blocks). This takes a long time.
9. Meanwhile, the cmdThread picks up the second registration command from the 
queue and starts registration again (retrieveNamespaceInfo, clear ibr, 
register). After it completes, another full block report is scheduled.
10. The BRThread finishes the full block report and reschedules the next full 
block report to a random time within the configured FBR window .
Any blocks written between the moment the incremental block report (IBR) is 
sent in step 8 and the moment clear ibr executes in step 9 will not be reported 
to the NN until the next full block report.  If all DataNodes holding replicas 
of a given block are in this state, the NN never receives any report for that 
block within the gap window.


> Block missing due to re-registration race after failover in rolling upgrade
> ---------------------------------------------------------------------------
>
>                 Key: HDFS-17969
>                 URL: https://issues.apache.org/jira/browse/HDFS-17969
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: datanode
>            Reporter: lei w
>            Priority: Major
>
> Our cluster upgrade follows this procedure:
> 1. The active NameNode is switched to standby (failover).
> 2. The (now standby) node is upgraded while the other node serves as active.
> 3. After the upgrade completes, the upgraded node is switched back to active.
> Block missing was observed under the following race condition, which occurs 
> when the upgraded node is switched back to active.
> Root cause analysis:
> 1. NN restart/failover completes; the in-memory datanodemap is reset.
> 2. DN Actor thread sends a heartbeat; since the NN no longer recognizes this 
> DN, the heartbeat response carries a registration command. Every DN receives 
> this registration command after the failover.
> 3. The Actor thread enqueues the command into the command queue, which is 
> consumed by the cmdThread.
> 4. The cmdThread starts registration. Since the NN is busy at this point 
> (processing block reports from all DNs), the retrieveNamespaceInfo call 
> preceding registration takes a long time and may even time out.
> 5. Because step 4 can take longer than the heartbeat interval, the Actor 
> thread sends the next heartbeat before registration finishes. Upon success, 
> the NN returns another registration command, which is enqueued again.
> 6. The first registration succeeds; the cmdThread schedules a full block 
> report.
> 7. The Actor thread notices the scheduled full block report and triggers full 
> block report (covering both incremental and stored blocks). This takes a long 
> time.
> 8. Meanwhile, the cmdThread picks up the second registration command from the 
> queue and starts registration again (retrieveNamespaceInfo, clear ibr, 
> register). After it completes, another full block report is scheduled.
> 9. The ActorThread finishes the full block report and reschedules the next 
> full block report to a random time within the configured FBR window .
> Any blocks written between the moment the incremental block report (IBR) is 
> sent in step 7 and the moment clear ibr executes in step 8 will not be 
> reported to the NN until the next full block report.  If all DataNodes 
> holding replicas of a given block are in this state, the NN never receives 
> any report for that block within the gap window.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to