[ 
https://issues.apache.org/jira/browse/HDFS-17906?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

dzcxzl updated HDFS-17906:
--------------------------
    Description: 
 

HDFS-16942 introduced `InvalidBlockReportLeaseException`, which the NameNode 
now throws back to the DataNode via RPC when a block report is rejected due to 
an invalid lease. On a DataNode that also includes HDFS-16942, the exception is 
caught and `fullBlockReportLeaseId` is reset to 0, allowing the DN to request a 
new lease on the next heartbeat and retry.

However, during a rolling upgrade where the NameNode has been upgraded (with 
HDFS-16942) but DataNodes are still running an older version (without 
HDFS-16942), the old DataNode code does not have the 
`InvalidBlockReportLeaseException` handling branch in 
`BPServiceActor.offerService()`. This causes the DN to enter an infinite 
failure loop where it can never successfully send a full block report.

  was:
## Description

HDFS-16942 introduced `InvalidBlockReportLeaseException`, which the NameNode 
now throws back to the DataNode via RPC when a block report is rejected due to 
an invalid lease. On a DataNode that also includes HDFS-16942, the exception is 
caught and `fullBlockReportLeaseId` is reset to 0, allowing the DN to request a 
new lease on the next heartbeat and retry.

However, during a rolling upgrade where the NameNode has been upgraded (with 
HDFS-16942) but DataNodes are still running an older version (without 
HDFS-16942), the old DataNode code does not have the 
`InvalidBlockReportLeaseException` handling branch in 
`BPServiceActor.offerService()`. This causes the DN to enter an infinite 
failure loop where it can never successfully send a full block report.

## Root Cause

In `BPServiceActor.offerService()`, the logic works as follows:

1. The DN requests a block report lease during heartbeat only when 
`fullBlockReportLeaseId == 0`:

```java
boolean requestBlockReportLease = (fullBlockReportLeaseId == 0) &&
        scheduler.isBlockReportDue(startTime);
```

2. After sending the block report, `fullBlockReportLeaseId` is reset to 0:

```java
if ((fullBlockReportLeaseId != 0) || forceFullBr) {
    cmds = blockReport(fullBlockReportLeaseId);
    fullBlockReportLeaseId = 0;  // not reached if blockReport() throws
}
```

3. When the upgraded NN throws `InvalidBlockReportLeaseException`, 
`blockReport()` propagates the exception. The `fullBlockReportLeaseId = 0` line 
after the call is never executed.

4. The exception is caught by the generic `RemoteException` catch block. The 
old DN code does not recognize `InvalidBlockReportLeaseException`, so 
`fullBlockReportLeaseId` remains set to the stale invalid value.

5. On the next heartbeat iteration, because `fullBlockReportLeaseId != 0`, 
`requestBlockReportLease` is `false` — the DN does not request a new lease. It 
then attempts `blockReport()` again with the same stale lease, which the NN 
rejects again. This repeats indefinitely.


> Rolling upgrade: old DataNodes get stuck in infinite invalid block report 
> lease loop
> ------------------------------------------------------------------------------------
>
>                 Key: HDFS-17906
>                 URL: https://issues.apache.org/jira/browse/HDFS-17906
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: datanode
>            Reporter: dzcxzl
>            Priority: Major
>
>  
> HDFS-16942 introduced `InvalidBlockReportLeaseException`, which the NameNode 
> now throws back to the DataNode via RPC when a block report is rejected due 
> to an invalid lease. On a DataNode that also includes HDFS-16942, the 
> exception is caught and `fullBlockReportLeaseId` is reset to 0, allowing the 
> DN to request a new lease on the next heartbeat and retry.
> However, during a rolling upgrade where the NameNode has been upgraded (with 
> HDFS-16942) but DataNodes are still running an older version (without 
> HDFS-16942), the old DataNode code does not have the 
> `InvalidBlockReportLeaseException` handling branch in 
> `BPServiceActor.offerService()`. This causes the DN to enter an infinite 
> failure loop where it can never successfully send a full block report.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to