[ 
https://issues.apache.org/jira/browse/HDFS-17964?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Aleksei Ieshin updated HDFS-17964:
----------------------------------
    Description: 
Observer NameNodes fatally reject delegation tokens they have not yet tailed.   
                                                                                
              
                                                                                
                                                                                
                *DelegationTokenSecretManager.retriableRetrievePassword* and 
*FSNamesystem.verifyToken* wrap *InvalidToken* in a *RetriableException* only 
when *inTransitionToActive()*. An observer in steady state never is, so a token 
the active issued moments ago fails with 

{code:java}
InvalidToken: Token for real user: <user>, can't be found in cache      
{code}

HDFS-5322 added that guard in 2.3.0, before Observer NameNodes existed 
(HDFS-12943). Widening it to cover observers fixes both paths.

h6. Details:
The common trigger is *WebHDFS* *op=OPEN*, the one operation with a second hop: 
the NameNode redirects to a DataNode, and that DataNode validates the token 
against a NameNode chosen by its own *ConfiguredFailoverProxyProvider*. Nothing 
the client configures influences that choice, so in a three-NameNode HA setup 
with one observer a substantial fraction of reads land on the observer and fail 
with HTTP 403.                                                                  
                                                                                
         
                                                                                
                                                                                
                
Fix: treat an observer as possibly behind on tailing at both sites, so a failed 
lookup yields *RetriableException*. Callers already handle it - that is what 
HDFS-5322 relies on. No config, no new mechanism.

  was:
Observer NameNodes fatally reject delegation tokens they have not yet tailed.   
                                                                                
              
                                                                                
                                                                                
                *DelegationTokenSecretManager.retriableRetrievePassword* and 
*FSNamesystem.verifyToken* wrap *InvalidToken* in a *RetriableException* only 
when *inTransitionToActive()*. An observer in steady state never is, so a token 
the active issued moments ago fails with 

{code:java}
InvalidToken: Token for real user: <user>, can't be found in cache      
{code}

HDFS-5322 added that guard in 2.3.0, before Observer NameNodes existed 
(HDFS-12943). Widening it to cover observers fixes both paths.

h6. Details:
The common trigger is *WebHDFS* *op=OPEN*, the one operation with a second hop: 
 The common trigger is WebHDFS op=OPEN, the one operation with a second hop: 
the NameNode redirects to a DataNode, and that DataNode validates the token 
against a NameNode chosen by its own *ConfiguredFailoverProxyProvider*. Nothing 
the client configures influences that choice, so in a three-NameNode HA setup 
with one observer a substantial fraction of reads land on the observer and fail 
with HTTP 403.                                                                  
                                                                                
         
                                                                                
                                                                                
                
Fix: treat an observer as possibly behind on tailing at both sites, so a failed 
lookup yields *RetriableException*. Callers already handle it - that is what 
HDFS-5322 relies on. No config, no new mechanism.


> HDFS delegation token not found in cache errors seen on Observer NameNodes
> --------------------------------------------------------------------------
>
>                 Key: HDFS-17964
>                 URL: https://issues.apache.org/jira/browse/HDFS-17964
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: hdfs
>            Reporter: Aleksei Ieshin
>            Priority: Major
>
> Observer NameNodes fatally reject delegation tokens they have not yet tailed. 
>                                                                               
>                   
>                                                                               
>                                                                               
>                     *DelegationTokenSecretManager.retriableRetrievePassword* 
> and *FSNamesystem.verifyToken* wrap *InvalidToken* in a *RetriableException* 
> only when *inTransitionToActive()*. An observer in steady state never is, so 
> a token the active issued moments ago fails with 
> {code:java}
> InvalidToken: Token for real user: <user>, can't be found in cache      
> {code}
> HDFS-5322 added that guard in 2.3.0, before Observer NameNodes existed 
> (HDFS-12943). Widening it to cover observers fixes both paths.
> h6. Details:
> The common trigger is *WebHDFS* *op=OPEN*, the one operation with a second 
> hop: the NameNode redirects to a DataNode, and that DataNode validates the 
> token against a NameNode chosen by its own *ConfiguredFailoverProxyProvider*. 
> Nothing the client configures influences that choice, so in a three-NameNode 
> HA setup with one observer a substantial fraction of reads land on the 
> observer and fail with HTTP 403.                                              
>                                                                               
>                                
>                                                                               
>                                                                               
>                     
> Fix: treat an observer as possibly behind on tailing at both sites, so a 
> failed lookup yields *RetriableException*. Callers already handle it - that 
> is what HDFS-5322 relies on. No config, no new mechanism.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to