[ 
https://issues.apache.org/jira/browse/HDFS-17897?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ZhenyuLi updated HDFS-17897:
----------------------------
    Description: 
HDFS-12931 added handling for InvalidEncryptionKeyException in 
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
with encrypted data transfer. However, the parallel striped file path 
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.

 
Both paths call {{{}DFSClient.connectToDN(){}}}, which performs a SASL 
handshake using a cached {{DataEncryptionKey}} (DEK). After NameNode key 
rotation, the client may still hold a cached DEK derived from an old 
{{BlockKey}} that has already expired and been removed from DataNodes (as 
described in HDFS-12931), causing the handshake to fail with 
{{{}InvalidEncryptionKeyException{}}}.
 

In the replicated path, this exception is caught in HDFS-12931's fix, 
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block 
is retried. In the striped path, the exception falls through to the generic 
catch (IOException) block, which only logs a warning. The stale DEK is never 
cleared, so every DataNode in the block group fails with the same error. The 
operation fails permanently — even client side retries will reuse the same 
stale cached DEK.

Proposed Fix: Add catch (InvalidEncryptionKeyException) in 
checksumBlockGroup(), mirroring the existing handling in checksumBlock().

I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to 
reproduce it.

  was:
HDFS-12931 added handling for InvalidEncryptionKeyException in 
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
with encrypted data transfer. However, the parallel striped file path 
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.

 
Both paths call DFSClient.connectToDN(), which performs a SASL handshake using 
a cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has 
been removed from the DataNode (there is a time gap when the DataNode isn't 
updated with the new keys after key rotation, as described in HDFS-12931), the 
handshake fails with InvalidEncryptionKeyException.
 

In the replicated path, this exception is caught in HDFS-12931's fix, 
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block 
is retried. In the striped path, the exception falls through to the generic 
catch (IOException) block, which only logs a warning. The stale DEK is never 
cleared, so every DataNode in the block group fails with the same error. The 
operation fails permanently — even client side retries will reuse the same 
stale cached DEK.

Proposed Fix: Add catch (InvalidEncryptionKeyException) in 
checksumBlockGroup(), mirroring the existing handling in checksumBlock().

I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to 
reproduce it.


> Handle InvalidEncryptionKeyException during striped file checksum 
> ------------------------------------------------------------------
>
>                 Key: HDFS-17897
>                 URL: https://issues.apache.org/jira/browse/HDFS-17897
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: encryption
>    Affects Versions: 3.1.2
>            Reporter: ZhenyuLi
>            Priority: Major
>              Labels: pull-request-available
>
> HDFS-12931 added handling for InvalidEncryptionKeyException in 
> ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
> with encrypted data transfer. However, the parallel striped file path 
> StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
>  
> Both paths call {{{}DFSClient.connectToDN(){}}}, which performs a SASL 
> handshake using a cached {{DataEncryptionKey}} (DEK). After NameNode key 
> rotation, the client may still hold a cached DEK derived from an old 
> {{BlockKey}} that has already expired and been removed from DataNodes (as 
> described in HDFS-12931), causing the handshake to fail with 
> {{{}InvalidEncryptionKeyException{}}}.
>  
> In the replicated path, this exception is caught in HDFS-12931's fix, 
> clearDataEncryptionKey() is called to invalidate the cached DEK, and the 
> block is retried. In the striped path, the exception falls through to the 
> generic catch (IOException) block, which only logs a warning. The stale DEK 
> is never cleared, so every DataNode in the block group fails with the same 
> error. The operation fails permanently — even client side retries will reuse 
> the same stale cached DEK.
> Proposed Fix: Add catch (InvalidEncryptionKeyException) in 
> checksumBlockGroup(), mirroring the existing handling in checksumBlock().
> I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to 
> reproduce it.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to