[ 
https://issues.apache.org/jira/browse/HDFS-17897?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ZhenyuLi updated HDFS-17897:
----------------------------
    Description: 
HDFS-12931 added handling for InvalidEncryptionKeyException in 
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
with encrypted data transfer. However, the parallel striped file path 
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.

 
Both paths call DFSClient.connectToDN(), which performs a SASL handshake using 
a cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has 
been removed from the DataNode (there is a time gap when the DataNode isn't 
updated with the new keys after key rotation, as described in HDFS-12931), the 
handshake fails with InvalidEncryptionKeyException.
 

In the replicated path, this exception is caught in HDFS-12931's fix, 
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block 
is retried. In the striped path, the exception falls through to the generic 
catch (IOException) block, which only logs a warning. The stale DEK is never 
cleared, so every DataNode in the block group fails with the same error. The 
operation fails permanently — even client side retries will reuse the same 
stale cached DEK.

Proposed Fix: Add catch (InvalidEncryptionKeyException) in 
checksumBlockGroup(), mirroring the existing handling in checksumBlock().

  was:
HDFS-12931 added handling for InvalidEncryptionKeyException in 
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
with encrypted data transfer. However, the parallel striped file path 
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.

Both paths call DFSClient.connectToDN() which performs a SASL handshake using a 
cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has 
been removed from the DataNode (e.g., after
  NameNode restart or key rotation), the handshake fails with 
InvalidEncryptionKeyException.

  In the replicated path, this exception is caught, clearDataEncryptionKey() is 
called to invalidate the cached DEK, and the block is retried. In the striped 
path, the exception falls through to the
  generic catch (IOException) block, which only logs a warning. The stale DEK 
is never cleared, so every DataNode in the block group fails with the same 
error. The operation fails permanently — even
  user-level retries will reuse the same stale cached DEK.

  Fix: Add catch (InvalidEncryptionKeyException) in checksumBlockGroup(), 
mirroring the existing handling in checksumBlock().


> Handle InvalidEncryptionKeyException during striped file checksum 
> ------------------------------------------------------------------
>
>                 Key: HDFS-17897
>                 URL: https://issues.apache.org/jira/browse/HDFS-17897
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: encryption
>    Affects Versions: 3.1.2
>            Reporter: ZhenyuLi
>            Priority: Major
>
> HDFS-12931 added handling for InvalidEncryptionKeyException in 
> ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum() 
> with encrypted data transfer. However, the parallel striped file path 
> StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
>  
> Both paths call DFSClient.connectToDN(), which performs a SASL handshake 
> using a cached DataEncryptionKey (DEK). When the DEK references a BlockKey 
> that has been removed from the DataNode (there is a time gap when the 
> DataNode isn't updated with the new keys after key rotation, as described in 
> HDFS-12931), the handshake fails with InvalidEncryptionKeyException.
>  
> In the replicated path, this exception is caught in HDFS-12931's fix, 
> clearDataEncryptionKey() is called to invalidate the cached DEK, and the 
> block is retried. In the striped path, the exception falls through to the 
> generic catch (IOException) block, which only logs a warning. The stale DEK 
> is never cleared, so every DataNode in the block group fails with the same 
> error. The operation fails permanently — even client side retries will reuse 
> the same stale cached DEK.
> Proposed Fix: Add catch (InvalidEncryptionKeyException) in 
> checksumBlockGroup(), mirroring the existing handling in checksumBlock().



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to