[
https://issues.apache.org/jira/browse/HDFS-17897?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ZhenyuLi updated HDFS-17897:
----------------------------
Description:
HDFS-12931 added handling for InvalidEncryptionKeyException in
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
with encrypted data transfer. However, the parallel striped file path
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
Both paths call DFSClient.connectToDN(), which performs a SASL handshake using
a cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has
been removed from the DataNode (there is a time gap when the DataNode isn't
updated with the new keys after key rotation, as described in HDFS-12931), the
handshake fails with InvalidEncryptionKeyException.
In the replicated path, this exception is caught in HDFS-12931's fix,
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block
is retried. In the striped path, the exception falls through to the generic
catch (IOException) block, which only logs a warning. The stale DEK is never
cleared, so every DataNode in the block group fails with the same error. The
operation fails permanently — even client side retries will reuse the same
stale cached DEK.
Proposed Fix: Add catch (InvalidEncryptionKeyException) in
checksumBlockGroup(), mirroring the existing handling in checksumBlock().
was:
HDFS-12931 added handling for InvalidEncryptionKeyException in
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
with encrypted data transfer. However, the parallel striped file path
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
Both paths call DFSClient.connectToDN() which performs a SASL handshake using a
cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has
been removed from the DataNode (e.g., after
NameNode restart or key rotation), the handshake fails with
InvalidEncryptionKeyException.
In the replicated path, this exception is caught, clearDataEncryptionKey() is
called to invalidate the cached DEK, and the block is retried. In the striped
path, the exception falls through to the
generic catch (IOException) block, which only logs a warning. The stale DEK
is never cleared, so every DataNode in the block group fails with the same
error. The operation fails permanently — even
user-level retries will reuse the same stale cached DEK.
Fix: Add catch (InvalidEncryptionKeyException) in checksumBlockGroup(),
mirroring the existing handling in checksumBlock().
> Handle InvalidEncryptionKeyException during striped file checksum
> ------------------------------------------------------------------
>
> Key: HDFS-17897
> URL: https://issues.apache.org/jira/browse/HDFS-17897
> Project: Hadoop HDFS
> Issue Type: Bug
> Components: encryption
> Affects Versions: 3.1.2
> Reporter: ZhenyuLi
> Priority: Major
>
> HDFS-12931 added handling for InvalidEncryptionKeyException in
> ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
> with encrypted data transfer. However, the parallel striped file path
> StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
>
> Both paths call DFSClient.connectToDN(), which performs a SASL handshake
> using a cached DataEncryptionKey (DEK). When the DEK references a BlockKey
> that has been removed from the DataNode (there is a time gap when the
> DataNode isn't updated with the new keys after key rotation, as described in
> HDFS-12931), the handshake fails with InvalidEncryptionKeyException.
>
> In the replicated path, this exception is caught in HDFS-12931's fix,
> clearDataEncryptionKey() is called to invalidate the cached DEK, and the
> block is retried. In the striped path, the exception falls through to the
> generic catch (IOException) block, which only logs a warning. The stale DEK
> is never cleared, so every DataNode in the block group fails with the same
> error. The operation fails permanently — even client side retries will reuse
> the same stale cached DEK.
> Proposed Fix: Add catch (InvalidEncryptionKeyException) in
> checksumBlockGroup(), mirroring the existing handling in checksumBlock().
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]