[
https://issues.apache.org/jira/browse/HDFS-17897?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ZhenyuLi updated HDFS-17897:
----------------------------
Description:
HDFS-12931 added handling for InvalidEncryptionKeyException in
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
with encrypted data transfer. However, the parallel striped file path
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
Both paths call {{{}DFSClient.connectToDN(){}}}, which performs a SASL
handshake using a cached {{DataEncryptionKey}} (DEK). After NameNode key
rotation, the client may still hold a cached DEK derived from an old
{{BlockKey}} that has already expired and been removed from DataNodes (as
described in HDFS-12931), causing the handshake to fail with
{{{}InvalidEncryptionKeyException{}}}.
In the replicated path, this exception is caught in HDFS-12931's fix,
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block
is retried. In the striped path, the exception falls through to the generic
catch (IOException) block, which only logs a warning. The stale DEK is never
cleared, so every DataNode in the block group fails with the same error. The
operation fails permanently — even client side retries will reuse the same
stale cached DEK.
Proposed Fix: Add catch (InvalidEncryptionKeyException) in
checksumBlockGroup(), mirroring the existing handling in checksumBlock().
I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to
reproduce it.
was:
HDFS-12931 added handling for InvalidEncryptionKeyException in
ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
with encrypted data transfer. However, the parallel striped file path
StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
Both paths call DFSClient.connectToDN(), which performs a SASL handshake using
a cached DataEncryptionKey (DEK). When the DEK references a BlockKey that has
been removed from the DataNode (there is a time gap when the DataNode isn't
updated with the new keys after key rotation, as described in HDFS-12931), the
handshake fails with InvalidEncryptionKeyException.
In the replicated path, this exception is caught in HDFS-12931's fix,
clearDataEncryptionKey() is called to invalidate the cached DEK, and the block
is retried. In the striped path, the exception falls through to the generic
catch (IOException) block, which only logs a warning. The stale DEK is never
cleared, so every DataNode in the block group fails with the same error. The
operation fails permanently — even client side retries will reuse the same
stale cached DEK.
Proposed Fix: Add catch (InvalidEncryptionKeyException) in
checksumBlockGroup(), mirroring the existing handling in checksumBlock().
I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to
reproduce it.
> Handle InvalidEncryptionKeyException during striped file checksum
> ------------------------------------------------------------------
>
> Key: HDFS-17897
> URL: https://issues.apache.org/jira/browse/HDFS-17897
> Project: Hadoop HDFS
> Issue Type: Bug
> Components: encryption
> Affects Versions: 3.1.2
> Reporter: ZhenyuLi
> Priority: Major
> Labels: pull-request-available
>
> HDFS-12931 added handling for InvalidEncryptionKeyException in
> ReplicatedFileChecksumComputer.checksumBlock() to support getFileChecksum()
> with encrypted data transfer. However, the parallel striped file path
> StripedFileNonStripedChecksumComputer.checksumBlockGroup() was not updated.
>
> Both paths call {{{}DFSClient.connectToDN(){}}}, which performs a SASL
> handshake using a cached {{DataEncryptionKey}} (DEK). After NameNode key
> rotation, the client may still hold a cached DEK derived from an old
> {{BlockKey}} that has already expired and been removed from DataNodes (as
> described in HDFS-12931), causing the handshake to fail with
> {{{}InvalidEncryptionKeyException{}}}.
>
> In the replicated path, this exception is caught in HDFS-12931's fix,
> clearDataEncryptionKey() is called to invalidate the cached DEK, and the
> block is retried. In the striped path, the exception falls through to the
> generic catch (IOException) block, which only logs a warning. The stale DEK
> is never cleared, so every DataNode in the block group fails with the same
> error. The operation fails permanently — even client side retries will reuse
> the same stale cached DEK.
> Proposed Fix: Add catch (InvalidEncryptionKeyException) in
> checksumBlockGroup(), mirroring the existing handling in checksumBlock().
> I wrote a unit test testStripedFileChecksumWithInvalidEncryptionKey to
> reproduce it.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]