[
https://issues.apache.org/jira/browse/HDFS-17897?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18067826#comment-18067826
]
ASF GitHub Bot commented on HDFS-17897:
---------------------------------------
JHSUYU commented on PR #8364:
URL: https://github.com/apache/hadoop/pull/8364#issuecomment-4113540450
Hi @jojochuang, I noticed that you previously created an issue on JIRA
(https://issues.apache.org/jira/browse/HDFS-16136), but nobody took it. I
observed that this issue still occurs in other features, as shown in this PR.
The HDFS fault tolerance mechanisms could not even mask the data encryption key
error in these cases, causing applications to fail. I am very willing to fix
all such issues.
> Handle InvalidEncryptionKeyException during striped file checksum
> ------------------------------------------------------------------
>
> Key: HDFS-17897
> URL: https://issues.apache.org/jira/browse/HDFS-17897
> Project: Hadoop HDFS
> Issue Type: Bug
> Components: encryption
> Affects Versions: 3.1.2, 3.2.3, 3.3.6, 3.4.3
> Reporter: ZhenyuLi
> Assignee: ZhenyuLi
> Priority: Major
> Labels: pull-request-available
>
> *Description:*
> HDFS-12931 added handling for {{InvalidEncryptionKeyException}} in
> {{ReplicatedFileChecksumComputer.checksumBlock()}} to support
> {{getFileChecksum()}} with encrypted data transfer. However, the parallel
> striped file path
> {{StripedFileNonStripedChecksumComputer.checksumBlockGroup()}} was not
> updated.
> Both paths call {{{}DFSClient.connectToDN(){}}}, which performs a SASL
> handshake using a cached {{DataEncryptionKey}} (DEK). After NameNode key
> rotation, the client may still hold a cached DEK derived from an old
> {{BlockKey}} that has already expired and been removed from DataNodes (as
> described in HDFS-12931), causing the handshake to fail with
> {{{}InvalidEncryptionKeyException{}}}.
> In the replicated path, this exception is caught by the handling added in
> HDFS-12931: {{clearDataEncryptionKey()}} is called to invalidate the cached
> DEK, and the block operation is retried with a freshly negotiated key. In the
> striped path, the exception falls through to the generic {{catch
> (IOException)}} block, which only logs a warning and moves on to the next
> DataNode in the block group. Because the stale DEK is never cleared, every
> subsequent {{connectToDN()}} call within the same block group fails with the
> same error. The operation fails permanently — even client-side retries reuse
> the same stale cached DEK.
> *Proposed Fix:* Add {{catch (InvalidEncryptionKeyException)}} in
> {{{}checksumBlockGroup(){}}}, mirroring the existing handling in
> {{{}checksumBlock(){}}}: call {{clearDataEncryptionKey()}} to invalidate the
> cached DEK, then retry the block group.
> *Testing:* Added unit test
> {{testStripedFileChecksumWithInvalidEncryptionKey}} to reproduce the failure
> and verify the fix.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]