This is an automated email from the ASF dual-hosted git repository. davsclaus pushed a commit to branch quick-fix/tika-text-decode-javadoc in repository https://gitbox.apache.org/repos/asf/camel.git
commit 6faa18685bab9fe831f1a011f50520aa4341a3da Author: Claus Ibsen <[email protected]> AuthorDate: Mon Sep 28 10:01:02 2026 +0200 chore: camel-langchain4j-ingest - TikaTextDecode javadoc no longer says convertBodyTo is outranked by the charset header CAMEL-25034 made convertBodyTo with a charset also override the CamelCharsetName header while converting, so the javadoc and test comment were out of date. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Signed-off-by: Claus Ibsen <[email protected]> --- .../apache/camel/component/langchain4j/ingest/TikaTextDecode.java | 6 +++--- .../camel/component/langchain4j/ingest/TikaTextDecodeTest.java | 3 +-- 2 files changed, 4 insertions(+), 5 deletions(-) diff --git a/components/camel-ai/camel-langchain4j-ingest/src/main/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecode.java b/components/camel-ai/camel-langchain4j-ingest/src/main/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecode.java index d692999af4ca..557473bf73ab 100644 --- a/components/camel-ai/camel-langchain4j-ingest/src/main/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecode.java +++ b/components/camel-ai/camel-langchain4j-ingest/src/main/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecode.java @@ -27,9 +27,9 @@ import org.apache.camel.Processor; * {@code tikaParseOutputEncoding}. This step exists because a plain {@code getBody(String.class)} resolves its charset * through the exchange: the {@code CamelCharsetName} <em>header</em> first, then the exchange property — so any route * that lets a message-supplied {@code CamelCharsetName} header survive up to the conversion hands the decode to whoever - * sent the message. Reading the raw bytes and decoding as the pinned charset never consults that heuristic. Note that - * {@code convertBodyTo(String.class, "UTF-8")} would not be a safe substitute: the option sets the exchange - * <em>property</em>, which the injected header outranks. + * sent the message. Reading the raw bytes and decoding as the pinned charset never consults that heuristic. Since + * CAMEL-25034, {@code convertBodyTo(String.class, "UTF-8")} also overrides that header while it converts, but this step + * keeps the decode self-contained and does not depend on how the route converts the body. * * <p> * camel-tika itself no longer forwards Camel-namespace metadata names from the parsed document (CAMEL-24423), so the diff --git a/components/camel-ai/camel-langchain4j-ingest/src/test/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecodeTest.java b/components/camel-ai/camel-langchain4j-ingest/src/test/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecodeTest.java index 5213fac9d15f..c89cdd0b9821 100644 --- a/components/camel-ai/camel-langchain4j-ingest/src/test/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecodeTest.java +++ b/components/camel-ai/camel-langchain4j-ingest/src/test/java/org/apache/camel/component/langchain4j/ingest/TikaTextDecodeTest.java @@ -34,8 +34,7 @@ class TikaTextDecodeTest { * The reason this step exists: getBody(String) resolves its charset through the exchange — the CamelCharsetName * header first, then the property — so a message-supplied CamelCharsetName header would steer the decode and mangle * the extracted text. The decode reads raw bytes and applies the pinned UTF-8 regardless. The header (not the - * property) is injected here on purpose: convertBodyTo(String, "UTF-8") only sets the property, which the header - * outranks — proving that option would be no substitute. + * property) is injected here on purpose, as the header outranks the property when a body is converted. */ @Test void decodesThePinnedUtf8DespiteAnInjectedCharsetHeader() throws Exception {
