This is an automated email from the ASF dual-hosted git repository.
davsclaus pushed a commit to branch main
in repository https://gitbox.apache.org/repos/asf/camel.git
The following commit(s) were added to refs/heads/main by this push:
new 88f829ac4bc3 camel-langchain4j-ingest - Document the Tika parser
module requirement (#26905)
88f829ac4bc3 is described below
commit 88f829ac4bc3778997335d26c1dce35b7967e541
Author: Jiří Ondrušek <[email protected]>
AuthorDate: Mon Sep 28 10:03:44 2026 +0200
camel-langchain4j-ingest - Document the Tika parser module requirement
(#26905)
Co-Authored-By: Claude Opus 5.5 <[email protected]>
---
.../org/apache/camel/catalog/docs/langchain4j-ingest-component.adoc | 4 ++++
.../src/main/docs/langchain4j-ingest-component.adoc | 4 ++++
2 files changed, 8 insertions(+)
diff --git
a/catalog/camel-catalog/src/generated/resources/org/apache/camel/catalog/docs/langchain4j-ingest-component.adoc
b/catalog/camel-catalog/src/generated/resources/org/apache/camel/catalog/docs/langchain4j-ingest-component.adoc
index b56060446aa9..c0a73d526ec0 100644
---
a/catalog/camel-catalog/src/generated/resources/org/apache/camel/catalog/docs/langchain4j-ingest-component.adoc
+++
b/catalog/camel-catalog/src/generated/resources/org/apache/camel/catalog/docs/langchain4j-ingest-component.adoc
@@ -171,6 +171,10 @@ from("file:/var/data/reports?noop=true&readLock=changed")
.to("langchain4j-ingest:reports");
----
+Each format needs its Tika parser module on the classpath: camel-tika ships
the HTML and text
+modules only, so a PDF needs `org.apache.tika:tika-parser-pdf-module`. A
format without its
+module parses to nothing and the document is answered `EMPTY`.
+
`TikaTextDecode` ships with this component for exactly this route shape: it
decodes Tika's
text output as the pinned UTF-8 without consulting the exchange — a plain
`getBody(String.class)` resolves its charset from the `CamelCharsetName`
*header* first (then
diff --git
a/components/camel-ai/camel-langchain4j-ingest/src/main/docs/langchain4j-ingest-component.adoc
b/components/camel-ai/camel-langchain4j-ingest/src/main/docs/langchain4j-ingest-component.adoc
index b56060446aa9..c0a73d526ec0 100644
---
a/components/camel-ai/camel-langchain4j-ingest/src/main/docs/langchain4j-ingest-component.adoc
+++
b/components/camel-ai/camel-langchain4j-ingest/src/main/docs/langchain4j-ingest-component.adoc
@@ -171,6 +171,10 @@ from("file:/var/data/reports?noop=true&readLock=changed")
.to("langchain4j-ingest:reports");
----
+Each format needs its Tika parser module on the classpath: camel-tika ships
the HTML and text
+modules only, so a PDF needs `org.apache.tika:tika-parser-pdf-module`. A
format without its
+module parses to nothing and the document is answered `EMPTY`.
+
`TikaTextDecode` ships with this component for exactly this route shape: it
decodes Tika's
text output as the pinned UTF-8 without consulting the exchange — a plain
`getBody(String.class)` resolves its charset from the `CamelCharsetName`
*header* first (then