[
https://issues.apache.org/jira/browse/TIKA-2610?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16404643#comment-16404643
]
Aleksei Udalov commented on TIKA-2610:
--------------------------------------
Will be happy to contribute to the fix if the Team agrees it makes sense.
> Extend HtmlMapper isDiscardElement method with Attributes parameter
> -------------------------------------------------------------------
>
> Key: TIKA-2610
> URL: https://issues.apache.org/jira/browse/TIKA-2610
> Project: Tika
> Issue Type: Improvement
> Components: parser
> Affects Versions: 1.17
> Reporter: Aleksei Udalov
> Priority: Major
>
> Currently, if we want to discard HTML elements by attribute value/existence,
> an example from one of our projects
> {code:html}
> <div data-meta-no-index>Some content to be ignored by custom search indexer
> (Tika parser)</div>
> {code}
> it's required to implement a custom handler with logic very similar to what
> we have in org.apache.tika.parser.html.HtmlHandler. While it can be easily
> done by keep using HtmlHandler, but setting an instance of HtmlMapper with
> (newly added) isDiscardElement(String name, Attributes attributes) method
> overridden into the ParseContext.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)