kamthorn opened a new issue, #16721:
URL: https://github.com/apache/lucene/issues/16721
### Description
Currently, `ThaiTokenizer` relies exclusively on
`java.text.BreakIterator.getWordInstance(Locale.forLanguageTag("th"))` which
uses a hardcoded dictionary embedded in the JRE.
#### Problem:
Users cannot supply custom vocabularies, domain-specific terminology,
technical jargon, proper nouns, or loanwords (e.g. "พารากอน", "คลาวด์เนทีฟ",
"คนขับรถ", "เอไอ"). Consequently, these terms are frequently fragmented into
unnatural sub-words or single characters, making exact search matching and
phrase queries difficult or imprecise.
By contrast, other Asian language analyzers in Lucene (such as Kuromoji for
Japanese and Nori for Korean) provide user dictionary support.
#### Proposed Solution:
1. Update `ThaiTokenizer` to accept an optional `CharArraySet
userDictionary`. When matching terms, entries in the user dictionary take
precedence over default `BreakIterator` segmentation boundaries.
2. Update `ThaiTokenizerFactory` to implement `ResourceLoaderAware`,
accepting a `dictionary` parameter (e.g. `<tokenizer
class="solr.ThaiTokenizerFactory" dictionary="custom_words.txt"/>`).
3. Maintain full backwards compatibility when no dictionary is specified.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]