kamthorn opened a new issue, #16721:
URL: https://github.com/apache/lucene/issues/16721

   ### Description
   
   Currently, `ThaiTokenizer` relies exclusively on 
`java.text.BreakIterator.getWordInstance(Locale.forLanguageTag("th"))` which 
uses a hardcoded dictionary embedded in the JRE.
   
   #### Problem:
   Users cannot supply custom vocabularies, domain-specific terminology, 
technical jargon, proper nouns, or loanwords (e.g. "พารากอน", "คลาวด์เนทีฟ", 
"คนขับรถ", "เอไอ"). Consequently, these terms are frequently fragmented into 
unnatural sub-words or single characters, making exact search matching and 
phrase queries difficult or imprecise.
   
   By contrast, other Asian language analyzers in Lucene (such as Kuromoji for 
Japanese and Nori for Korean) provide user dictionary support.
   
   #### Proposed Solution:
   1. Update `ThaiTokenizer` to accept an optional `CharArraySet 
userDictionary`. When matching terms, entries in the user dictionary take 
precedence over default `BreakIterator` segmentation boundaries.
   2. Update `ThaiTokenizerFactory` to implement `ResourceLoaderAware`, 
accepting a `dictionary` parameter (e.g. `<tokenizer 
class="solr.ThaiTokenizerFactory" dictionary="custom_words.txt"/>`).
   3. Maintain full backwards compatibility when no dictionary is specified.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to