kamthorn opened a new pull request, #16727:
URL: https://github.com/apache/lucene/pull/16727

   ### Description
   Fixes #10153 (LUCENE-9112).
   
   `SegmentingTokenizerBase` uses a hardcoded 1,024-character buffer 
(`BUFFERMAX = 1024`) and only recognizes newlines (`\r`, `\n`, `\u0085`, 
`\u2028`, `\u2029`) as unambiguous sentence break positions in `isSafeEnd()`.
   
   In Thai, spaces are used as clause/sentence boundaries rather than newlines, 
and long paragraphs frequently exceed 1,024 characters without a newline. When 
`findSafeEnd()` cannot find a newline in the 1,024-character buffer, 
`usableLength` is set to `length` (1024), cutting text abruptly at the buffer 
boundary. This causes Thai words spanning across the boundary (e.g. 
`มหาวิทยาลัย` spanning indices 1020–1031) to be truncated into invalid word 
fragments (`มหา` and `วิทยาลัย`).
   
   ### Solution
   1. **Configurable Buffer Size in `SegmentingTokenizerBase`**:
      - Added constructors accepting `int bufferSize` to 
`SegmentingTokenizerBase`.
      - Preserves `BUFFERMAX` (1024) as default for complete backward 
compatibility.
      - Allows users and subclasses to configure larger or smaller buffers as 
discussed in #10153.
   
   2. **Thai Safe Boundary Detection in `ThaiTokenizer`**:
      - Overrode `isSafeEnd(char ch)` in `ThaiTokenizer` to recognize 
whitespace (`Character.isWhitespace(ch)`).
      - In Thai orthography, whitespace is always an unambiguous word/clause 
boundary (words never contain whitespace).
      - When refilling, the tokenizer safely stops at the last whitespace 
before buffer exhaustion, preventing words from being split across buffer 
boundaries.
      - Exposed `bufferSize` in `ThaiTokenizer` constructors and 
`ThaiTokenizerFactory` (`bufferSize` argument).
   
   3. **Tests**:
      - Added `testCustomBufferSize()` in `TestSegmentingTokenizerBase` with a 
small buffer size (16).
      - Added `testLongTextAcrossBufferBoundary()` in `TestThaiTokenizer` 
verifying that words spanning across the default 1,024-char boundary remain 
intact.
      - Added `testCustomBufferSize()` and `testFactoryWithBufferSize()` in 
`TestThaiTokenizer`.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to