kamthorn opened a new pull request, #16727:
URL: https://github.com/apache/lucene/pull/16727
### Description
Fixes #10153 (LUCENE-9112).
`SegmentingTokenizerBase` uses a hardcoded 1,024-character buffer
(`BUFFERMAX = 1024`) and only recognizes newlines (`\r`, `\n`, `\u0085`,
`\u2028`, `\u2029`) as unambiguous sentence break positions in `isSafeEnd()`.
In Thai, spaces are used as clause/sentence boundaries rather than newlines,
and long paragraphs frequently exceed 1,024 characters without a newline. When
`findSafeEnd()` cannot find a newline in the 1,024-character buffer,
`usableLength` is set to `length` (1024), cutting text abruptly at the buffer
boundary. This causes Thai words spanning across the boundary (e.g.
`มหาวิทยาลัย` spanning indices 1020–1031) to be truncated into invalid word
fragments (`มหา` and `วิทยาลัย`).
### Solution
1. **Configurable Buffer Size in `SegmentingTokenizerBase`**:
- Added constructors accepting `int bufferSize` to
`SegmentingTokenizerBase`.
- Preserves `BUFFERMAX` (1024) as default for complete backward
compatibility.
- Allows users and subclasses to configure larger or smaller buffers as
discussed in #10153.
2. **Thai Safe Boundary Detection in `ThaiTokenizer`**:
- Overrode `isSafeEnd(char ch)` in `ThaiTokenizer` to recognize
whitespace (`Character.isWhitespace(ch)`).
- In Thai orthography, whitespace is always an unambiguous word/clause
boundary (words never contain whitespace).
- When refilling, the tokenizer safely stops at the last whitespace
before buffer exhaustion, preventing words from being split across buffer
boundaries.
- Exposed `bufferSize` in `ThaiTokenizer` constructors and
`ThaiTokenizerFactory` (`bufferSize` argument).
3. **Tests**:
- Added `testCustomBufferSize()` in `TestSegmentingTokenizerBase` with a
small buffer size (16).
- Added `testLongTextAcrossBufferBoundary()` in `TestThaiTokenizer`
verifying that words spanning across the default 1,024-char boundary remain
intact.
- Added `testCustomBufferSize()` and `testFactoryWithBufferSize()` in
`TestThaiTokenizer`.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]