github-actions[bot] commented on code in PR #68539:
URL: https://github.com/apache/doris/pull/68539#discussion_r4121001282
##########
be/src/format/file_reader/new_plain_text_line_reader.cpp:
##########
@@ -287,6 +288,74 @@ inline bool NewPlainTextLineReader::update_eof() {
return _eof;
}
+Status NewPlainTextLineReader::skip_split_prefix(size_t split_start, const
std::string& delimiter,
+ bool* eof, const
io::IOContext* io_ctx,
+ size_t* skipped_lines) {
+ DCHECK_EQ(_total_read_bytes, 0);
+ DCHECK_EQ(_output_buf_limit, 0);
+ bool overlaps = false;
+ for (size_t shift = 1; shift < delimiter.size(); ++shift) {
+ if (delimiter.compare(shift, delimiter.size() - shift, delimiter, 0,
+ delimiter.size() - shift) == 0) {
+ overlaps = true;
+ break;
+ }
+ }
+
+ if (overlaps && _decompressor == nullptr) {
+ // A fixed lookbehind can start in an overlapping delimiter chain
(e.g. three newlines
+ // with a two-newline delimiter). Find a byte that cannot belong to
any delimiter, then
+ // replay greedy matches from immediately after it. Keep scratch
bounded even for long
+ // delimiter runs; ordinary delimiters do not need this extra I/O.
+ std::array<bool, 256> delimiter_bytes {};
+ for (unsigned char byte : delimiter) {
+ delimiter_bytes[byte] = true;
+ }
+ constexpr size_t max_lookbehind_size = 64 * 1024;
+ std::vector<char> buffer(1024);
+ size_t sync_offset = _current_offset;
+ bool synchronized = false;
+ while (sync_offset > 0 && !synchronized) {
Review Comment:
[P1] Avoid replaying the whole prefix for each overlapping split. With
delimiter `\n\n` and a long run of `\n`, this loop cannot find a
synchronization byte until before the run (or file start). Every later CSV/JSON
split scans that prefix backward, then `skip_split_prefix()` rereads it
forward. Plain files are split into ranges by FE; a 10 GiB run with 64 MiB
ranges incurs roughly 1.6 TiB of redundant input work, so scan cost grows
quadratically with split count. Keep these files unsplit or reuse
delimiter-boundary metadata so each split does not rescan the whole prefix.
##########
be/src/format/file_reader/new_plain_text_line_reader.cpp:
##########
@@ -287,6 +288,74 @@ inline bool NewPlainTextLineReader::update_eof() {
return _eof;
}
+Status NewPlainTextLineReader::skip_split_prefix(size_t split_start, const
std::string& delimiter,
+ bool* eof, const
io::IOContext* io_ctx,
+ size_t* skipped_lines) {
+ DCHECK_EQ(_total_read_bytes, 0);
+ DCHECK_EQ(_output_buf_limit, 0);
+ bool overlaps = false;
+ for (size_t shift = 1; shift < delimiter.size(); ++shift) {
+ if (delimiter.compare(shift, delimiter.size() - shift, delimiter, 0,
Review Comment:
[P2] Make delimiter overlap detection linear or bound delimiter length. For
a user-supplied delimiter of N-1 `a` bytes followed by `b`, every suffix
comparison reaches its last byte before mismatching, so this loop examines
N(N-1)/2 bytes. A 100 KiB delimiter costs about 5 GiB of comparisons on every
non-first split before any file read; FE CSV/JSON properties impose no length
limit. This delimiter has no overlap, so the cost occurs even when the
backward-search branch is skipped.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]