raghav-reglobe opened a new pull request, #68527:
URL: https://github.com/apache/doris/pull/68527
### What problem does this PR solve?
Issue Number: close #63887
Related PR: #63889
Problem Summary:
A selective predicate plus a struct sub-field in the projection on a Parquet
table under lazy materialization crashes the BE with `SIGSEGV address not
mapped to object (@0x0)` in `ScalarColumnReader::gen_filter_map` (the stack in
#63887). We hit it on an Iceberg table with a text primary key: `SELECT id,
s['a'], s['b'] FROM t WHERE id IN ('k1','k2','k3','k4')` took down every BE the
query touched, and the FE's automatic retries took down the next ones.
Predicate-free scans of the same struct column are fine.
How the null map is produced, in `RowGroupReader::_do_lazy_read`:
1. a predicate batch whose rows all fail the predicate is cached as
filter-all (`_cached_filtered_rows`) and the loop continues;
2. when the row group's remaining pages are all pruned by the page index,
the next predicate read returns 0 rows with eof, and the loop breaks **before**
`filter_map_ptr` is reassigned, so it still points at the previous filter-all
map;
3. `_rebuild_filter_map` re-initializes that map as `init(nullptr,
cached_rows, true)` and the lazy columns are read with it;
4. a scalar lazy column asks `FilterMap::can_filter_all()` first and skips;
a nested lazy column projects the map onto its repetition levels
(`gen_filter_map`) and indexes the null data pointer.
An `IN` list on a text key gives the page index exactly this layout (a
candidate page whose rows all miss, followed by pruned pages to the end of the
row group), which is why the crash is deterministic for that query shape.
`init(nullptr, n, true)` has a single producer, `_rebuild_filter_map`, and it
is only reached through the 0-row break, so this is the only path to the crash.
### Design
Two changes, one per layer:
- `_do_lazy_read`: a 0-row predicate read with cached filter-all rows ends
the batch as eof and accounts the cached rows as lazy-read-filtered. There is
nothing to align the lazy columns against, so no lazy read happens at all, the
same way the existing filter-all-at-eof branch already returns. This removes
the producer of the null map.
- `gen_filter_map`: a filter-all map, or any map without data, projects to
an all-filtered nested map without reading the parent, the decision the scalar
path takes through `can_filter_all()`. #63889 adds this guard at the call site;
here it sits inside the projection, which is now the public static
`ScalarColumnReader::gen_nested_filter_map` so it can be unit-tested and so any
caller is covered.
Either change alone stops the crash; the first is the root cause, the second
is the safety net. If the maintainers prefer to land #63889 first, this PR
rebases onto it cleanly and keeps only the `_do_lazy_read` change and the tests.
### Release note
Fix a BE crash (SIGSEGV in `ScalarColumnReader::gen_filter_map`) when a
lazy-materialized Parquet scan combines a selective predicate with a nested
(struct/array/map) column and the row group's trailing pages are pruned by the
page index.
### Check List (For Author)
- Test
- [x] Unit Test: `parquet_nested_filter_map_test.cpp` (data-less
filter-all map, all-zero filter-all map, a level window, the selective
projection, a window starting at a later row) and
`FilterMapTest.test_filter_all_without_data` (the contract of `init(nullptr, n,
true)`).
- [x] Manual test: the crashing query above, on an FE + BE built from
this branch and on a production cluster running this fix on top of 4.1: rows
returned, both BEs alive, no crash marker in `be.out`; the same query on the
unpatched image reproduces the crash every time.
- Behavior changed: No.
- Does this need documentation: No.
### Check List (For Reviewer who merge this PR)
- [ ] Confirm the release note
- [ ] Confirm test cases
- [ ] Confirm document
- [ ] Add branch pick label
🤖 Generated with [Claude Code](https://claude.com/claude-code)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]