HappenLee opened a new pull request, #68553:
URL: https://github.com/apache/doris/pull/68553
### What problem does this PR solve?
Issue Number: None
Related PR: #68446
Ordinary IN scans currently send the complete point-key range list to every
selected tablet. FE tablet pruning reduces the tablet set but each remaining
tablet still probes keys belonging to other buckets. For example, 55 keys
across 43 selected tablets produce 2,365 tablet/range associations.
This change groups final exact point ranges by the table's hash bucket
before creating scanners. FE attaches the full bucket count and ordinal from
the existing tablet map; BE uses the same raw-byte CRC32 as VARCHAR
distribution. Ordinary scanners and both parallel splitting strategies consume
bucket-local ranges, and empty groups skip scanner/segment loading. Residual
predicates and versioned read sources stay on the normal scan path.
Initial scope: unpartitioned base indexes with one non-null VARCHAR storage
key equal to the sole HASH distribution column. Other schemas, rollups, special
scans, and non-point/coalesced ranges retain the existing path.
`enable_scan_key_bucket_prune` defaults to true and provides a per-session A/B
switch. Optional Thrift metadata preserves mixed-version behavior. Two Profile
counters expose tablet/range associations before and after grouping.
### Performance validation
Same RELEASE binary, local single FE/BE on a shared 192-logical-CPU host; 1M
MOW rows, 128 buckets, eight insert batches, auto compaction disabled, 4KB
string payload. Each query selects address/chain/payload for an IN list and
`chain = 1`. Parallel scan enabled; short-circuit, SQL cache and query cache
disabled. Each cell has 50 samples; paired OFF/ON order is randomized and all
1,000 timed query results match.
`doris_cache_cold` clears the dedicated BE's Doris caches before each
sample. OS caches remain uncontrolled; this does not model cold object-storage
reads. Compare OFF/ON within a row, not across cache modes.
| Cache | IN keys | OFF p50 ms | ON p50 ms | p50 reduction | OFF p95 ms | ON
p95 ms |
|---|---:|---:|---:|---:|---:|---:|
| hot | 4 | 12.306 | 12.327 | -0.2% | 14.861 | 14.358 |
| hot | 9 | 11.956 | 11.738 | 1.8% | 15.550 | 14.778 |
| hot | 30 | 13.324 | 12.665 | 4.9% | 16.209 | 17.543 |
| hot | 55 | 16.221 | 13.691 | 15.6% | 17.982 | 14.849 |
| hot | 100 | 77.752 | 37.955 | 51.2% | 100.982 | 46.330 |
| doris_cache_cold | 4 | 8.263 | 8.257 | 0.1% | 9.886 | 9.912 |
| doris_cache_cold | 9 | 9.583 | 9.320 | 2.7% | 13.308 | 11.846 |
| doris_cache_cold | 30 | 13.207 | 11.947 | 9.5% | 15.255 | 13.696 |
| doris_cache_cold | 55 | 16.720 | 14.012 | 16.2% | 19.495 | 16.214 |
| doris_cache_cold | 100 | 94.023 | 44.950 | 52.2% | 225.355 | 108.331 |
For 55 keys, range associations drop from 2365 to 55 (97.7%); end-to-end p50
drops about 16%. Separately captured warm profiles show
`GenerateRowRangeByKeysTime` 13.560 → 1.653 ms and cumulative `ScannerCpuTime`
46.460 → 31.784 ms (single samples). Small lists show little gain; hot 30-key
p95 increased in this run, so this is not a universal tail-latency improvement.
<details>
<summary>Workload construction and timing controls</summary>
```sql
CREATE TABLE scan_key_bucket_perf (
address VARCHAR(80) NOT NULL, chain INT NOT NULL, payload STRING
) UNIQUE KEY(address) DISTRIBUTED BY HASH(address) BUCKETS 128
PROPERTIES("replication_num"="1", "enable_unique_key_merge_on_write"="true",
"disable_auto_compaction"="true");
```
Insert eight batches of 125,000 rows from `numbers("number"="125000")`. In
batch `b`, let `i = number + b * 125000`; address is `concat('address',
lpad(cast(i as string), 9, '0'))`, chain is `i % 3`, and payload is
`repeat(sha2(cast(i as string), 256), 64)`.
For list size `n`, keys are generated with Python
`random.Random(68446+n).sample(range(1000000), n)` and the same address
formatting. Query all three columns using the generated IN list and `chain =
1`. Set `max_scan_key_num=1024`, `max_pushdown_conditions_per_column=1024`,
`enable_parallel_scan=true`, `enable_short_circuit_query=false`,
`enable_sql_cache=false`, and `enable_query_cache=false`. Toggle
`enable_scan_key_bucket_prune` per trial. Hot mode uses five warm-ups per
setting; cold mode calls `/api/clear_cache/all` on the dedicated BE before each
sample. Fetch all rows, measure with `perf_counter_ns`, and compare sorted
result checksums. Profiles are collected separately with both paths warmed.
</details>
### Release note
Improve eligible VARCHAR IN scans by pruning foreign-bucket point ranges
before scanner creation. The `enable_scan_key_bucket_prune` session variable
controls the optimization.
### Check List (For Author)
- Test
- [x] Regression test: generated and replayed
`query_p0/test_scan_key_bucket_prune`; 35 result blocks covering ON/OFF,
parallel ON/OFF, range coalescing, residual predicates, DUP rows, MOW
updates/deletes, and unsupported nullable/composite/partitioned layouts. Eight
setting combinations agree; generated results also match independent Python
references.
- [x] Unit Test: 10 BE tests under ASAN
(`ScanKeyBucketPrunerTest.*:ScannerLateArrivalRfTest.*`) and 14 FE tests
(`OlapScanNodeTest`).
- [x] Manual test: RELEASE BE/FE builds, FE checkstyle, clang-format
16.0.6, build/header hygiene, and the A/B measurements above.
- Behavior changed: Yes, eligible scanners receive bucket-local point
ranges; SQL results and storage formats are unchanged.
- Does this need documentation: No, internal scan optimization; switch
description is included in session-variable metadata.
Static analysis: the repository clang-tidy script was run with commands
expanded from the unity compilation database. Its current result is blocked by
a pre-existing unmatched `NOLINTEND` in `be/src/core/types.h`; full details
will be updated when the run finishes. No diagnostics on changed lines have
been reported so far.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]