adriangb commented on PR #24086: URL: https://github.com/apache/datafusion/pull/24086#issuecomment-5857231608
# Benchmark summary: `4a9f276` (speculative read-ahead + MemoryPool accounting) `read_ahead_bytes` = 100 MB on the branch side only. GKE runs: adriangbot `c4a-highmem-16`, compared with merge-base `e8e41ae`. "Speedup" is the ratio of the per-query time totals. ## With `SIMULATE_LATENCY`, `pushdown_filters=false` | suite | query total base → branch | speedup | faster / slower / same | peak memory base → branch | |---|---|---|---|---| | tpch_sf1 | 19.09 s → 10.94 s | **1.75x** | 21 / 0 / 1 | 873 MiB → 1.0 GiB | | tpch_sf10 | 119.52 s → 17.01 s | **7.03x** | 22 / 0 / 0 | 4.1 GiB → 4.0 GiB | | tpcds_sf1 | 65.10 s → 65.37 s | 1.00x | 9 / 12 / 78 | 1.0 GiB → 1.0 GiB | | clickbench_partitioned | 87.01 s → 44.39 s | **1.96x** | 36 / 1 / 6 | 16.2 GiB → 17.2 GiB | ## With `SIMULATE_LATENCY`, `pushdown_filters=true` | suite | query total base → branch | speedup | faster / slower / same | peak memory base → branch | |---|---|---|---|---| | tpch_sf1 | 41.47 s → 11.11 s | **3.73x** | 22 / 0 / 0 | 623 MiB → 748 MiB | | tpcds_sf1 | 124.94 s → 72.70 s | **1.72x** | 94 / 2 / 3 | 1.1 GiB → 1.1 GiB | | clickbench_partitioned | 108.67 s → 39.45 s | **2.75x** | 41 / 0 / 2 | 9.0 GiB → 10.3 GiB | | tpch_sf1, `DATAFUSION_RUNTIME_MEMORY_LIMIT=1G` | 42.28 s → 11.06 s | **3.82x** | 22 / 0 / 0 | 641 MiB → 697 MiB | ## History of the pushdown case | revision | read-ahead of conditional ranges | tpch_sf1 | tpcds_sf1 | |---|---|---|---| | `4896a0a` | skipped | 0.49x | 0.14x | | `3d7d334` (diagnostic) | skipped / fetched | 0.50x / 3.79x | 0.14x / 1.74x | | `4a9f276` (this) | always fetched | **3.73x** | **1.72x** | When conditional ranges are skipped, each window of `batch_size` rows waits for one round trip per predicate plus one. Row-group mode pays about two per row group. DuckDB (eager by default, lazy per row group only after filters prove selective), ClickHouse MergeTree on remote storage (prefetches PREWHERE and other columns together) and Arrow C++ `pre_buffer` also prefetch post-filter columns on object storage. ## Notes - The memory accounting has no measurable cost: the no-pushdown results match the previous revision (1.76x, 6.95x, 1.98x). - The 1 GB pool run has no failures. Its peak memory stayed under the limit, so it does not show how often read-ahead backs off. The unit test with a 64 KiB pool covers the backoff. - tpcds Q27 is slower in every latency run (1.31x on `32fdd03`, 1.55x here with pushdown: 141 → 218 ms). Q44 is 1.38x here. I will profile Q27 locally. - Without pushdown, tpcds has 12 queries slower by 8–17% and 9 faster, with a total of 1.00x. This looks like noise, but I will check it with the Q27 profile. Runs: [pushdown + latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458045), [latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458126), [1 GB pool](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458219). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
