adriangb commented on PR #24086:
URL: https://github.com/apache/datafusion/pull/24086#issuecomment-5857231608

   # Benchmark summary: `4a9f276` (speculative read-ahead + MemoryPool 
accounting)
   
   `read_ahead_bytes` = 100 MB on the branch side only. GKE runs: adriangbot 
`c4a-highmem-16`, compared with merge-base `e8e41ae`. "Speedup" is the ratio of 
the per-query time totals.
   
   ## With `SIMULATE_LATENCY`, `pushdown_filters=false`
   
   | suite | query total base → branch | speedup | faster / slower / same | 
peak memory base → branch |
   |---|---|---|---|---|
   | tpch_sf1 | 19.09 s → 10.94 s | **1.75x** | 21 / 0 / 1 | 873 MiB → 1.0 GiB |
   | tpch_sf10 | 119.52 s → 17.01 s | **7.03x** | 22 / 0 / 0 | 4.1 GiB → 4.0 
GiB |
   | tpcds_sf1 | 65.10 s → 65.37 s | 1.00x | 9 / 12 / 78 | 1.0 GiB → 1.0 GiB |
   | clickbench_partitioned | 87.01 s → 44.39 s | **1.96x** | 36 / 1 / 6 | 16.2 
GiB → 17.2 GiB |
   
   ## With `SIMULATE_LATENCY`, `pushdown_filters=true`
   
   | suite | query total base → branch | speedup | faster / slower / same | 
peak memory base → branch |
   |---|---|---|---|---|
   | tpch_sf1 | 41.47 s → 11.11 s | **3.73x** | 22 / 0 / 0 | 623 MiB → 748 MiB |
   | tpcds_sf1 | 124.94 s → 72.70 s | **1.72x** | 94 / 2 / 3 | 1.1 GiB → 1.1 
GiB |
   | clickbench_partitioned | 108.67 s → 39.45 s | **2.75x** | 41 / 0 / 2 | 9.0 
GiB → 10.3 GiB |
   | tpch_sf1, `DATAFUSION_RUNTIME_MEMORY_LIMIT=1G` | 42.28 s → 11.06 s | 
**3.82x** | 22 / 0 / 0 | 641 MiB → 697 MiB |
   
   ## History of the pushdown case
   
   | revision | read-ahead of conditional ranges | tpch_sf1 | tpcds_sf1 |
   |---|---|---|---|
   | `4896a0a` | skipped | 0.49x | 0.14x |
   | `3d7d334` (diagnostic) | skipped / fetched | 0.50x / 3.79x | 0.14x / 1.74x 
|
   | `4a9f276` (this) | always fetched | **3.73x** | **1.72x** |
   
   When conditional ranges are skipped, each window of `batch_size` rows waits 
for one round trip per predicate plus one. Row-group mode pays about two per 
row group. DuckDB (eager by default, lazy per row group only after filters 
prove selective), ClickHouse MergeTree on remote storage (prefetches PREWHERE 
and other columns together) and Arrow C++ `pre_buffer` also prefetch 
post-filter columns on object storage.
   
   ## Notes
   
   - The memory accounting has no measurable cost: the no-pushdown results 
match the previous revision (1.76x, 6.95x, 1.98x).
   - The 1 GB pool run has no failures. Its peak memory stayed under the limit, 
so it does not show how often read-ahead backs off. The unit test with a 64 KiB 
pool covers the backoff.
   - tpcds Q27 is slower in every latency run (1.31x on `32fdd03`, 1.55x here 
with pushdown: 141 → 218 ms). Q44 is 1.38x here. I will profile Q27 locally.
   - Without pushdown, tpcds has 12 queries slower by 8–17% and 9 faster, with 
a total of 1.00x. This looks like noise, but I will check it with the Q27 
profile.
   
   Runs: [pushdown + 
latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458045),
 
[latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458126),
 [1 GB 
pool](https://github.com/apache/datafusion/pull/24086#issuecomment-5852458219).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to