adriangb commented on PR #24086: URL: https://github.com/apache/datafusion/pull/24086#issuecomment-5848039299
# Benchmark summary: `32fdd03` (push decoder with `FetchGranularity::Batch`) Read-ahead window 100 MB. GKE runs: adriangbot `c4a-highmem-16`, compared with merge-base `e8e41ae`. "Query total" is the sum of per-query times from the bot. These runs use the `DF_FETCH_*` variables, which `32fdd03` reads. The current head reads `datafusion.execution.parquet.read_ahead_bytes` instead. ## With `SIMULATE_LATENCY` | suite | query total base → branch | speedup | faster / slower / same | wall base → branch | peak memory base → branch | previous revision (`16c5ba3`, sync reader) | |---|---|---|---|---|---|---| | tpch_sf1 | 19.22 s → 10.90 s | **1.76x** | 21 / 0 / 1 | 110 s → 65 s | 856 MiB → 944 MiB | 1.73x | | tpch_sf10 | 118.78 s → 17.08 s | **6.95x** | 22 / 0 / 0 | 630 s → 95 s | 4.0 GiB → 4.4 GiB | 7.04x | | tpcds_sf1 | 65.44 s → 65.50 s | 1.00x | 3 / 5 / 91 | 370 s → 370 s | 1.1 GiB → 1.1 GiB | 1.01x | | clickbench_partitioned | 86.25 s → 43.56 s | **1.98x** | 38 / 0 / 5 | 450 s → 230 s | 16.4 GiB → 16.5 GiB | 1.96x | Slower queries: tpcds Q27 1.31x (142 → 186 ms), Q54 1.14x, Q74 1.11x, Q67 1.09x, Q99 1.07x. ## Without latency | suite | query total base → branch | change | faster / slower / same | peak memory base → branch | previous revision | |---|---|---|---|---|---| | tpch_sf1 | 0.74 s → 0.74 s | 1.00x | 0 / 0 / 22 | 1.2 GiB → 1.3 GiB | 1.01x | | tpch_sf10 | 6.31 s → 6.53 s | 0.97x | 1 / 6 / 15 | 5.0 GiB → 4.4 GiB | not run | | tpcds_sf1 | 8.79 s → 8.80 s | 1.00x | 0 / 1 / 98 | 1.7 GiB → 2.1 GiB | 0.98x | | clickbench_partitioned | 20.10 s → 20.04 s | 1.00x | 1 / 2 / 40 | 18.1 GiB → 18.6 GiB | 0.97x | Slower queries: tpch_sf10 Q5, Q9, Q10, Q14, Q19, Q20 (1.05x–1.08x). tpcds Q28 1.10x. clickbench Q36 1.06x, Q39 1.07x. ## Notes - The speedups with latency did not change when the sync reader was replaced by the push decoder. - Without latency, ClickBench went from 0.97x (6 slower queries, +11% peak memory) to 1.00x (2 slower queries, +3% peak memory). - tpch_sf10 without latency is 0.97x. This suite was not run on the previous revision. To find out if this is noise or decode overhead, an A/A control and a rerun are requested below. - These suites run with `pushdown_filters=false`, so they do not use the new streaming path for filtered scans. A run with `pushdown_filters=true` on both sides is requested below. Runs with latency: [tpch](https://github.com/apache/datafusion/pull/24086#issuecomment-5841259150), [tpch10](https://github.com/apache/datafusion/pull/24086#issuecomment-5841330120), [tpcds](https://github.com/apache/datafusion/pull/24086#issuecomment-5841330578), [clickbench](https://github.com/apache/datafusion/pull/24086#issuecomment-5841335335). Without: [tpch](https://github.com/apache/datafusion/pull/24086#issuecomment-5841241056), [tpch10](https://github.com/apache/datafusion/pull/24086#issuecomment-5841304023), [tpcds](https://github.com/apache/datafusion/pull/24086#issuecomment-5841250467), [clickbench](https://github.com/apache/datafusion/pull/24086#issuecomment-5841275003). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
