shrivasshankar opened a new pull request, #25790:
URL: https://github.com/apache/datafusion/pull/25790

   ## Which issue does this PR close?
   
   - Closes #25613.
   
   ## Rationale for this change
   
   Decimal columns were excluded from interval analysis, so any filter touching
   one fell back to the default 20% selectivity. For the Q6-style predicate in 
the
   issue, that meant estimating ~1.2M rows when only ~114K match.
   
   ## What changes are included in this PR?
   
   - `FilterExec` statistics use a new `check_statistics_support`, which also
     accepts Decimal32/64/128/256.
   - `next_value` / `prev_value` handle decimals, so strict comparisons get 
tight
     bounds (`x < 24.00` becomes `x <= 23.99`). At the limit of a decimal's
     precision the bound becomes unbounded instead of overflowing.
   - `get_extreme_value!` supports Decimal32/64, which previously hit
     `unreachable!()`.
   - `CastExpr` now propagates constraints back through lossless decimal 
widening
     casts. Without this, a literal with a different scale than the column (e.g.
     `a >= 0.055` on `Decimal(15, 2)`, coerced to `Decimal128(30, 15)`) left the
     column range untouched and the estimate at 100%.
   
   `check_support` and `is_datatype_supported` are unchanged, so symmetric hash
   join pruning behaves as before. Decimal support there could be a follow-up.
   
   ## What is the testing strategy for this PR?
   
   - Unit tests for decimal next/prev values across all four widths, including
     precision boundaries like 9.99 / -9.99 for `Decimal(3, 2)`, plus strict
     comparison bounds and Decimal32/64 overflow handling.
   - `test_filter_statistics_decimal_expr` and
     `test_filter_statistics_decimal_mixed_scale_expr` check row estimates and
     tightened min/max for same-scale and mixed-scale predicates. The 
mixed-scale
     test mirrors the plan the SQL planner produces.
   - New cases in `test_cast_constraint_propagation`, plus an exhaustive check
     that propagating through a decimal widening cast never excludes a valid 
input.
   - `test_custom_filter_selectivity` relied on decimals being unsupported, so I
     switched it to `Utf8`.
   - No sqllogictest or TPC-H/TPC-DS plan outputs changed.
   
   Reproducer from the issue (TPC-H SF1, #25570 applied locally):
   
   | SELECT                    |    Before |   After |  Actual |
   | ------------------------- | --------: | ------: | ------: |
   | Date range                |   869,535 | 869,535 | 909,455 |
   | Date + Decimal conditions | 1,200,243 | 111,291 | 114,160 |
   
   Worth a closer look in review:
   - The cast change extends the allowlist added in #25531. Decimal widening is
     lossless and order-preserving, and casting bounds back rounds to a
     neighboring value, which only widens the range. That reasoning is what
     makes it safe, so I'd appreciate a second opinion on it.
   - Decimal arithmetic inside filters (e.g. `price * (1 - discount) > x`) now
     goes through interval analysis too. The existing tests pass, but I didn't 
add
     tests specifically for that.
   - A decimal column with no min/max stats now gets selectivity 1.0 instead of
     the default, which matches how integer columns already behave.
   
   ## Are there any user-facing changes?
   
   Better row estimates for filters on decimal columns, which may change some
   plan choices. One new public function, `check_statistics_support`; no 
breaking
   changes.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to