Messages by Thread
-
[I] Add integer boundary coverage to in-memory cache pruning tests [datafusion-comet]
via GitHub
-
[PR] perf(parquet): skip Bloom reads for fully matched row groups [datafusion]
via GitHub
-
[PR] fix: size the native runtime for standalone executors without spark.executor.cores [datafusion-comet]
via GitHub
-
[PR] fix: Preserve column identifiers in DataFrame fill methods [datafusion]
via GitHub
-
[I] INSERT nullability: unpartitioned path skips the runtime NULL check; partitioned path rejects nullable sources at plan time [datafusion-iceberg]
via GitHub
-
[PR] fix: build release jars from the same commit as the native libraries [datafusion-comet]
via GitHub
-
[PR] fix: make benchmark cases over ShortType columns exercise Comet [datafusion-comet]
via GitHub
-
[I] date_bin: define the semantics of negative strides [datafusion]
via GitHub
-
[I] date_bin: month strides on s/ms/us timestamps return NULL outside the chrono DateTime range [datafusion]
via GitHub
-
[PR] docs: add a Celeborn integration guide [datafusion-comet]
via GitHub
-
[PR] fix: derive the native memory limit the way YARN and Kubernetes size executor containers [datafusion-comet]
via GitHub
-
[PR] fix: warn about unreleased native memory only after a task's last native plan closes [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: divide_dt_interval [datafusion-comet]
via GitHub
-
Re: [PR] feat: support divide_dt_interval with codegen dispatch [datafusion-comet]
via GitHub
-
[PR] fix: release input Arrow streams that native never takes [datafusion-comet]
via GitHub
-
Re: [PR] fix(parquet): cache deferred page index loads [datafusion]
via GitHub
-
Re: [PR] feat(aggregate): cost-aware partial-aggregation skip (opt-in) [datafusion]
via GitHub
-
[PR] fix: serialize large-offset string and binary vectors with 32-bit offsets [datafusion-comet]
via GitHub
-
Re: [I] [DISCUSS] Is Comet ready to move to top-level Apache Comet project? [datafusion-comet]
via GitHub
-
[PR] fix: count hash-based JVM shuffle batches as shuffle bytes written, not spill [datafusion-comet]
via GitHub
-
[PR] fix: cap spark.comet.shuffle.jvm.batchSize at spark.comet.batchSize where it is read [datafusion-comet]
via GitHub
-
[PR] perf: compare bucket group indexes [datafusion]
via GitHub
-
Re: [PR] perf: stop redoing per-chunk setup on the hash join probe path [datafusion]
via GitHub
-
[PR] fix: check accelerated mapInArrow output against the declared schema [datafusion-comet]
via GitHub
-
Re: [PR] perf: cut schema check and subquery rescan overhead in optimizer [datafusion]
via GitHub
-
Re: [PR] perf: skip functional dependency work when inputs carry none [datafusion]
via GitHub
-
[PR] chore: remove AlignedArrowStreamReader and fix the Native to JVM FFI docs [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Spilling coverage gaps vs Spark and DuckDB: window, cross join, and pool-driven reclaim [datafusion]
via GitHub
-
[I] Native aggregate fails after spilling when groups are few and large (collect_list, collect_set) [datafusion-comet]
via GitHub
-
[I] Hash aggregate with a few large groups fails after spilling: spill batches are sized by rows, not bytes [datafusion]
via GitHub
-
[PR] fix: fail a native plan whose producer task is cancelled before its stream ends [datafusion-comet]
via GitHub
-
[PR] fix: support JVM columnar shuffle with spark.shuffle.checksum.enabled=false [datafusion-comet]
via GitHub
-
Re: [I] feat: support Spark HyperLogLog sketch functions (hll_sketch_agg, hll_union_agg, hll_sketch_estimate, hll_union) [datafusion-comet]
via GitHub
-
Re: [PR] fix: bump DataFusion to 54.1.0 and adapt to RecursiveQuery API change [datafusion-python]
via GitHub
-
Re: [PR] fix: fall back for ANSI and TRY integer sums over sliding windows [datafusion-comet]
via GitHub
-
[PR] fix: install Comet's cache serializer only when Comet and native execution are enabled [datafusion-comet]
via GitHub
-
[I] TopKAggregation drops the NULL group under ORDER BY <group key> LIMIT n [datafusion]
via GitHub
-
[PR] chore(deps): align with current iceberg-rust revision [datafusion-iceberg]
via GitHub
-
Re: [PR] test: cover matching nested nulls in collect aggregates [datafusion-comet]
via GitHub
-
[PR] docs: [branch-1.1] credit co-authors in the 1.1.0 changelog [datafusion-comet]
via GitHub
-
[PR] fix: credit co-authors of a PR in the generated changelog [datafusion-comet]
via GitHub
-
[PR] docs: add a backporting policy to the contributor guide [datafusion-comet]
via GitHub
-
Re: [I] `FFI_ExecutionPlan::new` silently discards transparent wrapper nodes [datafusion]
via GitHub
-
Re: [I] docs: add a user-facing Celeborn integration guide [datafusion-comet]
via GitHub
-
[PR] fix: make Scala logging consistent and log exceptions with their stack traces [datafusion-comet]
via GitHub
-
[PR] feat: optimize count over non-null arguments [datafusion]
via GitHub
-
[PR] test: consolidate partition metrics tests and clarify snapshot docs [datafusion]
via GitHub
-
[PR] feat: native Parquet writes to S3 [datafusion-comet]
via GitHub
-
[PR] fix: truncate timestamps around DST transitions the way Spark does [datafusion-comet]
via GitHub
-
Re: [I] Add efficient partition-specific metrics access to ExecutionPlan [datafusion]
via GitHub
-
[PR] fix: warn when native and JVM timezone databases differ [datafusion-comet]
via GitHub
-
[I] Optimize COUNT over provably non-null arguments as a nullary aggregate [datafusion]
via GitHub
-
Re: [PR] Rewrite `AVG(expr)` --> `SUM(expr) / COUNT(expr)` when components can be shared [datafusion]
via GitHub
-
[PR] docs: explain operator plan identity and exchange reuse [datafusion-comet]
via GitHub
-
[PR] fix: normalize session timezone IDs before passing them to native code [datafusion-comet]
via GitHub
-
Re: [PR] feat: support AtLeastNNonNulls natively [datafusion-comet]
via GitHub
-
[PR] perf: optimize native `CASE WHEN` and `IF` (up to 12x faster) [datafusion-comet]
via GitHub
-
Re: [PR] fix: preserve duplicate named_struct fields in codegen dispatch [datafusion-comet]
via GitHub
-
Re: [PR] feat: Use DataFusion date_trunc for compatible trunc and date_trunc paths [datafusion-comet]
via GitHub
-
Re: [PR] fix: enforce null-key rejection and mapKeyDedupPolicy in native map construction [datafusion-comet]
via GitHub
-
[PR] fix: fall back from the native CSV scan for timestamps outside UTC [datafusion-comet]
via GitHub
-
[PR] feat: prototype frozen final aggregation buckets [datafusion]
via GitHub
-
[PR] fix: count the days partition transform in UTC [datafusion-comet]
via GitHub
-
Re: [PR] fix: make wide date-to-timestamp casts safe [datafusion-comet]
via GitHub
-
[PR] ci: Checkout repo first in the Release Version Labeler workflow [datafusion]
via GitHub
-
[PR] fix: relabel CASE and COALESCE timestamp branches instead of panicking [datafusion-comet]
via GitHub
-
Re: [PR] fix: fall back for Parquet datetime rebasing [datafusion-comet]
via GitHub
-
Re: [PR] fix: refuse codegen dispatch for a TRY cast that can put a null key in a map [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] log partial memory grants at DEBUG and drop the memory usage dump (#6269) [datafusion-comet]
via GitHub
-
[PR] fix: keep the input's timezone label on native date_trunc results [datafusion-comet]
via GitHub
-
Re: [PR] ci: key TPC dataset caches on pinned generators [datafusion-comet]
via GitHub
-
[PR] docs: [branch-1.1] generate release docs for 1.1.0 [datafusion-comet]
via GitHub
-
Re: [PR] fix: preserve DPP filters during transition revert [datafusion-comet]
via GitHub
-
Re: [PR] refactor: remove unused cast compatibility flag [datafusion-comet]
via GitHub
-
Re: [PR] fix: reverse ordering and range in NegativeExpr::get_properties [datafusion-comet]
via GitHub
-
Re: [PR] perf: build spark_size LargeList lengths from i64 offsets [datafusion-comet]
via GitHub
-
[PR] perf: reuse Variant reconstruction buffers [datafusion-comet]
via GitHub
-
[PR] docs: simplify the memory tuning page [datafusion-comet]
via GitHub
-
[I] Wrong results: a list `unnest` keeps its input's GROUP BY key, so a later DISTINCT or GROUP BY drops rows [datafusion]
via GitHub
-
[PR] [branch-55] perf: Avoid copying when materializing output in OrderedPartialAggregateStream (#25312) [datafusion]
via GitHub
-
[PR] [branch-55] fix: prune row groups when file statistics collapse the predicate to a constant (#24770) [datafusion]
via GitHub
-
[PR] fix: label native timestamp_seconds results as UTC timestamps [datafusion-comet]
via GitHub
-
Re: [PR] feat: support NullType output types in codegen dispatch [datafusion-comet]
via GitHub
-
[PR] docs: check timezone handling against the contributor guide in the PR review skills [datafusion-comet]
via GitHub
-
Re: [PR] fix: fall back to Spark for negative-scale decimal casts and arithmetic [datafusion-comet]
via GitHub
-
[I] Validate cached file statistics and Parquet metadata with `e_tag`, not just size and mtime [datafusion]
via GitHub
-
[PR] fix: zero sliced boolean offsets at every level before exporting to the JVM [datafusion-comet]
via GitHub
-
[PR] docs(chaos): don't link to the feature-gated k8s module [datafusion-ballista]
via GitHub
-
[PR] perf:settle never-matching streamed rows against the buffered extreme… [datafusion]
via GitHub
-
Re: [PR] perf: add direct native reads for broadcast exchange [datafusion-comet]
via GitHub
-
Re: [PR] feat: support PivotFirst aggregate for the optimized PIVOT fast path [datafusion-comet]
via GitHub
-
[I] days transform is evaluated in the session timezone, while hours and Iceberg use UTC [datafusion-comet]
via GitHub
-
[PR] test: [branch-1.0] wait for Parquet write plan callbacks (#6108) [datafusion-comet]
via GitHub
-
Re: [I] MySQL `->` / `->>` bind too loosely against arithmetic and bitwise operators [datafusion-sqlparser-rs]
via GitHub
-
[PR] docs: add contributor guide page on timezone handling [datafusion-comet]
via GitHub
-
[PR] bench: classic pwmj sql benchmark [datafusion]
via GitHub