Messages by Thread
-
-
Re: [I] Improve performance of first/last aggregates [datafusion-comet]
via GitHub
-
[PR] oss: add benches remaining scalars [datafusion-comet]
via GitHub
-
[PR] refactor: deprecate statistics providers that duplicate operator estimates [datafusion]
via GitHub
-
[PR] fix: [branch-1.1] defer Parquet conversion errors until a row group is decoded, as Spark does (#6515) [datafusion-comet]
via GitHub
-
Re: [PR] fix: Set Substrait output_type on window functions and LIKE [datafusion]
via GitHub
-
[PR] fix: keep decimal fractions in modulo by one [datafusion]
via GitHub
-
Re: [I] Fuse operations in `equal_rows_arr` [datafusion]
via GitHub
-
Re: [PR] Add a fuzz target that flags superlinear parsing and printing [datafusion-sqlparser-rs]
via GitHub
-
[I] Final aggregate under AQE reads its shuffle through the JVM decoder instead of native direct read [datafusion-comet]
via GitHub
-
[PR] Js/store partitioning in dynamic filters [datafusion]
via GitHub
-
Re: [PR] fix: mark avg, bit_and/or/xor, stddev and variance as order-insensitive [datafusion]
via GitHub
-
Re: [I] Spark `xxhash64` ignores the running hash for dictionary-encoded values inside structs and lists [datafusion]
via GitHub
-
[PR] Show decimal add bug [datafusion]
via GitHub
-
Re: [PR] perf: bounded distinct count optimization [datafusion]
via GitHub
-
[PR] fix: hash every NaN like the canonical NaN in Spark xxhash64 [datafusion]
via GitHub
-
[PR] fix: array_concat panics or fails on non-default nested lists [datafusion]
via GitHub
-
Re: [PR] Capture the `COPY ... FROM STDIN` payload verbatim [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] ci: shard the Iceberg extensions test task [datafusion-comet]
via GitHub
-
Re: [PR] Add ordered bind placeholder dialect metadata [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] Support PostgreSQL `ALTER DEFAULT PRIVILEGES` [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] PostgreSQL: parse `CREATE GROUP` and `DROP GROUP` [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] feat: stage ORDER BY expression evaluation [datafusion]
via GitHub
-
[PR] perf: PartitionedTopKRank on the shared store with decide-then-gather [datafusion]
via GitHub
-
Re: [I] TPC-H SF1000 results for Ballista main (DataFusion 55.1.0): 0.72x of Spark 3.5, 1.45x of Spark + Comet, 15% of time in planning (#2497) [datafusion-ballista]
via GitHub
-
Re: [I] Proposal for more efficient disk-based shuffle mechanism [datafusion-ballista]
via GitHub
-
[I] PostgreSQL DO statement [datafusion-sqlparser-rs]
via GitHub
-
[PR] ci: Enforce PR limit to 3 for non-committers [datafusion]
via GitHub
-
[PR] fix: keep Spark's cache format when Kryo would reject Comet's or Comet disables itself [datafusion-comet]
via GitHub
-
[PR] perf: normalize float hash join build keys once [datafusion]
via GitHub
-
Re: [I] CI should run Ballista integration tests [datafusion-ballista]
via GitHub
-
Re: [I] Implement dynamic discrete pruning through a join [datafusion]
via GitHub
-
Re: [PR] fix: align nested collection buffer nullability before spilling [datafusion-comet]
via GitHub
-
Re: [PR] Plan struct colon access as get_field [datafusion]
via GitHub
-
Re: [I] fix: string-to-timestamp does not trim ISO control characters, and leading '+' returns null under ANSI [datafusion-comet]
via GitHub
-
Re: [I] panic: GROUP BY or SELECT DISTINCT over a zero-field (empty) struct key [datafusion]
via GitHub
-
Re: [PR] fix: gate the power and log rewrites on the base's value, not its nullability [datafusion]
via GitHub
-
Re: [PR] fix: skip the null-aware join when a NOT IN correlation duplicates the IN equality [datafusion]
via GitHub
-
Re: [PR] ci: drop non-existent suites from the Linux and macOS test matrices [datafusion-comet]
via GitHub
-
Re: [PR] fix: avoid query failures from inferred numeric cast bounds [datafusion]
via GitHub
-
Re: [PR] fix: preserve sorts for wrapping integer negation [datafusion]
via GitHub
-
[I] Support rolling upgrades without an outage window [datafusion-ballista]
via GitHub
-
Re: [I] Consolidate built-in statistics estimation in one place [datafusion]
via GitHub
-
[PR] feat: record why Spark scans a relation cached in Comet's format when Comet is off [datafusion-comet]
via GitHub
-
Re: [I] Preserve primitive Iceberg pruning beside unsupported complex null conjuncts [datafusion-comet]
via GitHub
-
Re: [I] Q10 SF1000 regression after #2315: SortPreservingMergeExec exhausts fair memory pool under AQE default-on [datafusion-ballista]
via GitHub
-
Re: [I] Avoid single-task execution for null-aware anti joins (NOT IN subqueries) [datafusion-ballista]
via GitHub
-
Re: [I] Implement fast path for QueryStageExec when writing 1 shuffle partition [datafusion-ballista]
via GitHub
-
Re: [I] Implement hash partitioned aggregation in Ballista [datafusion-ballista]
via GitHub
-
Re: [I] The type of Timestamp(Nanosecond, None) >= Date32 of binary physical should be same [datafusion-ballista]
via GitHub
-
Re: [I] The type of Float64 >= Int64 of binary physical should be same [datafusion-ballista]
via GitHub
-
Re: [I] Integration test script should test the Python bindings [datafusion-ballista]
via GitHub
-
[PR] fix: plan and run repeated SUM(x + literal) on decimals and intervals [datafusion]
via GitHub
-
Re: [I] Move ballista-proto to separate module/crate [datafusion-ballista]
via GitHub
-
Re: [I] Scheduler infinite loop after failed/canceled job [datafusion-ballista]
via GitHub
-
Re: [I] Add python crate to workspace [datafusion-ballista]
via GitHub
-
Re: [I] Stop showing very verbose streaming errors [datafusion-ballista]
via GitHub
-
Re: [I] Implement memory management in shuffle writer [datafusion-ballista]
via GitHub
-
Re: [I] Configuration guide should be generated from code [datafusion-ballista]
via GitHub
-
Re: [I] Ballista: Executor must return statistics in CompletedTask / CompletedJob [datafusion-ballista]
via GitHub
-
Re: [I] Broadcast join promotion drops a sort order required by a parent SortMergeJoinExec, silently producing wrong results [datafusion-ballista]
via GitHub
-
Re: [I] Enable benchmark data validation for distributed execution [datafusion-ballista]
via GitHub
-
Re: [I] Optimize shuffle before coalesce [datafusion-ballista]
via GitHub
-
Re: [I] Provide integration with Grafana / Prometheus [datafusion-ballista]
via GitHub
-
Re: [I] [Python] Allow configuration options to be set when creating BallistaContext [datafusion-ballista]
via GitHub
-
Re: [I] Add "Getting Started" docs to repo [datafusion-ballista]
via GitHub
-
Re: [PR] feat: enable Comet's in-memory cache by default [datafusion-comet]
via GitHub
-
Re: [I] Ballista: Fix hacks around concurrency=2 to force hash-partitioned joins [datafusion-ballista]
via GitHub
-
Re: [I] Ballista should serialize Parquet statistics [datafusion-ballista]
via GitHub
-
Re: [I] Remove `datafusion.proto` and use the `datafusion-proto` api for logical plan serde [datafusion-ballista]
via GitHub
-
Re: [I] Improve new top-level README for this project [datafusion-ballista]
via GitHub
-
[I] S3 credential SPI: follow-ups from #6478 (hostless Iceberg metadata locations, docs, tests) [datafusion-comet]
via GitHub
-
Re: [I] Account for partitions in local TopK statistics [datafusion]
via GitHub
-
Re: [I] date_trunc rejects scalar timestamps that array evaluation accepts [datafusion]
via GitHub
-
[PR] refactor: move shuffle planning into dedicated module [datafusion-comet]
via GitHub
-
[PR] test: strengthen scalar UDF type recovery regression coverage [datafusion]
via GitHub
-
Re: [I] Distributed execution returns incorrect results for several TPC-DS queries (static planner) [datafusion-ballista]
via GitHub
-
Re: [I] Job hangs indefinitely instead of failing when all executors are lost [datafusion-ballista]
via GitHub
-
[PR] Align statistics cache sharing test comments with reset-free contract [datafusion]
via GitHub
-
[PR] fix: keep decimal and temporal types in nvl and ifnull [datafusion]
via GitHub
-
Re: [I] Implement FFI versions of runtime environment and execution properties [datafusion]
via GitHub
-
[PR] chore: Hide internal public utility APIs [datafusion]
via GitHub
-
Re: [PR] fix: use NDV for unresolved scalar subquery selectivity instead of 20% fallback [datafusion]
via GitHub
-
Re: [PR] fix: use session statistics registry in dfbench reports [datafusion]
via GitHub
-
Re: [PR] feat: add opt-in probe selection for partitioned inner hash joins [datafusion]
via GitHub
-
[PR] refactor: Implement optimizer hook for combining partial/final aggregate in `AggregateExec` [datafusion]
via GitHub
-
Re: [PR] Add ComposedNamedPhysicalExtensionCodec [datafusion]
via GitHub
-
Re: [PR] fix: infer LIMIT and OFFSET parameter types [datafusion]
via GitHub
-
Re: [PR] Feat(parquet) : Introduce support for optionally writing Distinct Values to statistics [datafusion]
via GitHub
-
Re: [PR] perf: adapt generic hash join allocation to key cardinality [datafusion]
via GitHub
-
Re: [PR] chore: guard RangeExpr protobuf fields exhaustively [datafusion]
via GitHub