Messages by Thread
-
[PR] perf: charge the local shuffle writer's buffers to the memory pool [datafusion-comet]
via GitHub
-
[I] [EPIC] Timezone handling bugs [datafusion-comet]
via GitHub
-
[I] Native IF fails with "column types must match schema types" when its branches differ in nested nullability [datafusion-comet]
via GitHub
-
Re: [PR] feat: add native distinct-combined collect_list aggregate support [datafusion-comet]
via GitHub
-
[I] Native timezone rules come from chrono-tz's bundled tzdata, so results differ from Spark when the JVM's tzdata differs [datafusion-comet]
via GitHub
-
[I] Native CSV V2 scan parses timestamps as UTC and ignores the session timezone [datafusion-comet]
via GitHub
-
[I] Native date_trunc labels its output with the session timezone, so comparing it with another timestamp fails in Etc/UTC sessions [datafusion-comet]
via GitHub
-
[I] Session timezone IDs such as GMT+8, Z and PST make native timestamp expressions fail [datafusion-comet]
via GitHub
-
[I] CASE and COALESCE over timestamps panic when the branches have different Arrow timezones [datafusion-comet]
via GitHub
-
[I] Native timestamp_seconds returns a TIMESTAMP_NTZ-typed array, so downstream expressions ignore the session timezone [datafusion-comet]
via GitHub
-
[PR] fix: fall back when struct field names collide case-insensitively [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.0] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm (#6025) [datafusion-comet]
via GitHub
-
Re: [I] [FEATURE] Native scan support for VariantType columns (Iceberg + Spark 4.0) [datafusion-comet]
via GitHub
-
Re: [I] [FEATURE] Shredded Parquet Reader/Writer support for Variant type [datafusion-comet]
via GitHub
-
[PR] test: run microbenchmarks in extended CI [datafusion]
via GitHub
-
[PR] fix: arrays_zip with duplicate field names falls back to Spark [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm (#6025) [datafusion-comet]
via GitHub
-
[I] LATERAL subquery fails with a schema error when a derived table alias differs from the table name [datafusion]
via GitHub
-
[I] Limit the package root to the common entry points [datafusion-python]
via GitHub
-
[PR] fix: pull a correlated filter only through plan nodes known to be safe [datafusion]
via GitHub
-
[PR] docs: redraw the home page architecture diagram as a simple SVG [datafusion-comet]
via GitHub
-
[I] Bug triage results: 2026-09-28 [datafusion-comet]
via GitHub
-
[PR] PostgreSQL: Remove test `parse_create_table_with_inherit` [datafusion-sqlparser-rs]
via GitHub
-
[PR] docs: document casts involving arrays, structs, and maps [datafusion-comet]
via GitHub
-
Re: [I] Native S3 scan on EKS/IRSA turns a transient STS throttle into a hard 403 storm [datafusion-comet]
via GitHub
-
[I] Nested DATE to numeric casts in structs and maps return the day count or fail [datafusion-comet]
via GitHub
-
[PR] fix: route nested DATE to numeric casts in structs and maps through the codegen dispatcher [datafusion-comet]
via GitHub
-
[PR] feat: [branch-1.1] built-in S3 credential provider adapters for the native Parquet scan (#6023) [datafusion-comet]
via GitHub
-
[I] Support Spark 4.2 geometry type + functions [datafusion-comet]
via GitHub
-
Re: [PR] chore(deps): bump pydata-sphinx-theme from 0.20.0 to 0.21.0 in the all-uv-deps group across 1 directory [datafusion]
via GitHub
-
[PR] fix: Estimate string equality and LIKE filter selectivity [datafusion]
via GitHub
-
Re: [I] Charge the native shuffle writer's write buffers to the memory pool [datafusion-comet]
via GitHub
-
[I] Native Parquet writes on Spark 3.4/3.5 leave INSERT INTO targets stale until REFRESH TABLE [datafusion-comet]
via GitHub
-
Re: [I] Add support for spilling in hash joins [datafusion-comet]
via GitHub
-
[I] Wrong results: a correlated filter under the right side of an ASOF JOIN is pulled above the join [datafusion]
via GitHub
-
Re: [I] size, arrays_zip, map_from_arrays and array_append return wrong answers for a nondeterministic child [datafusion-comet]
via GitHub
-
[PR] fix: publish metrics on the update interval for native blocks without a JVM input [datafusion-comet]
via GitHub
-
[I] Native blocks without a JVM input publish SQL metrics on every batch, ignoring spark.comet.metrics.updateInterval [datafusion-comet]
via GitHub
-
Re: [I] Map lookups with float, collated or complex keys fall back to Spark (`map_col[key]`, `element_at`) [datafusion-comet]
via GitHub
-
Re: [PR] fix: give a multi-set aggregate the grouping set index Substrait defines [datafusion]
via GitHub
-
Re: [I] Make Comet even friendlier to agentic development [datafusion-comet]
via GitHub
-
[PR] test: [branch-1.1] backport the Iceberg write report and the mid-write retry test (#6155, #6111) [datafusion-comet]
via GitHub
-
Re: [PR] feat: give PhysicalExtensionCodec the full decode context [datafusion]
via GitHub
-
[PR] test: fix flaky mixed field id directory test on Spark 3.4 and 3.5 [datafusion-comet]
via GitHub
-
Re: [I] Native Parquet S3 scan fails on credential provider classes that Spark/Hadoop accept [datafusion-comet]
via GitHub
-
[I] Run benchmarks for arrow-rs bump [datafusion]
via GitHub
-
Re: [I] Lack of benchmarks for RepartionExec [datafusion]
via GitHub
-
[PR] feat(core): run queries that mix information_schema with other tables on the cluster [datafusion-ballista]
via GitHub
-
Re: [I] Support of new Syntax [datafusion-sqlparser-rs]
via GitHub
-
Re: [I] `ORDER BY` parser [datafusion-sqlparser-rs]
via GitHub
-
Re: [I] Discussion: Should Comet add geospatial (ST_*) function support? [datafusion-comet]
via GitHub
-
Re: [I] Test dual-impl (native + codegen-dispatch) expressions consistently across the full routing matrix [datafusion-comet]
via GitHub
-
Re: [I] Do another audit sweep for string collation differences [datafusion-comet]
via GitHub
-
Re: [I] Improve RewriteJoin logic to calculate hash table size [datafusion-comet]
via GitHub
-
Re: [I] Fork DataFusion SMJ in Comet repo [datafusion-comet]
via GitHub
-
Re: [PR] fix: `AT TIME ZONE` on a timezone-aware timestamp returns a naive timestamp [datafusion]
via GitHub
-
Re: [I] Use coarse time when calculating metrics refresh interval [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: make_dt_interval [datafusion-comet]
via GitHub
-
Re: [I] [datafusion-spark] Test integrating datafusion-spark code into comet [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Implement expressions as ScalarUDFImpl [datafusion-comet]
via GitHub
-
Re: [I] Add tests for map types to CometFuzzTestSuite [datafusion-comet]
via GitHub
-
Re: [I] [blog post] Using Arrow/DataFusion across JVM/Rust boundary [datafusion-comet]
via GitHub
-
Re: [I] [blog post] ANSI support [datafusion-comet]
via GitHub
-
Re: [I] [blog post] Complex type support [datafusion-comet]
via GitHub
-
Re: [I] [blog post] Spark compatibility testing [datafusion-comet]
via GitHub
-
Re: [I] Proposal: Port most microbenchmarks to PySpark [datafusion-comet]
via GitHub
-
Re: [I] Improve fallback reporting for aggregates [datafusion-comet]
via GitHub
-
Re: [I] Respect Spark compression settings in native Parquet writer [datafusion-comet]
via GitHub
-
[PR] [DO NOT MERGE] Test apache/arrow-rs#10859 (fuse consecutive same-projection filters) with pushdown tests and ClickBench [datafusion]
via GitHub
-
[PR] chore: start 1.2.0 development [datafusion-comet]
via GitHub
-
Re: [I] Benchmark automation [datafusion-comet]
via GitHub
-
[I] 邀请 arrow-datafusion 加入 GithubStarMate,让更多人发现你的作品 [datafusion]
via GitHub
-
Re: [I] [Research] Use custom cost model when deciding between SMJ and SHJ [datafusion-comet]
via GitHub
-
Re: [I] Explore cost-based optimizations [datafusion-comet]
via GitHub
-
[PR] fix(proto): logical plan deserialization stack overflow protection [datafusion]
via GitHub
-
Re: [I] [EPIC] Provide JVM/codegen-dispatch implementations for Incompatible expressions so they never fall back by default [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Optimize performance for slow expressions [datafusion-comet]
via GitHub
-
[PR] fix: retry a native acquire that Spark failed after dropping the task's entry [datafusion-comet]
via GitHub
-
[I] A JVM consumer's parked page allocation fails when another consumer of the task empties its balance [datafusion-comet]
via GitHub
-
[PR] docs: refresh stale roadmap entries [datafusion-comet]
via GitHub
-
[PR] docs: update TPC-DS benchmark results for Comet 1.1.0 on Spark 4.2 [datafusion-comet]
via GitHub
-
[PR] perf: [branch-1.1] cache Iceberg FileIO per executor instead of building one per task (#6106) [datafusion-comet]
via GitHub
-
[PR] fix: [branch-1.1] decline native Iceberg writes with a custom location provider (#6216) [datafusion-comet]
via GitHub
-
Re: [PR] fix: use null counts to estimate IS NULL and IS NOT NULL [datafusion]
via GitHub
-
[I] DataFrame::fill_null and fill_nan fail on column names with uppercase letters or dots [datafusion]
via GitHub
-
Re: [I] perf: shuffle reader and writer performance review (native, JVM columnar, reader, Celeborn) [datafusion-comet]
via GitHub
-
[I] CSV/NDJSON range scans issue one sequential 16 KiB GET per 16 KiB of a long record at a range boundary [datafusion]
via GitHub
-
Re: [PR] fix: account for every partition in local TopK and limit statistics [datafusion]
via GitHub
-
Re: [I] `datafusion-examples` should only use the `datafusion` crate [datafusion]
via GitHub
-
Re: [I] Track PR release version with labels/milestones [datafusion]
via GitHub
-
Re: [I] Improve coverage of tests for array expressions [datafusion-comet]
via GitHub
-
[PR] docs: stop listing GROUPS window frames as a fallback [datafusion-comet]
via GitHub
-
Re: [I] perf: optimize native shuffle writer (redundant copies, per-block allocations) [datafusion-comet]
via GitHub
-
Re: [I] Replace small hand-rolled element loops with arrow kernels and arity helpers [datafusion-comet]
via GitHub
-
Re: [I] Custom Authentication & External File Systems [datafusion-comet]
via GitHub
-
Re: [I] sum_int: use arrow sum kernels and collapse integer type dispatch [datafusion-comet]
via GitHub
-
Re: [I] [EPIC] Implement all Spark date/time expressions [datafusion-comet]
via GitHub
-
Re: [I] Support expressions already implemented in datafusion-spark crate [datafusion-comet]
via GitHub
-
Re: [I] Native scans cannot propagate JVM-side Spark accumulators [datafusion-comet]
via GitHub
-
Re: [I] Change UDF signature to use ColumnarValue rather than raw Arrow types [datafusion-comet]
via GitHub
-
Re: [PR] chore(`PartialReduceHashAggregateStream`): convert to async generators and cleanup [datafusion]
via GitHub