rishabh1244 opened a new pull request, #25740:
URL: https://github.com/apache/datafusion/pull/25740

   ## Which issue does this PR close?
   
   - Closes #22993
   
   ## Rationale for this change
   
   DataFusion has no `map_agg` aggregate function, even though equivalent 
functions exist in
   other engines (Trino/Presto `map_agg`, Spark `map_from_arrays`, PostgreSQL 
`json_object_agg`).
   Users currently have no way to aggregate `(key, value)` pairs from many rows 
into a single
   `Map` value — the map equivalent of `array_agg`.
   
   ## What changes are included in this PR?
   
   - New `map_agg(key, value)` aggregate UDAF in 
`datafusion/functions-aggregate/src/map_agg.rs`:
     - `MapAgg` implements `AggregateUDFImpl` with `Signature::any(2, 
Volatility::Immutable)`
       and a `Map(entries { key: non-null, value: nullable })` return type
     - `MapAggAccumulator` collects key/value batches, concatenates them in 
`evaluate()` and
       emits a single-row `MapArray` scalar
     - `state()` / `merge_batch()` support partial aggregation, so `map_agg` 
works with
       `GROUP BY` in multi-partition (Partial + Final) plans
     - Empty input produces a `NULL` map (same convention as `array_agg`)
     - Null keys are rejected with `map key cannot be null`, matching the 
`map()` UDF, since
       Arrow maps cannot represent null keys
   - Registered in `all_default_aggregate_functions()` and exported via 
`expr_fn::map_agg`
   - Regenerated `docs/source/user-guide/sql/aggregate_functions.md`
   
   ## What is the testing strategy for this PR?
   
   Unit tests in `datafusion/functions-aggregate/src/map_agg.rs`:
   
   - `map_agg_builds_map_from_batches` — two batches concatenated, 
entries/values/null values checked
   - `map_agg_empty_group_produces_null_map` — empty input returns a null map 
scalar
   - `map_agg_null_key_errors` — null key is rejected
   - `map_agg_state_merge_roundtrip` — `state()` output merged back yields the 
same map
     (exercises the Partial + Final path)
   - `map_agg_update_batch_rejects_mismatched_lengths` — key/value length 
mismatch errors
   - plus existing name / return_type / registration tests
   
   Verified end-to-end with the CLI, including the example from the issue:
   
   ```sql
   SELECT k, map_agg(v, m)
   FROM (VALUES
       ('a', 1, 10),
       ('a', 2, 20),
       ('b', 3, 30)
   ) t(k, v, m)
   GROUP BY k
   ORDER BY k;
   +---+------------------+
   | k | map_agg(t.v,t.m) |
   +---+------------------+
   | a | {1: 10, 2: 20}   |
   | b | {3: 30}          |
   +---+------------------+
   
   Also verified NULL for empty input, the null-key error, and EXPLAIN shows
   Partial + Final AggregateExec nodes for the grouped query.
   
   Are there any user-facing changes?
   
   Yes — map_agg(key, value) becomes available in SQL, and it is documented in
   docs/source/user-guide/sql/aggregate_functions.md.
   Known limitations to discuss in review:
   
   - Duplicate keys are currently kept as-is, not rejected. The map() UDF 
errors on
   duplicates (map key must be unique); aligning map_agg would be a small 
follow-up.
   
   - ORDER BY inside the aggregate (as array_agg supports) is not implemented; 
it can
   be added separately if wanted.
   
   ## What is the testing strategy for this PR?
   
   - Unit tests in `datafusion/functions-aggregate/src/map_agg.rs`: map 
building across
     batches, empty input → NULL, null-key rejection, `state()`/`merge_batch()` 
roundtrip,
     key/value length validation, plus name/return_type/registration tests.
   - sqllogictest coverage in `datafusion/sqllogictest/test_files/map_agg.slt`:
   
   ```sql
   SELECT k, map_agg(v, m)
   FROM (VALUES
       ('a', 1, 10),
       ('a', 2, 20),
       ('b', 3, 30)
   ) t(k, v, m)
   GROUP BY k
   ORDER BY k;
   +---+-----------------+
   | k | map_agg(v, m)   |
   +---+-----------------+
   | a | {1: 10, 2: 20}  |
   | b | {3: 30}         |
   +---+-----------------+
     Also covered there: NULL values retained, NULL key rejected, empty input → 
NULL,
     and multi-batch aggregation (batch_size = 1).


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to