This is an automated email from the ASF dual-hosted git repository.
mrhhsg pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris-website.git
The following commit(s) were added to refs/heads/master by this push:
new bf9892167d0 docs: clarify MAP_AGG duplicate-key handling (#4175)
bf9892167d0 is described below
commit bf9892167d06541fac3cfcb28b4e0ff571653ebe
Author: HappenLee <[email protected]>
AuthorDate: Mon Sep 28 15:04:02 2026 +0800
docs: clarify MAP_AGG duplicate-key handling (#4175)
`MAP_AGG` keeps the first value encountered for a key and ignores later
values for the same key, including when merging partial states. Document
this behavior and explain why processing order does not guarantee the
earliest inserted value or a stable winner across queries. Add a
duplicate-key example showing both possible outcomes.
Update the English and Chinese dev, 4.x and 3.x pages together. Also
fence the syntax and remove an orphaned English example whose output
does not match the page's sample data.
Related: https://github.com/apache/doris/pull/68496
## Versions
- [x] dev
- [x] 4.x
- [x] 3.x
- [ ] 2.1 or older (not covered by version/language sync gate)
## Languages
- [x] Chinese
- [x] English
## Docs Checklist
- [x] Checked by AI
- [x] Test Cases Built (new duplicate-key SQL example executed locally)
- [x] Updated required version and language counterparts, or explained
why not
- [x] If only one language changed, confirmed whether source/translation
counterparts need sync (both languages updated)
## Validation
- All six changed pages compile with the MDX compiler.
- Version/language synchronization check: no findings.
- Documentation lint: no new findings compared with the same six pages
at the base commit (30 existing findings reduced to 22).
- `git diff --check` passes.
- Duplicate-key example executed on the PR FE with an existing BE at
execution version 14; returned `{1:"a"}` as one allowed outcome. Other
versions were checked against implementation semantics, not separate
running clusters.
---
.../sql-functions/aggregate-functions/map-agg.md | 33 ++++++++++++++--------
.../sql-functions/aggregate-functions/map-agg.md | 27 +++++++++++++++++-
.../sql-functions/aggregate-functions/map-agg.md | 27 +++++++++++++++++-
.../sql-functions/aggregate-functions/map-agg.md | 27 +++++++++++++++++-
.../sql-functions/aggregate-functions/map-agg.md | 27 +++++++++++++++++-
.../sql-functions/aggregate-functions/map-agg.md | 33 ++++++++++++++--------
6 files changed, 148 insertions(+), 26 deletions(-)
diff --git a/docs/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/docs/sql-manual/sql-functions/aggregate-functions/map-agg.md
index 9070c27816b..cad416f1b12 100644
--- a/docs/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++ b/docs/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
The MAP_AGG function is used to form a mapping structure based on key-value
pairs from multiple rows of data.
+For duplicate keys, `MAP_AGG` keeps the value encountered first during
aggregation and ignores subsequent values for a key that already exists. When
partial aggregate states are merged, the value already present in the
destination state is also retained.
+
+“Encountered first” refers to processing order, not insertion order. Scanning,
parallel execution, and the order in which partial states are merged can affect
which value is retained, so the value selected for a duplicate key is not
guaranteed to be the same across queries. An outer `ORDER BY` only sorts result
rows; it does not determine which value is retained for a duplicate key.
+
## Syntax
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## Parameters
@@ -78,17 +84,22 @@ select map_agg(`n_name`, `n_nationkey` % 5) from `nation`
where n_nationkey is n
| {} |
+--------------------------------------+
```
-select n_regionkey, map_agg(`n_name`, `n_nationkey` % 5) from `nation` group
by `n_regionkey`;
+
+The following query contains a duplicate key, so the result contains only one
key-value pair. The result can be `{1:"a"}` or `{1:"b"}`, depending on which
value is processed first. One possible output is:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
```
```text
-+-------------+------------------------------------------------------------------------+
-| n_regionkey | map_agg(`n_name`, (`n_nationkey` % 5))
|
-+-------------+------------------------------------------------------------------------+
-| 2 | {"INDIA":3, "INDONESIA":4, "JAPAN":2, "CHINA":3, "VIETNAM":1}
|
-| 0 | {"ALGERIA":0, "ETHIOPIA":0, "KENYA":4, "MOROCCO":0,
"MOZAMBIQUE":1} |
-| 3 | {"FRANCE":1, "GERMANY":2, "ROMANIA":4, "RUSSIA":2, "UNITED
KINGDOM":3} |
-| 1 | {"ARGENTINA":1, "BRAZIL":2, "CANADA":3, "PERU":2, "UNITED
STATES":4} |
-| 4 | {"EGYPT":4, "IRAN":0, "IRAQ":1, "JORDAN":3, "SAUDI ARABIA":0}
|
-+-------------+------------------------------------------------------------------------+
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
```
diff --git
a/i18n/zh-CN/docusaurus-plugin-content-docs/current/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/i18n/zh-CN/docusaurus-plugin-content-docs/current/sql-manual/sql-functions/aggregate-functions/map-agg.md
index a00643ec2b3..2b2fe0d7ec8 100644
---
a/i18n/zh-CN/docusaurus-plugin-content-docs/current/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++
b/i18n/zh-CN/docusaurus-plugin-content-docs/current/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
MAP_AGG 函数用于根据多行数据中的键值对形成一个映射结构。
+对于相同的 key,`MAP_AGG` 保留聚合过程中先遇到的 value;如果 key 已存在,则忽略后来遇到的
value。合并局部聚合结果时,同样保留目标聚合状态中已有 key 的 value。
+
+这里的“先遇到”指执行时的处理顺序,不代表数据写入顺序。扫描、并行执行和局部聚合结果的合并顺序都可能影响最终保留的 value,因此重复 key 对应的
value 不保证在多次查询之间保持一致。查询外层的 `ORDER BY` 只对结果行排序,不能决定重复 key 保留哪个 value。
+
## 语法
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## 参数说明
@@ -79,3 +85,22 @@ select map_agg(`n_name`, `n_nationkey` % 5) from `nation`
where n_nationkey is n
| {} |
+--------------------------------------+
```
+
+下面的查询包含两个相同的 key,结果只包含一个键值对。输出可能为 `{1:"a"}` 或 `{1:"b"}`,取决于执行时先处理哪个
value。以下是一种可能的输出:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
+```
+
+```text
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
+```
diff --git
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
index 36f363e0250..6cf9d7f2d2f 100644
---
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
MAP_AGG 函数用于根据多行数据中的键值对形成一个映射结构。
+对于相同的 key,`MAP_AGG` 保留聚合过程中先遇到的 value;如果 key 已存在,则忽略后来遇到的
value。合并局部聚合结果时,同样保留目标聚合状态中已有 key 的 value。
+
+这里的“先遇到”指执行时的处理顺序,不代表数据写入顺序。扫描、并行执行和局部聚合结果的合并顺序都可能影响最终保留的 value,因此重复 key 对应的
value 不保证在多次查询之间保持一致。查询外层的 `ORDER BY` 只对结果行排序,不能决定重复 key 保留哪个 value。
+
## 语法
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## 参数说明
@@ -108,3 +114,22 @@ select n_regionkey, map_agg(`n_name`, `n_nationkey` % 5)
from `nation` group by
| 4 | {"EGYPT":4, "IRAN":0, "IRAQ":1, "JORDAN":3, "SAUDI ARABIA":0}
|
+-------------+------------------------------------------------------------------------+
```
+
+下面的查询包含两个相同的 key,结果只包含一个键值对。输出可能为 `{1:"a"}` 或 `{1:"b"}`,取决于执行时先处理哪个
value。以下是一种可能的输出:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
+```
+
+```text
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
+```
diff --git
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
index 75e80bb4b9d..c5896894691 100644
---
a/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++
b/i18n/zh-CN/docusaurus-plugin-content-docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
MAP_AGG 函数用于根据多行数据中的键值对形成一个映射结构。
+对于相同的 key,`MAP_AGG` 保留聚合过程中先遇到的 value;如果 key 已存在,则忽略后来遇到的
value。合并局部聚合结果时,同样保留目标聚合状态中已有 key 的 value。
+
+这里的“先遇到”指执行时的处理顺序,不代表数据写入顺序。扫描、并行执行和局部聚合结果的合并顺序都可能影响最终保留的 value,因此重复 key 对应的
value 不保证在多次查询之间保持一致。查询外层的 `ORDER BY` 只对结果行排序,不能决定重复 key 保留哪个 value。
+
## 语法
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## 参数说明
@@ -80,3 +86,22 @@ select map_agg(`n_name`, `n_nationkey` % 5) from `nation`
where n_nationkey is n
+--------------------------------------+
```
+
+下面的查询包含两个相同的 key,结果只包含一个键值对。输出可能为 `{1:"a"}` 或 `{1:"b"}`,取决于执行时先处理哪个
value。以下是一种可能的输出:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
+```
+
+```text
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
+```
diff --git
a/versioned_docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/versioned_docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
index b24acbf42bf..51104e35c59 100644
---
a/versioned_docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++
b/versioned_docs/version-3.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
The MAP_AGG function is used to form a mapping structure based on key-value
pairs from multiple rows of data.
+For duplicate keys, `MAP_AGG` keeps the value encountered first during
aggregation and ignores subsequent values for a key that already exists. When
partial aggregate states are merged, the value already present in the
destination state is also retained.
+
+“Encountered first” refers to processing order, not insertion order. Scanning,
parallel execution, and the order in which partial states are merged can affect
which value is retained, so the value selected for a duplicate key is not
guaranteed to be the same across queries. An outer `ORDER BY` only sorts result
rows; it does not determine which value is retained for a duplicate key.
+
## Syntax
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## Parameters
@@ -108,3 +114,22 @@ select n_regionkey, map_agg(`n_name`, `n_nationkey` % 5)
from `nation` group by
| 4 | {"EGYPT":4, "IRAN":0, "IRAQ":1, "JORDAN":3, "SAUDI ARABIA":0}
|
+-------------+------------------------------------------------------------------------+
```
+
+The following query contains a duplicate key, so the result contains only one
key-value pair. The result can be `{1:"a"}` or `{1:"b"}`, depending on which
value is processed first. One possible output is:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
+```
+
+```text
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
+```
diff --git
a/versioned_docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
b/versioned_docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
index d12878d9372..eede503d815 100644
---
a/versioned_docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
+++
b/versioned_docs/version-4.x/sql-manual/sql-functions/aggregate-functions/map-agg.md
@@ -10,9 +10,15 @@
The MAP_AGG function is used to form a mapping structure based on key-value
pairs from multiple rows of data.
+For duplicate keys, `MAP_AGG` keeps the value encountered first during
aggregation and ignores subsequent values for a key that already exists. When
partial aggregate states are merged, the value already present in the
destination state is also retained.
+
+“Encountered first” refers to processing order, not insertion order. Scanning,
parallel execution, and the order in which partial states are merged can affect
which value is retained, so the value selected for a duplicate key is not
guaranteed to be the same across queries. An outer `ORDER BY` only sorts result
rows; it does not determine which value is retained for a duplicate key.
+
## Syntax
-`MAP_AGG(<expr1>, <expr2>)`
+```sql
+MAP_AGG(<expr1>, <expr2>)
+```
## Parameters
@@ -78,17 +84,22 @@ select map_agg(`n_name`, `n_nationkey` % 5) from `nation`
where n_nationkey is n
| {} |
+--------------------------------------+
```
-select n_regionkey, map_agg(`n_name`, `n_nationkey` % 5) from `nation` group
by `n_regionkey`;
+
+The following query contains a duplicate key, so the result contains only one
key-value pair. The result can be `{1:"a"}` or `{1:"b"}`, depending on which
value is processed first. One possible output is:
+
+```sql
+SELECT MAP_AGG(k, v) AS result
+FROM (
+ SELECT 1 AS k, 'a' AS v
+ UNION ALL
+ SELECT 1 AS k, 'b' AS v
+) AS input;
```
```text
-+-------------+------------------------------------------------------------------------+
-| n_regionkey | map_agg(`n_name`, (`n_nationkey` % 5))
|
-+-------------+------------------------------------------------------------------------+
-| 2 | {"INDIA":3, "INDONESIA":4, "JAPAN":2, "CHINA":3, "VIETNAM":1}
|
-| 0 | {"ALGERIA":0, "ETHIOPIA":0, "KENYA":4, "MOROCCO":0,
"MOZAMBIQUE":1} |
-| 3 | {"FRANCE":1, "GERMANY":2, "ROMANIA":4, "RUSSIA":2, "UNITED
KINGDOM":3} |
-| 1 | {"ARGENTINA":1, "BRAZIL":2, "CANADA":3, "PERU":2, "UNITED
STATES":4} |
-| 4 | {"EGYPT":4, "IRAN":0, "IRAQ":1, "JORDAN":3, "SAUDI ARABIA":0}
|
-+-------------+------------------------------------------------------------------------+
++---------+
+| result |
++---------+
+| {1:"a"} |
++---------+
```
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]