764276020 opened a new issue, #68538:
URL: https://github.com/apache/doris/issues/68538

   ### Environment
   - Doris **3.1.4**, **cloud / compute-storage-separation** mode (FDB-backed 
MetaService)
   - Large log table: 54+ TB, 184 daily partitions, ~23-39 buckets per 
partition (5754 tablets)
   - Operation: `ALTER TABLE coremail.<t> ADD INDEX ... USING INVERTED ...` 
(heavy schema-change path, `SchemaChangeJobV2` / shadow index)
   
   ### Symptom
   The schema change fails with a MetaService transaction-size error and the 
whole job is then CANCELLED:
   
   ```
   schema change tasks failed, error reason: task type: ALTER, status_code: 
INVALID_ARGUMENT,
   status_message: [(doris-dwcloud-be-w4sata-20...)[INVALID_ARGUMENT]failed to 
commit tablet job:
   failed to commit job kv, err=Transaction exceeds byte limit],
   backendId: Backend [id=1743164044730, host=doris-dwcloud-be-w4sata-20...]
   ```
   
   Then `SHOW ALTER TABLE COLUMN` shows the job as `CANCELLED` with that 
message.
   
   ### Root cause (code)
   `CloudSchemaChangeJob`'s COMMIT is executed by MetaService 
`finish_tablet_job` and everything is packed into **one FDB transaction**:
   
   - `cloud/src/meta-service/meta_service_job.cpp:1787` – "move rowsets 
[2-alter_version] to recycle" (one `RecycleRowsetPB` with a full 
`RowsetMetaCloudPB` per rowset)
   - `cloud/src/meta-service/meta_service_job.cpp:1963` – `for (size_t i = 0; i 
< schema_change.txn_ids().size(); ++i)` converts **every** tmp rowset to a 
formal rowset in the same txn
   - commit failure surfaces as `failed to commit job kv` (`:2204` in master, 
`:1641` in 3.1.4)
   - `cloud/src/meta-store/txn_kv_error.h:55` – `Transaction exceeds byte 
limit` (FDB 10 MB hard limit)
   
   Transaction size therefore scales with the number of rowsets of the tablet. 
Schema change writes `tmp rowset -> formal rowset` **1:1**, so a tablet with 
many small rowsets (frequent small loads + compaction lag) deterministically 
blows the 10 MB limit. There is **no batching, no `approximate_bytes` pre-check 
and no resumable split** in this path (searched for `approximate_bytes` / 
`max_txn_commit_byte` / batch in that function: none).
   
   For contrast, the **load** path already handles exactly this problem:
   - `cloud/src/common/config.h:356` `enable_cloud_txn_lazy_commit = true`
   - `cloud/src/common/config.h:358` `txn_lazy_commit_rowsets_thresold = 1000`
   - `cloud/src/common/config.h:362` `txn_lazy_max_rowsets_per_batch = 1000`
   - `cloud/src/meta-service/txn_lazy_committer.cpp` commits in batches of <= 
1000 rowsets
   
   So "one txn must not hold too many rowsets" is known to the project; the 
schema-change commit path just was not covered.
   
   ### Impact
   - The failing tablet makes the job **permanently un-runnable** 
(deterministic, not transient): re-running the ALTER reproduces it.
   - No retry at the FE level: `AlterJobV2.getRetryTimes()` 
(`fe/fe-core/.../alter/AlterJobV2.java:267-275`) only retries 
`DELETE_BITMAP_LOCK_ERROR` / `NETWORK_ERROR`, and replica_num = 1 in cloud, so 
the first failed tablet task cancels the whole job (hours of conversion work 
lost, plus the rewritten objects in object storage).
   - On our 54 TB / 86 TB tables this makes ADD INDEX practically impossible 
without pre-compaction.
   
   ### Suggested fix
   Batch the COMMIT like the load path does, e.g. convert tmp->formal rowset 
and write the recycle records in chunks of N rowsets per transaction, making 
the operation resumable/idempotent (so the BE can continue instead of failing 
the task); or at least split when the txn size approaches the limit.
   
   ### Workaround we use
   Run `ADMIN COMPACT TABLE <db>.<tbl> PARTITION (<p>) WHERE type='BASE';` for 
the affected partitions first, to reduce the rowset count per tablet, then 
re-run the ALTER (one table at a time).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to