Jeongmin Kim created HBASE-30459:
------------------------------------
Summary: Add an option to account compaction throughput control by
actual output (post-compression) bytes
Key: HBASE-30459
URL: https://issues.apache.org/jira/browse/HBASE-30459
Project: HBase
Issue Type: Improvement
Components: Compaction
Reporter: Jeongmin Kim
Compaction throughput limits are consumed in {{Compactor#performCompaction}} by
{{cell.getSerializedSize()}} — the pre-encoding, pre-compression size of each
cell — while the bytes actually written to disk are post-DATA_BLOCK_ENCODING /
post-compression. Under the same limit, a store whose data encodes/compresses R
times runs its compactions at roughly 1/R of the configured disk write rate.
This works against both goals of throughput control: keeping compactions inside
a predictable resource envelope, and letting them actually use the resources
that envelope grants. The better the encoding and compression work, the more of
the budget is charged for bytes that never reach the disk, and the compaction
sleeps instead of writing. In practice — we observed this in production — a
compaction of a highly compressible store spends most of its wall time in
throttle sleep while the disk stays nearly idle: even a compaction that is
small on disk takes a long time, occupies a longCompactions thread throughout,
and the compaction queue backs up. And since limits are per-RegionServer (there
is no per-table or per-family limit), stores holding incompressible payloads
consume the same budget 1:1, so mixed workloads additionally become unfair
between families.
Proposal: an opt-in configuration
{{hbase.hstore.compaction.throughput.control.by.output}} (default
{{{}false{}}}, current behavior unchanged). When enabled and the sink exposes
its output position, {{performCompaction}} calls
{{ThroughputController#control}} with the delta of the writer's output position
after each cell-batch append, instead of the cells' serialized sizes. Notes on
the implementation we have been running:
* {{StoreFileWriter#getPos}} also includes the historical file writer when
{{{}hbase.enable.historical.compaction.files=true{}}}, so the whole disk write
load is accounted; {{AbstractMultiFileWriter}} gains a {{getPos}} summing its
lower writers (stripe / date-tiered compactors).
* The position only advances when a block is encoded/compressed and flushed to
the output stream, so the controller sees actual disk bytes. The control check
interval ({{{}controlPerSize{}}}, by default the throughput lower bound) is
much larger than block sizes (32K ~ 128K), so block-granularity jumps are
absorbed by the accounting.
* Progress accounting, the shipped() cadence and {{CloseChecker}} stay on cell
serialized sizes — only the throughput accounting unit changes.
* Sinks that do not expose an output position (e.g.
{{{}DefaultMobStoreCompactor{}}}, which overrides {{performCompaction}} anyway)
keep the existing accounting; an unexpected sink falls back with a warn log.
* Bytes written at close time (remaining inline chunks, root index, file info,
trailer) are not accounted — the existing accounting does not see them either.
* Input-side load (reading and decompressing the source storefiles) is not
accounted; for major compactions input ≈ output on disk, so the practical gap
is small.
* A side benefit: the controller's finish log ("... average throughput is
...") then reports disk MB/s, which is directly comparable with device
throughput.
With the option enabled, the controller charges only what is actually written:
compactions of compressible stores can use the full configured disk budget, so
their duration becomes proportional to their on-disk size, and stores with
incompressible payloads behave exactly as before — output accounting only
differs where encoding/compression does. The implementation has been running in
production with no regressions observed.
A PR against master follows.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)