Jeongmin Kim created HBASE-30459:
------------------------------------

             Summary: Add an option to account compaction throughput control by 
actual output (post-compression) bytes
                 Key: HBASE-30459
                 URL: https://issues.apache.org/jira/browse/HBASE-30459
             Project: HBase
          Issue Type: Improvement
          Components: Compaction
            Reporter: Jeongmin Kim


Compaction throughput limits are consumed in {{Compactor#performCompaction}} by 
{{cell.getSerializedSize()}} — the pre-encoding, pre-compression size of each 
cell — while the bytes actually written to disk are post-DATA_BLOCK_ENCODING / 
post-compression. Under the same limit, a store whose data encodes/compresses R 
times runs its compactions at roughly 1/R of the configured disk write rate.

This works against both goals of throughput control: keeping compactions inside 
a predictable resource envelope, and letting them actually use the resources 
that envelope grants. The better the encoding and compression work, the more of 
the budget is charged for bytes that never reach the disk, and the compaction 
sleeps instead of writing. In practice — we observed this in production — a 
compaction of a highly compressible store spends most of its wall time in 
throttle sleep while the disk stays nearly idle: even a compaction that is 
small on disk takes a long time, occupies a longCompactions thread throughout, 
and the compaction queue backs up. And since limits are per-RegionServer (there 
is no per-table or per-family limit), stores holding incompressible payloads 
consume the same budget 1:1, so mixed workloads additionally become unfair 
between families.

Proposal: an opt-in configuration 
{{hbase.hstore.compaction.throughput.control.by.output}} (default 
{{{}false{}}}, current behavior unchanged). When enabled and the sink exposes 
its output position, {{performCompaction}} calls 
{{ThroughputController#control}} with the delta of the writer's output position 
after each cell-batch append, instead of the cells' serialized sizes. Notes on 
the implementation we have been running:
 * {{StoreFileWriter#getPos}} also includes the historical file writer when 
{{{}hbase.enable.historical.compaction.files=true{}}}, so the whole disk write 
load is accounted; {{AbstractMultiFileWriter}} gains a {{getPos}} summing its 
lower writers (stripe / date-tiered compactors).
 * The position only advances when a block is encoded/compressed and flushed to 
the output stream, so the controller sees actual disk bytes. The control check 
interval ({{{}controlPerSize{}}}, by default the throughput lower bound) is 
much larger than block sizes (32K ~ 128K), so block-granularity jumps are 
absorbed by the accounting.
 * Progress accounting, the shipped() cadence and {{CloseChecker}} stay on cell 
serialized sizes — only the throughput accounting unit changes.
 * Sinks that do not expose an output position (e.g. 
{{{}DefaultMobStoreCompactor{}}}, which overrides {{performCompaction}} anyway) 
keep the existing accounting; an unexpected sink falls back with a warn log.
 * Bytes written at close time (remaining inline chunks, root index, file info, 
trailer) are not accounted — the existing accounting does not see them either.
 * Input-side load (reading and decompressing the source storefiles) is not 
accounted; for major compactions input ≈ output on disk, so the practical gap 
is small.
 * A side benefit: the controller's finish log ("... average throughput is 
...") then reports disk MB/s, which is directly comparable with device 
throughput.

With the option enabled, the controller charges only what is actually written: 
compactions of compressible stores can use the full configured disk budget, so 
their duration becomes proportional to their on-disk size, and stores with 
incompressible payloads behave exactly as before — output accounting only 
differs where encoding/compression does. The implementation has been running in 
production with no regressions observed.

A PR against master follows.

 



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to