[ 
https://issues.apache.org/jira/browse/HBASE-30459?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated HBASE-30459:
-----------------------------------
    Labels: pull-request-available  (was: )

> Add an option to account compaction throughput control by actual output 
> (post-compression) bytes
> ------------------------------------------------------------------------------------------------
>
>                 Key: HBASE-30459
>                 URL: https://issues.apache.org/jira/browse/HBASE-30459
>             Project: HBase
>          Issue Type: Improvement
>          Components: Compaction
>            Reporter: Jeongmin Kim
>            Priority: Major
>              Labels: pull-request-available
>
> Compaction throughput limits are consumed in {{Compactor#performCompaction}} 
> by {{cell.getSerializedSize()}} — the pre-encoding, pre-compression size of 
> each cell — while the bytes actually written to disk are 
> post-DATA_BLOCK_ENCODING / post-compression. Under the same limit, a store 
> whose data encodes/compresses R times runs its compactions at roughly 1/R of 
> the configured disk write rate.
> This works against both goals of throughput control: keeping compactions 
> inside a predictable resource envelope, and letting them actually use the 
> resources that envelope grants. The better the encoding and compression work, 
> the more of the budget is charged for bytes that never reach the disk, and 
> the compaction sleeps instead of writing. In practice — we observed this in 
> production — a compaction of a highly compressible store spends most of its 
> wall time in throttle sleep while the disk stays nearly idle: even a 
> compaction that is small on disk takes a long time, occupies a 
> longCompactions thread throughout, and the compaction queue backs up. And 
> since limits are per-RegionServer (there is no per-table or per-family 
> limit), stores holding incompressible payloads consume the same budget 1:1, 
> so mixed workloads additionally become unfair between families.
> Proposal: an opt-in configuration 
> {{hbase.hstore.compaction.throughput.control.by.output}} (default 
> {{{}false{}}}, current behavior unchanged). When enabled and the sink exposes 
> its output position, {{performCompaction}} calls 
> {{ThroughputController#control}} with the delta of the writer's output 
> position after each cell-batch append, instead of the cells' serialized 
> sizes. Notes on the implementation we have been running:
>  * {{StoreFileWriter#getPos}} also includes the historical file writer when 
> {{{}hbase.enable.historical.compaction.files=true{}}}, so the whole disk 
> write load is accounted; {{AbstractMultiFileWriter}} gains a {{getPos}} 
> summing its lower writers (stripe / date-tiered compactors).
>  * The position only advances when a block is encoded/compressed and flushed 
> to the output stream, so the controller sees actual disk bytes. The control 
> check interval ({{{}controlPerSize{}}}, by default the throughput lower 
> bound) is much larger than block sizes (32K ~ 128K), so block-granularity 
> jumps are absorbed by the accounting.
>  * Progress accounting, the shipped() cadence and {{CloseChecker}} stay on 
> cell serialized sizes — only the throughput accounting unit changes.
>  * Sinks that do not expose an output position (e.g. 
> {{{}DefaultMobStoreCompactor{}}}, which overrides {{performCompaction}} 
> anyway) keep the existing accounting; an unexpected sink falls back with a 
> warn log.
>  * Bytes written at close time (remaining inline chunks, root index, file 
> info, trailer) are not accounted — the existing accounting does not see them 
> either.
>  * Input-side load (reading and decompressing the source storefiles) is not 
> accounted; for major compactions input ≈ output on disk, so the practical gap 
> is small.
>  * A side benefit: the controller's finish log ("... average throughput is 
> ...") then reports disk MB/s, which is directly comparable with device 
> throughput.
> With the option enabled, the controller charges only what is actually 
> written: compactions of compressible stores can use the full configured disk 
> budget, so their duration becomes proportional to their on-disk size, and 
> stores with incompressible payloads behave exactly as before — output 
> accounting only differs where encoding/compression does. The implementation 
> has been running in production with no regressions observed.
> A PR against master follows.
>  



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to