[ 
https://issues.apache.org/jira/browse/HBASE-30461?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jose Luis López updated HBASE-30461:
------------------------------------
    Status: Patch Available  (was: Open)

> AssignRegionHandler aborts the RegionServer when a requested stop races a 
> region open
> -------------------------------------------------------------------------------------
>
>                 Key: HBASE-30461
>                 URL: https://issues.apache.org/jira/browse/HBASE-30461
>             Project: HBase
>          Issue Type: Bug
>          Components: Region Assignment, regionserver
>    Affects Versions: 2.6.7, 2.6.3, 3.0.0, 4.0.0-alpha-1
>            Reporter: Jose Luis López
>            Priority: Minor
>              Labels: pull-request-available
>
> When a RegionServer is asked to stop while an AssignRegionHandler is opening
> a region, the handler aborts the RegionServer, turning an orderly shutdown
> into an abort.
> HRegionServer#stop sets stopped=true immediately. An AssignRegionHandler that
> has already called HRegion.openHRegion then calls postOpenDeployTasks, whose
> first step is RSRpcServices#checkOpen:
>     if (server.isStopped()) {
>       throw new RegionServerStoppedException("Server ... stopping");
>     }
> Because this happens after the "PONR" comment in AssignRegionHandler#process,
> the exception reaches handleException, which aborts the server:
>     getServer().abort("Failed to open region ... and can not recover", t);
> The window is the whole region open, which can take seconds when it has to
> replay recovered edits, so an ordinary "stop regionserver" or rolling restart
> that overlaps an assignment can hit it.
> Observed in a mini-cluster with HBase 2.6.3:
>   22:52:02.488  ***** STOPPING region server '...,42315,...' *****
>   22:52:02.494  Opened f6f5937a47a606f0399f03df7faef64f
>   22:52:02.498  AssignRegionHandler: Fatal error occurred while opening
>                 region ..., aborting...
>                 RegionServerStoppedException: Server ... stopping
>                   at RSRpcServices.checkOpen(RSRpcServices.java:1571)
>                   at 
> HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
>                   at AssignRegionHandler.process(AssignRegionHandler.java:161)
>   22:52:02.501  ***** ABORTING region server ...: Failed to open region ...
>                 and can not recover *****
> Consequences of the abort, compared with the stop that was requested:
> - FATAL logging and reportRSFatalError to the master.
> - run() calls shutdownWAL(!abortRequested), so the WAL is not closed cleanly.
> - The abort timer is armed (see HBASE-30460, where it later halted a JVM
>   hosting a mini-cluster).
> - The process exits with an error ("HRegionServer Aborted").
> Separately, the region opened by the handler is leaked: addRegion() comes
> after postOpenDeployTasks(), so the HRegion is never in onlineRegions and
> closeUserRegions() never closes it. Its submittedRegionProcedures entry is
> not removed either (compare HBASE-29660). The deprecated OpenRegionHandler
> closes the region on this path; AssignRegionHandler does not.
> Aborting adds no safety here: checkOpen fails before the OPENED transition
> is reported, so the master still considers the region OPENING, and the
> region was never online, so it holds no edits. When the stopping server
> expires, the master's ServerCrashProcedure reassigns the region, just as it
> would after the abort.
> Proposed fix: in AssignRegionHandler#process, when postOpenDeployTasks throws
> RegionServerStoppedException (but not RegionServerAbortedException) and the
> server has been stopped, close the region that was just opened, remove it
> from regionsInTransitionInRS and submittedRegionProcedures, log at INFO, and
> return without aborting. All other failures after the PONR keep aborting.
> Test: a RegionObserver whose postOpen blocks on a latch holds a region open
> in HRegion.openHRegion; the test calls stop() on the RegionServer, releases
> the latch, waits for the RegionServer thread to end, and asserts that it was
> not aborted and that the region was closed.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to