Jose Luis López created HBASE-30461:
---------------------------------------
Summary: AssignRegionHandler aborts the RegionServer when a
requested stop races a region open
Key: HBASE-30461
URL: https://issues.apache.org/jira/browse/HBASE-30461
Project: HBase
Issue Type: Bug
Components: Region Assignment, regionserver
Affects Versions: 2.6.7, 2.6.3, 3.0.0, 4.0.0-alpha-1
Reporter: Jose Luis López
When a RegionServer is asked to stop while an AssignRegionHandler is opening
a region, the handler aborts the RegionServer, turning an orderly shutdown
into an abort.
HRegionServer#stop sets stopped=true immediately. An AssignRegionHandler that
has already called HRegion.openHRegion then calls postOpenDeployTasks, whose
first step is RSRpcServices#checkOpen:
if (server.isStopped()) {
throw new RegionServerStoppedException("Server ... stopping");
}
Because this happens after the "PONR" comment in AssignRegionHandler#process,
the exception reaches handleException, which aborts the server:
getServer().abort("Failed to open region ... and can not recover", t);
The window is the whole region open, which can take seconds when it has to
replay recovered edits, so an ordinary "stop regionserver" or rolling restart
that overlaps an assignment can hit it.
Observed in a mini-cluster with HBase 2.6.3:
22:52:02.488 ***** STOPPING region server '...,42315,...' *****
22:52:02.494 Opened f6f5937a47a606f0399f03df7faef64f
22:52:02.498 AssignRegionHandler: Fatal error occurred while opening
region ..., aborting...
RegionServerStoppedException: Server ... stopping
at RSRpcServices.checkOpen(RSRpcServices.java:1571)
at HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
at AssignRegionHandler.process(AssignRegionHandler.java:161)
22:52:02.501 ***** ABORTING region server ...: Failed to open region ...
and can not recover *****
Consequences of the abort, compared with the stop that was requested:
- FATAL logging and reportRSFatalError to the master.
- run() calls shutdownWAL(!abortRequested), so the WAL is not closed cleanly.
- The abort timer is armed (see HBASE-30460, where it later halted a JVM
hosting a mini-cluster).
- The process exits with an error ("HRegionServer Aborted").
Separately, the region opened by the handler is leaked: addRegion() comes
after postOpenDeployTasks(), so the HRegion is never in onlineRegions and
closeUserRegions() never closes it. Its submittedRegionProcedures entry is
not removed either (compare HBASE-29660). The deprecated OpenRegionHandler
closes the region on this path; AssignRegionHandler does not.
Aborting adds no safety here: checkOpen fails before the OPENED transition
is reported, so the master still considers the region OPENING, and the
region was never online, so it holds no edits. When the stopping server
expires, the master's ServerCrashProcedure reassigns the region, just as it
would after the abort.
Proposed fix: in AssignRegionHandler#process, when postOpenDeployTasks throws
RegionServerStoppedException (but not RegionServerAbortedException) and the
server has been stopped, close the region that was just opened, remove it
from regionsInTransitionInRS and submittedRegionProcedures, log at INFO, and
return without aborting. All other failures after the PONR keep aborting.
Test: a RegionObserver whose postOpen blocks on a latch holds a region open
in HRegion.openHRegion; the test calls stop() on the RegionServer, releases
the latch, waits for the RegionServer thread to end, and asserts that it was
not aborted and that the region was closed.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)