[
https://issues.apache.org/jira/browse/HBASE-30460?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jose Luis López updated HBASE-30460:
------------------------------------
Affects Version/s: 2.6.7
2.6.3
3.0.0
4.0.0-alpha-1
> RegionServer abort timer is never cancelled and halts the JVM after the
> RegionServer has shut down
> --------------------------------------------------------------------------------------------------
>
> Key: HBASE-30460
> URL: https://issues.apache.org/jira/browse/HBASE-30460
> Project: HBase
> Issue Type: Bug
> Components: regionserver
> Affects Versions: 3.0.0, 4.0.0-alpha-1, 2.6.3, 2.6.7
> Reporter: Jose Luis López
> Priority: Minor
>
> HRegionServer#abort calls scheduleAbortTimer(), which creates the
> "Abort regionserver monitor" Timer and schedules SystemExitWhenAbortTimeout
> to run after hbase.regionserver.abort.timeout (default 1200000 ms). The task
> prints a thread dump headed "Zombie HRegionServer" and calls
> Runtime.getRuntime().halt(1).
> The timer is never cancelled. It fires even when the abort completes
> normally and HRegionServer#run has returned. As introduced by HBASE-21325
> ("Force to terminate regionserver when abort hang in somewhere") and
> switched to halt by HBASE-21932, its purpose is to cap an abort that hangs.
> After a clean exit there is nothing left to cap.
> In a standalone RegionServer process this goes unnoticed, because the
> process exits once run() returns. When the RegionServer runs inside a JVM
> that outlives it, as with HBaseTestingUtility / MiniHBaseCluster, the timer
> halts that JVM 20 minutes after an abort that had already finished. Neither
> HBaseTestingUtility nor MiniHBaseCluster sets
> hbase.regionserver.abort.timeout or hbase.regionserver.abort.timeout.task.
> Observed downstream in Apache Hadoop, whose
> hadoop-yarn-server-timelineservice-hbase-tests module runs a mini-cluster
> in the Maven JVM (surefire forkCount=0), with HBase 2.6.3-hadoop3:
> 22:52:02.488 main: ***** STOPPING region server '…,42315,…' *****
> (HBaseTestingUtility#shutdownMiniCluster)
> 22:52:02.494 RS_OPEN_REGION: Opened f6f5937a47a606f0399f03df7faef64f
> 22:52:02.498 AssignRegionHandler: Fatal error occurred while opening
> region …, aborting…
> RegionServerStoppedException: Server … stopping
> at RSRpcServices.checkOpen(RSRpcServices.java:1571)
> at
> HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
> at AssignRegionHandler.process(AssignRegionHandler.java:161)
> 22:52:02.501 ***** ABORTING region server …,42315,…: Failed to open
> region … and can not recover *****
> 22:52:02.825 RS:0;…:42315 Exiting; stopping=…,42315,…; zookeeper
> connection closed. <- run() returned, shutDown=true
> 22:52:02.826 JVMClusterUtil: Shutdown of 1 master(s) and 1
> regionserver(s) complete
> …
> 23:12:02.528 Process Thread Dump: Zombie HRegionServer
> -> Runtime.halt(1); the whole Maven build exits with code 1
> The RegionServer had exited 20 minutes before it was declared a zombie.
> Proposed fix: cancel abortMonitor once the RegionServer has finished
> shutting down, e.g. at the end of HRegionServer#run after shutDown is set:
> if (abortMonitor != null) {
> abortMonitor.cancel();
> }
> An abort that really hangs never reaches that point, so the timer still
> fires in the case it exists for. A test can use
> hbase.regionserver.abort.timeout.task to install a task that records
> whether it ran. With a short hbase.regionserver.abort.timeout, it can abort
> a RegionServer in a mini-cluster, wait for the RegionServer thread to end,
> and assert the task never ran.
> Related, possibly a separate issue: AssignRegionHandler#handleException
> aborts the server when postOpenDeployTasks fails with
> RegionServerStoppedException because a requested stop is already under
> way. That turns an ordinary shutdown racing a region open into an abort,
> which is what armed the timer here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)