[ 
https://issues.apache.org/jira/browse/HBASE-30460?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Jose Luis López updated HBASE-30460:
------------------------------------
    Affects Version/s: 2.6.7
                       2.6.3
                       3.0.0
                       4.0.0-alpha-1

> RegionServer abort timer is never cancelled and halts the JVM after the 
> RegionServer has shut down
> --------------------------------------------------------------------------------------------------
>
>                 Key: HBASE-30460
>                 URL: https://issues.apache.org/jira/browse/HBASE-30460
>             Project: HBase
>          Issue Type: Bug
>          Components: regionserver
>    Affects Versions: 3.0.0, 4.0.0-alpha-1, 2.6.3, 2.6.7
>            Reporter: Jose Luis López
>            Priority: Minor
>
> HRegionServer#abort calls scheduleAbortTimer(), which creates the
> "Abort regionserver monitor" Timer and schedules SystemExitWhenAbortTimeout
> to run after hbase.regionserver.abort.timeout (default 1200000 ms). The task
> prints a thread dump headed "Zombie HRegionServer" and calls
> Runtime.getRuntime().halt(1).
> The timer is never cancelled. It fires even when the abort completes
> normally and HRegionServer#run has returned. As introduced by HBASE-21325
> ("Force to terminate regionserver when abort hang in somewhere") and
> switched to halt by HBASE-21932, its purpose is to cap an abort that hangs.
> After a clean exit there is nothing left to cap.
> In a standalone RegionServer process this goes unnoticed, because the
> process exits once run() returns. When the RegionServer runs inside a JVM
> that outlives it, as with HBaseTestingUtility / MiniHBaseCluster, the timer
> halts that JVM 20 minutes after an abort that had already finished. Neither
> HBaseTestingUtility nor MiniHBaseCluster sets
> hbase.regionserver.abort.timeout or hbase.regionserver.abort.timeout.task.
> Observed downstream in Apache Hadoop, whose
> hadoop-yarn-server-timelineservice-hbase-tests module runs a mini-cluster
> in the Maven JVM (surefire forkCount=0), with HBase 2.6.3-hadoop3:
>   22:52:02.488  main: ***** STOPPING region server '…,42315,…' *****
>                 (HBaseTestingUtility#shutdownMiniCluster)
>   22:52:02.494  RS_OPEN_REGION: Opened f6f5937a47a606f0399f03df7faef64f
>   22:52:02.498  AssignRegionHandler: Fatal error occurred while opening
>                 region …, aborting…
>                 RegionServerStoppedException: Server … stopping
>                   at RSRpcServices.checkOpen(RSRpcServices.java:1571)
>                   at 
> HRegionServer.postOpenDeployTasks(HRegionServer.java:2608)
>                   at AssignRegionHandler.process(AssignRegionHandler.java:161)
>   22:52:02.501  ***** ABORTING region server …,42315,…: Failed to open
>                 region … and can not recover *****
>   22:52:02.825  RS:0;…:42315 Exiting; stopping=…,42315,…; zookeeper
>                 connection closed.        <- run() returned, shutDown=true
>   22:52:02.826  JVMClusterUtil: Shutdown of 1 master(s) and 1
>                 regionserver(s) complete
>   …
>   23:12:02.528  Process Thread Dump: Zombie HRegionServer
>                 -> Runtime.halt(1); the whole Maven build exits with code 1
> The RegionServer had exited 20 minutes before it was declared a zombie.
> Proposed fix: cancel abortMonitor once the RegionServer has finished
> shutting down, e.g. at the end of HRegionServer#run after shutDown is set:
>     if (abortMonitor != null) {
>       abortMonitor.cancel();
>     }
> An abort that really hangs never reaches that point, so the timer still
> fires in the case it exists for. A test can use
> hbase.regionserver.abort.timeout.task to install a task that records
> whether it ran. With a short hbase.regionserver.abort.timeout, it can abort
> a RegionServer in a mini-cluster, wait for the RegionServer thread to end,
> and assert the task never ran.
> Related, possibly a separate issue: AssignRegionHandler#handleException
> aborts the server when postOpenDeployTasks fails with
> RegionServerStoppedException because a requested stop is already under
> way. That turns an ordinary shutdown racing a region open into an abort,
> which is what armed the timer here.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to