[
https://issues.apache.org/jira/browse/THRIFT-6392?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119882#comment-18119882
]
Sylwester Lachiewicz commented on THRIFT-6392:
----------------------------------------------
Root cause: {{FastestPoolJob}} in
[lib/d/src/thrift/codegen/async_client_pool.d|https://github.com/apache/thrift/blob/master/lib/d/src/thrift/codegen/async_client_pool.d]
stops registering completion callbacks at the first child future that has
already completed, but a child that failed with an RPC fault does not set the
pool's result, so the remaining children are never counted and the pool's
future never completes. {{client_pool_test}} hits it in the all-clients-fail
case when the 1 ms server answers before the pool is constructed. Fix and a
unit test in [PR #3972|https://github.com/apache/thrift/pull/3972].
> D client_pool_test occasionally hangs until the lib-d job times out
> -------------------------------------------------------------------
>
> Key: THRIFT-6392
> URL: https://issues.apache.org/jira/browse/THRIFT-6392
> Project: Thrift
> Issue Type: Bug
> Components: D - Library
> Reporter: Sylwester Lachiewicz
> Priority: Minor
> Time Spent: 10m
> Remaining Estimate: 0h
>
> {{lib/d/test/client_pool_test}} occasionally never finishes, so the lib-d job
> runs into its 60-minute limit. It happened twice in a row on [PR
> #3971|https://github.com/apache/thrift/pull/3971], which does not touch lib/d
> ([job 1|https://github.com/apache/thrift/actions/runs/36332377379], [job
> 2|https://github.com/apache/thrift/actions/runs/36337045757/job/108670208530]),
> while {{make -C lib/d check}} took about 2.5 minutes on the other PRs of the
> same day and on master.
> In both hung runs all 104 unit tests pass, then each of the six server
> threads logs, three seconds later:
> {noformat}
> src/thrift/server/simple.d:144: Client died unexpectedly:
> thrift.transport.base.TTransportException@src/thrift/transport/socket.d(340):
> Timed out
> thrift.transport.socket.TSocket.read(ubyte[])
> thrift.transport.buffered.TBufferedTransport.peek()
>
> thrift.server.simple.TSimpleServer.serve(thrift.util.cancellation.TCancellation)
> client_pool_test.ServerThread.run()
> {noformat}
> and nothing else is printed until the job is cancelled; the runner then kills
> an orphaned {{client_pool_test}}. A passing run prints none of these
> messages. So the test's own client side stopped while holding a connection to
> every server, and the servers' 3-second {{recvTimeout}} is the only timeout
> involved.
> The client side of the test has no bound anywhere, so any lost reply waits
> forever:
> * the synchronous clients are {{TSocket}}s without {{recvTimeout}};
> * the asynchronous tests use {{waitGet()}}, and the implicit {{waitGet}} of
> {{TFuture}}, instead of {{waitGet(Duration)}};
> * {{main}} ignores the result of {{sem.wait(dur!"seconds"(1))}}, so it goes
> on even if a server is not listening yet.
> What sets it off is not known yet. Bounding those waits would turn the hang
> into a test failure that names the call, and a step-level {{timeout-minutes}}
> on "Run make check for d" in build.yml would stop a hang well before the job
> limit. Related: THRIFT-4155.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)