[
https://issues.apache.org/jira/browse/THRIFT-6392?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jens Geyer resolved THRIFT-6392.
--------------------------------
Fix Version/s: 0.26.0
Assignee: Sylwester Lachiewicz
Resolution: Fixed
> D client_pool_test occasionally hangs until the lib-d job times out
> -------------------------------------------------------------------
>
> Key: THRIFT-6392
> URL: https://issues.apache.org/jira/browse/THRIFT-6392
> Project: Thrift
> Issue Type: Bug
> Components: D - Library
> Reporter: Sylwester Lachiewicz
> Assignee: Sylwester Lachiewicz
> Priority: Minor
> Fix For: 0.26.0
>
> Time Spent: 20m
> Remaining Estimate: 0h
>
> {{lib/d/test/client_pool_test}} occasionally never finishes, so the lib-d job
> runs into its 60-minute limit. It happened twice in a row on [PR
> #3971|https://github.com/apache/thrift/pull/3971], which does not touch lib/d
> ([job 1|https://github.com/apache/thrift/actions/runs/36332377379], [job
> 2|https://github.com/apache/thrift/actions/runs/36337045757/job/108670208530]),
> while {{make -C lib/d check}} took about 2.5 minutes on the other PRs of the
> same day and on master.
> In both hung runs all 104 unit tests pass, then each of the six server
> threads logs, three seconds later:
> {noformat}
> src/thrift/server/simple.d:144: Client died unexpectedly:
> thrift.transport.base.TTransportException@src/thrift/transport/socket.d(340):
> Timed out
> thrift.transport.socket.TSocket.read(ubyte[])
> thrift.transport.buffered.TBufferedTransport.peek()
>
> thrift.server.simple.TSimpleServer.serve(thrift.util.cancellation.TCancellation)
> client_pool_test.ServerThread.run()
> {noformat}
> and nothing else is printed until the job is cancelled; the runner then kills
> an orphaned {{client_pool_test}}. A passing run prints none of these
> messages. So the test's own client side stopped while holding a connection to
> every server, and the servers' 3-second {{recvTimeout}} is the only timeout
> involved.
> The client side of the test has no bound anywhere, so any lost reply waits
> forever:
> * the synchronous clients are {{TSocket}}s without {{recvTimeout}};
> * the asynchronous tests use {{waitGet()}}, and the implicit {{waitGet}} of
> {{TFuture}}, instead of {{waitGet(Duration)}};
> * {{main}} ignores the result of {{sem.wait(dur!"seconds"(1))}}, so it goes
> on even if a server is not listening yet.
> What sets it off is not known yet. Bounding those waits would turn the hang
> into a test failure that names the call, and a step-level {{timeout-minutes}}
> on "Run make check for d" in build.yml would stop a hang well before the job
> limit. Related: THRIFT-4155.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)