krylosov-aa opened a new pull request, #2054:
URL: https://github.com/apache/cloudberry/pull/2054

   ### What does this PR do?
   
   Problem: during teardown, the test restarts content 1 with `pg_ctl restart 
-w` and then immediately resets the `checkpoint` fault on all primaries.
   
   `pg_ctl -w` may return as soon as crash recovery starts, even if the segment 
is not ready to accept connections yet. `gp_inject_fault` connects directly to 
each segment and does not retry, so the reset can sometimes fail with an error 
like `connection to ... failed`.
   
   Regular queries do not have this issue because the dispatcher retries gang 
creation.
   
   Solution: reset the `checkpoint` fault before the second restart. The second 
restart is only used for teardown and does not need the fault. Also, content 1 
loses its fault state after the restart anyway.
   
   ### Type of Change
   - [x] Bug fix (non-breaking change)
   
   ### Test Plan
   Local checks on Linux ARM64
   
   ### Checklist
   - [x] Followed [contribution 
guide](https://cloudberry.apache.org/contribute/code)
   - [ ] Requested review from [cloudberry 
committers](https://github.com/orgs/apache/teams/cloudberry-committers)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to