lymerin opened a new issue, #7379: URL: https://github.com/apache/shenyu/issues/7379
### Is there an existing issue for this? - [x] I have searched the existing issues ### Current Behavior The Install k8s step in k8s-examples-http can reach cat /etc/rancher/k3s/k3s.yaml without retrying, then fail because the kubeconfig does not exist. In [this failed job](https://github.com/apache/shenyu/actions/runs/36542805118/job/109321997957#step:7:23), the step took about one second. Its log contains neither output from the k3s installer nor a k3s install failed on attempt message. It ends with: cat: /etc/rancher/k3s/k3s.yaml: No such file or directory The [workflow step](https://github.com/apache/shenyu/blob/86581fc24b504bfc1550595d421101902cc4dfe5/.github/workflows/k8s-examples-http.yml#L111-L132) breaks out of the retry loop when install_k3s returns zero. There are two ways this can happen without producing a kubeconfig: The step uses curl -sfL ... | sh - without specifying shell: bash. On Linux, GitHub Actions runs an unspecified shell as bash -e {0}, without pipefail. If curl fails without writing a script, sh - can exit successfully on empty input, making the pipeline return zero. [GitHub Actions shell documentation](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax#jobsjob_idstepsshell) The loop does not check for /etc/rancher/k3s/k3s.yaml before treating an installation attempt as successful. An installer that returns zero without producing that file also skips the remaining attempts. The CI log is consistent with the first path, but it does not reveal why curl failed, or conclusively distinguish a download failure from another installation failure. The offline reproduction below demonstrates both script defects. ### Expected Behavior A failed download or installer should trigger the existing three-attempt retry loop. The loop should proceed only after the install command succeeds and the kubeconfig exists. If all attempts fail, the step should report an installation failure; a curl error should remain visible in the log. ### Steps To Reproduce Run the following script with Bash. It is fully offline: the curl command is replaced with a temporary stub, no k3s installation occurs, and the kubeconfig path points to a temporary directory. It does not require root or read the machine's real kubeconfig. #!/usr/bin/env bash set -u tmp=$(mktemp -d) mkdir -p "$tmp/bin" "$tmp/home" printf 'Temporary sandbox: %s\n' "$tmp" cat > "$tmp/bin/curl" <<'STUB' #!/bin/sh case "$REPRO_CASE" in curl_fails) exit 22 ;; no_kubeconfig) printf 'echo "[INFO] simulated installer completed"\nexit 0\n' ;; esac STUB chmod +x "$tmp/bin/curl" for case in curl_fails no_kubeconfig; do printf '\nCase: %s\n' "$case" REPRO_CASE="$case" HOME="$tmp/home" \ KUBECONFIG_FILE="$tmp/k3s.yaml" PATH="$tmp/bin:$PATH" \ bash -e <<'STEP' install_k3s() { curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=v1.29.6+k3s2 K3S_KUBECONFIG_MODE=777 sh - } for attempt in 1 2 3; do if install_k3s; then break fi if [ "$attempt" = 3 ]; then echo "k3s install failed after $attempt attempts" exit 1 fi echo "k3s install failed on attempt $attempt" sleep $((attempt * 15)) done cat "$KUBECONFIG_FILE" STEP printf 'Exit code: %s\n' "$?" done Observed output (the temporary path varies): Case: curl_fails cat: /tmp/tmp.XXXX/k3s.yaml: No such file or directory Exit code: 1 Case: no_kubeconfig [INFO] simulated installer completed cat: /tmp/tmp.XXXX/k3s.yaml: No such file or directory Exit code: 1 Neither case prints an attempt-failed message or waits before trying to read the missing file. To compare the intended behavior, use bash -eo pipefail and require the kubeconfig to exist in the if install_k3s condition; both cases should then reach all three attempts and fail with k3s install failed after 3 attempts. ### Environment ```markdown ShenYu workflow code: master commit 86581fc24b504bfc1550595d421101902cc4dfe5 Failing PR run: commit 2f0d8ac60; its Install k8s step has the same relevant code Runner: ubuntu-latest k3s version requested by the workflow: v1.29.6+k3s2 Debug logs ``` ### Debug logs The [failed step](https://github.com/apache/shenyu/actions/runs/36542805118/job/109321997957#step:7:23) echoes the installation command at 2026-09-29 08:40:15.964 UTC and reports the missing kubeconfig at 08:40:16.058 UTC. There is no retry message between them. The full step duration is approximately one second. ### Anything else? uggested fix: Make pipeline failure observable, for example by setting shell: bash so the step runs with pipefail, or by downloading the installation script successfully before executing it. Use curl -sSfL to retain the error message while keeping normal output quiet. Require the installation command to succeed and /etc/rancher/k3s/k3s.yaml to exist before breaking out of the retry loop. PR verification note: The workflow's internal [path filter](https://github.com/apache/shenyu/blob/86581fc24b504bfc1550595d421101902cc4dfe5/.github/workflows/k8s-examples-http.yml#L80-L107) excludes .github/**; a PR that changes only this workflow may therefore skip the Install k8s step. Verification should explicitly demonstrate that the changed step ran, alongside the offline failure and success-path checks. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
