lymerin opened a new issue, #7379:
URL: https://github.com/apache/shenyu/issues/7379

   ### Is there an existing issue for this?
   
   - [x] I have searched the existing issues
   
   ### Current Behavior
   
   The Install k8s step in k8s-examples-http can reach cat 
/etc/rancher/k3s/k3s.yaml without retrying, then fail because the kubeconfig 
does not exist.
   
   In [this failed 
job](https://github.com/apache/shenyu/actions/runs/36542805118/job/109321997957#step:7:23),
 the step took about one second. Its log contains neither output from the k3s 
installer nor a k3s install failed on attempt message. It ends with:
   
   cat: /etc/rancher/k3s/k3s.yaml: No such file or directory
   
   The [workflow 
step](https://github.com/apache/shenyu/blob/86581fc24b504bfc1550595d421101902cc4dfe5/.github/workflows/k8s-examples-http.yml#L111-L132)
 breaks out of the retry loop when install_k3s returns zero. There are two ways 
this can happen without producing a kubeconfig:
   
   The step uses curl -sfL ... | sh - without specifying shell: bash. On Linux, 
GitHub Actions runs an unspecified shell as bash -e {0}, without pipefail. If 
curl fails without writing a script, sh - can exit successfully on empty input, 
making the pipeline return zero. [GitHub Actions shell 
documentation](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax#jobsjob_idstepsshell)
   
   The loop does not check for /etc/rancher/k3s/k3s.yaml before treating an 
installation attempt as successful. An installer that returns zero without 
producing that file also skips the remaining attempts.
   
   The CI log is consistent with the first path, but it does not reveal why 
curl failed, or conclusively distinguish a download failure from another 
installation failure. The offline reproduction below demonstrates both script 
defects.
   
   ### Expected Behavior
   
   A failed download or installer should trigger the existing three-attempt 
retry loop. The loop should proceed only after the install command succeeds and 
the kubeconfig exists. If all attempts fail, the step should report an 
installation failure; a curl error should remain visible in the log.
   
   ### Steps To Reproduce
   
   Run the following script with Bash. It is fully offline: the curl command is 
replaced with a temporary stub, no k3s installation occurs, and the kubeconfig 
path points to a temporary directory. It does not require root or read the 
machine's real kubeconfig.
   
   #!/usr/bin/env bash
   set -u
   tmp=$(mktemp -d)
   mkdir -p "$tmp/bin" "$tmp/home"
   printf 'Temporary sandbox: %s\n' "$tmp"
   
   cat > "$tmp/bin/curl" <<'STUB'
   #!/bin/sh
   case "$REPRO_CASE" in
     curl_fails) exit 22 ;;
     no_kubeconfig) printf 'echo "[INFO] simulated installer completed"\nexit 
0\n' ;;
   esac
   STUB
   chmod +x "$tmp/bin/curl"
   
   for case in curl_fails no_kubeconfig; do
     printf '\nCase: %s\n' "$case"
     REPRO_CASE="$case" HOME="$tmp/home" \
       KUBECONFIG_FILE="$tmp/k3s.yaml" PATH="$tmp/bin:$PATH" \
       bash -e <<'STEP'
   install_k3s() {
     curl -sfL https://get.k3s.io |
       INSTALL_K3S_VERSION=v1.29.6+k3s2 K3S_KUBECONFIG_MODE=777 sh -
   }
   for attempt in 1 2 3; do
     if install_k3s; then
       break
     fi
     if [ "$attempt" = 3 ]; then
       echo "k3s install failed after $attempt attempts"
       exit 1
     fi
     echo "k3s install failed on attempt $attempt"
     sleep $((attempt * 15))
   done
   cat "$KUBECONFIG_FILE"
   STEP
     printf 'Exit code: %s\n' "$?"
   done
   
   Observed output (the temporary path varies):
   
   Case: curl_fails
   cat: /tmp/tmp.XXXX/k3s.yaml: No such file or directory
   Exit code: 1
   
   Case: no_kubeconfig
   [INFO] simulated installer completed
   cat: /tmp/tmp.XXXX/k3s.yaml: No such file or directory
   Exit code: 1
   
   Neither case prints an attempt-failed message or waits before trying to read 
the missing file. To compare the intended behavior, use bash -eo pipefail and 
require the kubeconfig to exist in the if install_k3s condition; both cases 
should then reach all three attempts and fail with k3s install failed after 3 
attempts.
   
   ### Environment
   
   ```markdown
   ShenYu workflow code: master commit 86581fc24b504bfc1550595d421101902cc4dfe5
   
   Failing PR run: commit 2f0d8ac60; its Install k8s step has the same relevant 
code
   
   Runner: ubuntu-latest
   
   k3s version requested by the workflow: v1.29.6+k3s2
   
   Debug logs
   ```
   
   ### Debug logs
   
   The [failed 
step](https://github.com/apache/shenyu/actions/runs/36542805118/job/109321997957#step:7:23)
 echoes the installation command at 2026-09-29 08:40:15.964 UTC and reports the 
missing kubeconfig at 08:40:16.058 UTC. There is no retry message between them. 
The full step duration is approximately one second.
   
   ### Anything else?
   
   uggested fix: Make pipeline failure observable, for example by setting 
shell: bash so the step runs with pipefail, or by downloading the installation 
script successfully before executing it. Use curl -sSfL to retain the error 
message while keeping normal output quiet. Require the installation command to 
succeed and /etc/rancher/k3s/k3s.yaml to exist before breaking out of the retry 
loop.
   
   PR verification note: The workflow's internal [path 
filter](https://github.com/apache/shenyu/blob/86581fc24b504bfc1550595d421101902cc4dfe5/.github/workflows/k8s-examples-http.yml#L80-L107)
 excludes .github/**; a PR that changes only this workflow may therefore skip 
the Install k8s step. Verification should explicitly demonstrate that the 
changed step ran, alongside the offline failure and success-path checks.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to