runbook: the 14-minute 'Set up job' failure and its cause
ci / lint (push) Waiting to run
ci / types (push) Blocked by required conditions
ci / unit (push) Blocked by required conditions
ci / integration (push) Blocked by required conditions
ci / security (push) Blocked by required conditions
ci / dockerfile (push) Blocked by required conditions
ci / chart (push) Blocked by required conditions
ci / image (api) (push) Blocked by required conditions
ci / image (reconciler) (push) Blocked by required conditions
ci / image (worker) (push) Blocked by required conditions
ci / bump (push) Blocked by required conditions

Every job was failing at Set up job after ~14 minutes, then the first real step died
at 0s. It read like a broken action; it was an empty dind image cache re-pulling the
1.6GB act job image on every run.

dind has no volume for /var/lib/docker, so the cache lives in its writable layer and
dies with every pod restart. A restart-looping runner therefore never keeps one.
Warming it by hand took lint from failure to success with no code change.
This commit is contained in:
Nguyen Minh Phuc
2026-07-20 03:35:12 +00:00
parent c537073c21
commit f87d8d4d78
+28 -1
View File
@@ -97,7 +97,34 @@ kubectl's vendored `golang.org/x/net` and Go stdlib, inside the newest kubectl t
no version cleared them. kubectl was in that image for exactly one call, and helm already
does the same thing with a flag. Removing it removed the CVEs, a binary, and an adapter.
### 5. Verify the whole loop, not just the green checkmarks
### 5. When every job fails at `Set up job` after ~14 minutes
Symptom: `Set up job` runs for 1015 minutes and succeeds, then the first real step fails
instantly at 0s, and every downstream job is skipped. It looks like the action broke.
Cause: the dind sidecar's image cache is empty, so each run re-pulls the ~1.6GB act job
image (`ghcr.io/catthehacker/ubuntu:act-24.04`) before it can start. dind has no volume for
`/var/lib/docker` — the cache lives in its container writable layer and is destroyed on
every pod restart. A runner that is restart-looping therefore never keeps a cache, and each
job pays the full pull.
Check it:
```bash
kubectl -n gitea exec gitea-actions-runner-0 -c dind -- docker images
```
Empty output is the diagnosis. Warm it once by hand:
```bash
kubectl -n gitea exec gitea-actions-runner-0 -c dind -- docker pull \
ghcr.io/catthehacker/ubuntu:act-24.04@sha256:c710431fbad9eb3bcb102d04e5ff74fbd0ce6e383f78afebfb3770a1a817fdf9
```
The durable fix is to stop the runner restarting. Its `/data` PVC is ReadWriteOnce, so
every reschedule hits `Multi-Attach error` and the pod sits in Init until Longhorn detaches
from the old node. It is pinned to node2 in `oci-k8s/.../addons/tasks/main.yml` for exactly
that reason. A dedicated PVC for the image cache would survive restarts outright, but on
this cluster that volume faulted and blocked the runner, so it is deliberately not used.
### 6. Verify the whole loop, not just the green checkmarks
```bash
# the digest CI pushed