The dind entry landed ahead of the postgres one, and the postgres entry I added
earlier was numbered 7 while 'Verify the whole loop' already was. Now 7
postgres, 8 dind, 9 verify.
dind is killed by a probe whose only job is to check a socket exists, with a
1s timeout the nodes cannot always meet: 27 failures over 156 minutes.
Worth recording because it probably explains build failures already attributed
to something else. `DeadlineExceeded: no active session` was blamed on CPU
starvation and addressed by dropping runner capacity to 1; the likelier
mechanism is kubelet killing dind mid-build and taking the buildkit session
with it. Capacity reduced the load that trips the probe, which fits #17 passing
and #18-#21 failing anyway.
Written as a hypothesis with the command to confirm it, not as a conclusion.
The chart exposes no probe knobs, so the candidate fix is a Kyverno mutation in
Ansible, following the existing force-best-effort-cpu precedent.
The documented scale 0/1 recovery did not work this time: the volume returned
detached/faulted and refused to attach, so the pod sat in ContainerCreating.
auto-salvage: true cannot rescue it, because salvage happens during attach and
a faulted volume never gets that far.
The replica's failedAt timestamp is the only thing holding it faulted.
Clearing it brought the volume to attached/healthy and postgres to 1/1 in
under a minute, with repo data and CI history intact.
Adds the backup check first, which matters more now that Longhorn runs at one
replica and there is no second copy to fall back on.
The previous entry stated that restarting act_runner leaves every in-flight job
orphaned. That is wrong: run #16 had its remaining jobs re-dispatched to the
new pod and finished normally, while run #14 really was left stuck. Both
outcomes happen, so in_progress after a restart is ambiguous and the runner
logs are what settle it.
Also corrects the recovery advice. Gitea 1.26 has no cancel endpoint anywhere
in its swagger, and DELETE on a run returned 204 against a live run without
stopping it. The UI button is the only way to cancel.
Restarting act_runner leaves its in-flight jobs in_progress with nothing behind
them, and at capacity 1 one orphan blocks every later run. Gitea 1.26 has no
cancel endpoint; DELETE .../actions/runs/{index} does the job and takes the run
index, not the database id.
Every job was failing at Set up job after ~14 minutes, then the first real step died
at 0s. It read like a broken action; it was an empty dind image cache re-pulling the
1.6GB act job image on every run.
dind has no volume for /var/lib/docker, so the cache lives in its writable layer and
dies with every pod restart. A restart-looping runner therefore never keeps one.
Warming it by hand took lint from failure to success with no code change.
trivy on the worker image: 39 findings (2 CRITICAL) -> 18 -> 5 -> 0. Verified in CI
run #7 on 4a426db: 'Total: 0 (HIGH: 0, CRITICAL: 0)' for all three images.
The last five lived in kubectl's vendored golang.org/x/net and Go stdlib, inside the
newest kubectl published. No version cleared them; removing the binary did.
CORRECTNESS
- lost-lease race: complete()/fail() did not check ownership, so a worker whose
lease expired could mark a task done while another worker was running it, or
requeue a task someone else owned. Reproduced, fixed with a CAS on
(state, locked_by), pinned by two regression tests.
- worker died on report failure: _run_one's docstring claimed no exception
escapes the TaskGroup; fail()/complete() were outside the guarded block, so a
DB blip cancelled every sibling provision on the pod.
- claim query used an INNER join, which could strand a just-claimed task and
report 'queue empty'. LEFT join.
- InstanceRepo.set_error bypassed the state machine and had no callers. Deleted.
- handle_deprovision ignored its CAS result, so a wrong-state instance kept a
dangling endpoint and got re-provisioned by the drift check 60s later.
- handle_verify re-notified on every retry: five pages for one halt.
DEPLOY-BREAKING
- the migration Job could never succeed: no Dockerfile copied migrations/, and
migrate.py resolved the path relative to the source tree, which only works for
an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR.
- ServiceMonitor selector did not match the Service: API metrics never scraped.
- SvcforgeReconcilerStale fired permanently from every pod, because the gauge is
module-level and every service exports it as 0. Scoped to the reconciler job.
- SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m].
- the digest guard accepted the all-zeros placeholder.
- worker terminationGracePeriodSeconds was 60s against a 600s helm timeout.
DEAD CODE THAT SHOULD NOT HAVE BEEN
- adapters/k8s.py was never called, so tenant namespaces were never created and
the first provision for a new team would fail. Wired into handle_provision.
- adapters/redis.py was never imported by any service. Rate limiting is now wired
into the API, failing open.
- Settings.check_production() had no callers. Given an explicit environment and
called from every entrypoint.
OBSERVABILITY
- the API never called obs.setup(): no JSON logs, no trace correlation, log_json
silently inert.
- LogNotifier's structured fields were discarded by the stdlib->structlog bridge.
- bind_task_context cleared the 'service' binding for the life of every task.
- split tasks_failed into task_attempts_failed and tasks_dead_lettered.
SECURITY
- trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl
1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3,
which also closes a four-minor skew against the v1.35.3 cluster.
TESTS THAT COULD NOT FAIL
- the concurrency cap test passed on a fully serial worker.
- the alert/metric cross-check asserted a hardcoded list instead of reading the
chart, so it could not catch a rename on the chart side.
- fixed OTel tracer-provider pollution between test files.
DOCS
- ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD
guarantee (verified with --dry-run=server).
- AGENTS.md + CLAUDE.md.
- prose sweep for back-and-forth phrasing across 19 files.
trivy-action@v0.29.0 internally uses setup-trivy@v0.2.2, a tag removed
upstream (earliest published is now v0.2.6), so it cannot resolve on any
runner. Run trivy from a digest-pinned image instead, as gitleaks already is.
RUNBOOK gains a 'Setting up CI/CD from scratch' section with the traps that
actually cost time: GITEA_TOKEN is 401 at the package registry, the runner's
cache fails soft, service containers resolve by name not localhost.
Not pushed: pushing triggers a run, and the runner is being restarted by the
ansible change that enables its cache.