A full run under the new 600m cap, to confirm the checkout no longer loses its
race with the build for node1's single core, and that the resulting digest bump
deploys through the webhook without a nudge.
Registered as gogs type on all three ArgoCD-tracked repos, secret shared with
argocd-secret's webhook.gogs.secret. Records how to read a failed delivery from
both ends, and that a drifted secret fails validation silently -- which looks
exactly like having no webhook.
1b56cda blamed the Application-informer watch break after an API server kill.
That is not it. Watched a freshly restarted controller for 11 minutes on a fully
healthy API server:
03:37:03 startup refresh, 3 apps adds_total = 3
03:38:31 adds_total = 3
03:41:33 adds_total = 3
03:44:34 adds_total = 3
03:47:35 adds_total = 3 expiry is 2m0s, jitter 60s
The periodic poll does not fire, ever, with or without a prior disruption. The
68 adds the earlier pod accumulated were cluster events during a rollout, and it
went flat the moment the cluster quieted -- same shape, no watch break involved.
Refreshes come from startup and from cluster events only. A commit that changes
just the repo is never noticed. Every "auto-sync" seen on 2026-07-22 landed
within seconds of a controller restart, which is the startup refresh -- I read
one of those as proof the pipeline worked, and it was not.
Also records that no Gitea webhook exists, so nothing covers for the broken
poll. Wiring one is worth doing regardless of whether the poll is ever repaired.
Reproduced three times on 2026-07-22, and the controller's own workqueue
metrics say what it is:
workqueue_depth 0
workqueue_longest_running_processor_seconds 0
workqueue_adds_total 68 <- frozen
Empty queue, idle processors, no new adds. Nothing is blocked -- nothing is
being enqueued. The periodic app-resync timer stops firing, so no Application
is ever queued for refresh again.
Every occurrence followed an API server disruption, with the matching informer
line in the logs: watch ended with error, http2: client connection lost, on
*v1alpha1.Application. The informer re-lists; the resync does not resume.
That is why the controller reads healthy from every angle except reconcile
count, and why only a restart clears it. On a single-control-plane cluster
anything that kills kube-apiserver -- including memory reclaim stalling /livez
past its 80s threshold -- can silently stop all GitOps.
c691f4f told you to check argocd_redis_request_total and
workqueue_unfinished_work_seconds. Both read healthy through an 82-minute wedge
on 2026-07-22 -- 45 cache reads per 15m and a queue depth of zero, while
argocd_app_reconcile_count sat flat at 62 and three Applications were two
commits behind master.
The cache counter measures that the process is running. The queue gauge reads
zero because this wedge blocks goroutines outside the queue, so nothing is in
flight and nothing is being enqueued.
argocd_app_reconcile_count is the one that went flat, and it is the one to
check. The other two are now in the table of things that look like answers and
are not, with what each read during the outage.
The stuck-deploy section printed `.status.reconciledAt` as a diagnostic without
saying how to read it, which invites exactly the wrong conclusion: ArgoCD only
writes that field when the computed status changes, so on an idle cluster it
stops advancing while the controller is healthy. Measured 2026-07-22 -- six
samples 50s apart against a 120s reconciliation timeout, zero movement.
It only means something alongside a sync.revision that is behind master.
Replaces the "is it doing anything" step with two metrics that answer it
directly: the controller's Redis cache reads, which happen every refresh cycle
regardless of change, and the reconcile queue's unfinished-work seconds, which
is precisely what the 2026-07-21 webhook wedge looked like. Log-based check
kept as the no-Prometheus fallback.
Cost 11 hours yesterday and looked like nothing was wrong, because the
Application reported Synced the whole time — at a revision five commits
behind. Records the diagnosis that actually works (log-lines-per-hour, and
a flat goroutine count meaning blocked rather than idle), the Kyverno
failurePolicy root cause, and the two escalating fixes. Folds in the
hook-finalizer deadlock, which has now happened four times and was only
written down in chat.
The runner moved to node0 and the drift check reads release secrets via the
Kubernetes API in-cluster; the docs still described node2 and `helm list`.
- RUNBOOK: the durable image-cache fix is now the node0 hostPath store, not a
node2 pin; /data is NFS RWX, so the Multi-Attach wait is gone. Cross-references
entries 9 and 10.
- ARCHITECTURE: the reconciler's edge to the cluster is "list releases", not
"helm list" (helm is the out-of-cluster fallback).
- ARCHITECTURE: the state diagram and its prose described fail() moving every
dead-lettered instance to `failed`. Corrected to the per-kind behaviour — only
provision fails the instance; deprovision stays `deleting` for retry, upgrade
and verify stay `ready` — matching the fix in tasks.py.
Three caches on this runner were container-layer only, each found because
something was slow: dind's image store, the trivy vuln DB, and act's action
clones. All three now have real storage.
act clones actions with full history, not shallow — 66.7MB/538 commits for
setup-uv, 24.4MB/222 for actions/checkout, and this workflow uses five. After a
restart that made `Set up job` an 11-minute step with the job container sat
idle running `sleep` while the runner cloned GitHub. Includes the command to
tell those two apart.
Also records the cascade the dind fix creates. A persistent image store makes
dockerd scan on boot — 38s, 2m13s, or over 5 minutes depending on node load —
and two timeouts then fire: dind's startup probe, and the runner image's own
hardcoded `Docker wait timeout of 5m0s`, which the chart cannot configure. The
runner exits 1, restarts, and kills the running job, which looks like every
step failing at once with no error after a green `Set up job`.
It self-heals in about ten minutes at the cost of one CI run, so runner
restarts are now something to do deliberately rather than casually.
The dind entry landed ahead of the postgres one, and the postgres entry I added
earlier was numbered 7 while 'Verify the whole loop' already was. Now 7
postgres, 8 dind, 9 verify.
dind is killed by a probe whose only job is to check a socket exists, with a
1s timeout the nodes cannot always meet: 27 failures over 156 minutes.
Worth recording because it probably explains build failures already attributed
to something else. `DeadlineExceeded: no active session` was blamed on CPU
starvation and addressed by dropping runner capacity to 1; the likelier
mechanism is kubelet killing dind mid-build and taking the buildkit session
with it. Capacity reduced the load that trips the probe, which fits #17 passing
and #18-#21 failing anyway.
Written as a hypothesis with the command to confirm it, not as a conclusion.
The chart exposes no probe knobs, so the candidate fix is a Kyverno mutation in
Ansible, following the existing force-best-effort-cpu precedent.
The documented scale 0/1 recovery did not work this time: the volume returned
detached/faulted and refused to attach, so the pod sat in ContainerCreating.
auto-salvage: true cannot rescue it, because salvage happens during attach and
a faulted volume never gets that far.
The replica's failedAt timestamp is the only thing holding it faulted.
Clearing it brought the volume to attached/healthy and postgres to 1/1 in
under a minute, with repo data and CI history intact.
Adds the backup check first, which matters more now that Longhorn runs at one
replica and there is no second copy to fall back on.
The previous entry stated that restarting act_runner leaves every in-flight job
orphaned. That is wrong: run #16 had its remaining jobs re-dispatched to the
new pod and finished normally, while run #14 really was left stuck. Both
outcomes happen, so in_progress after a restart is ambiguous and the runner
logs are what settle it.
Also corrects the recovery advice. Gitea 1.26 has no cancel endpoint anywhere
in its swagger, and DELETE on a run returned 204 against a live run without
stopping it. The UI button is the only way to cancel.
Restarting act_runner leaves its in-flight jobs in_progress with nothing behind
them, and at capacity 1 one orphan blocks every later run. Gitea 1.26 has no
cancel endpoint; DELETE .../actions/runs/{index} does the job and takes the run
index, not the database id.
Every job was failing at Set up job after ~14 minutes, then the first real step died
at 0s. It read like a broken action; it was an empty dind image cache re-pulling the
1.6GB act job image on every run.
dind has no volume for /var/lib/docker, so the cache lives in its writable layer and
dies with every pod restart. A restart-looping runner therefore never keeps one.
Warming it by hand took lint from failure to success with no code change.
trivy on the worker image: 39 findings (2 CRITICAL) -> 18 -> 5 -> 0. Verified in CI
run #7 on 4a426db: 'Total: 0 (HIGH: 0, CRITICAL: 0)' for all three images.
The last five lived in kubectl's vendored golang.org/x/net and Go stdlib, inside the
newest kubectl published. No version cleared them; removing the binary did.
CORRECTNESS
- lost-lease race: complete()/fail() did not check ownership, so a worker whose
lease expired could mark a task done while another worker was running it, or
requeue a task someone else owned. Reproduced, fixed with a CAS on
(state, locked_by), pinned by two regression tests.
- worker died on report failure: _run_one's docstring claimed no exception
escapes the TaskGroup; fail()/complete() were outside the guarded block, so a
DB blip cancelled every sibling provision on the pod.
- claim query used an INNER join, which could strand a just-claimed task and
report 'queue empty'. LEFT join.
- InstanceRepo.set_error bypassed the state machine and had no callers. Deleted.
- handle_deprovision ignored its CAS result, so a wrong-state instance kept a
dangling endpoint and got re-provisioned by the drift check 60s later.
- handle_verify re-notified on every retry: five pages for one halt.
DEPLOY-BREAKING
- the migration Job could never succeed: no Dockerfile copied migrations/, and
migrate.py resolved the path relative to the source tree, which only works for
an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR.
- ServiceMonitor selector did not match the Service: API metrics never scraped.
- SvcforgeReconcilerStale fired permanently from every pod, because the gauge is
module-level and every service exports it as 0. Scoped to the reconciler job.
- SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m].
- the digest guard accepted the all-zeros placeholder.
- worker terminationGracePeriodSeconds was 60s against a 600s helm timeout.
DEAD CODE THAT SHOULD NOT HAVE BEEN
- adapters/k8s.py was never called, so tenant namespaces were never created and
the first provision for a new team would fail. Wired into handle_provision.
- adapters/redis.py was never imported by any service. Rate limiting is now wired
into the API, failing open.
- Settings.check_production() had no callers. Given an explicit environment and
called from every entrypoint.
OBSERVABILITY
- the API never called obs.setup(): no JSON logs, no trace correlation, log_json
silently inert.
- LogNotifier's structured fields were discarded by the stdlib->structlog bridge.
- bind_task_context cleared the 'service' binding for the life of every task.
- split tasks_failed into task_attempts_failed and tasks_dead_lettered.
SECURITY
- trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl
1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3,
which also closes a four-minor skew against the v1.35.3 cluster.
TESTS THAT COULD NOT FAIL
- the concurrency cap test passed on a fully serial worker.
- the alert/metric cross-check asserted a hardcoded list instead of reading the
chart, so it could not catch a rename on the chart side.
- fixed OTel tracer-provider pollution between test files.
DOCS
- ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD
guarantee (verified with --dry-run=server).
- AGENTS.md + CLAUDE.md.
- prose sweep for back-and-forth phrasing across 19 files.
trivy-action@v0.29.0 internally uses setup-trivy@v0.2.2, a tag removed
upstream (earliest published is now v0.2.6), so it cannot resolve on any
runner. Run trivy from a digest-pinned image instead, as gitleaks already is.
RUNBOOK gains a 'Setting up CI/CD from scratch' section with the traps that
actually cost time: GITEA_TOKEN is 401 at the package registry, the runner's
cache fails soft, service containers resolve by name not localhost.
Not pushed: pushing triggers a run, and the runner is being restarted by the
ansible change that enables its cache.