cbd070928116f6f7e3d508a8fa8f9aff1ba36d77
17 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd61eca12e |
runbook: correct 1b56cda -- the poll never runs, watch breaks are not the cause
ci / lint (push) Successful in 25s
ci / types (push) Successful in 54s
ci / unit (push) Successful in 42s
ci / security (push) Successful in 57s
ci / dockerfile (push) Successful in 7s
ci / chart (push) Successful in 10s
ci / integration (push) Successful in 59s
ci / image (api) (push) Successful in 1m38s
ci / image (reconciler) (push) Successful in 1m32s
ci / image (worker) (push) Successful in 1m34s
ci / bump (push) Successful in 13s
|
||
|
|
1b56cda231 |
runbook: name the ArgoCD wedge instead of describing its symptoms
ci / lint (push) Successful in 37s
ci / security (push) Successful in 1m10s
ci / types (push) Successful in 2m8s
ci / unit (push) Successful in 1m49s
ci / dockerfile (push) Successful in 7s
ci / chart (push) Failing after 50s
ci / integration (push) Successful in 1m12s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
Reproduced three times on 2026-07-22, and the controller's own workqueue
metrics say what it is:
workqueue_depth 0
workqueue_longest_running_processor_seconds 0
workqueue_adds_total 68 <- frozen
Empty queue, idle processors, no new adds. Nothing is blocked -- nothing is
being enqueued. The periodic app-resync timer stops firing, so no Application
is ever queued for refresh again.
Every occurrence followed an API server disruption, with the matching informer
line in the logs: watch ended with error, http2: client connection lost, on
*v1alpha1.Application. The informer re-lists; the resync does not resume.
That is why the controller reads healthy from every angle except reconcile
count, and why only a restart clears it. On a single-control-plane cluster
anything that kills kube-apiserver -- including memory reclaim stalling /livez
past its 80s threshold -- can silently stop all GitOps.
|
||
|
|
60ee0f1cbf |
runbook: correct the wedge-detection metrics
ci / lint (push) Successful in 2m37s
ci / types (push) Successful in 1m10s
ci / unit (push) Successful in 38s
ci / security (push) Successful in 55s
ci / dockerfile (push) Successful in 18s
ci / chart (push) Successful in 8s
ci / integration (push) Successful in 1m25s
ci / image (api) (push) Successful in 4m17s
ci / image (reconciler) (push) Successful in 2m8s
ci / image (worker) (push) Successful in 2m3s
ci / bump (push) Successful in 16s
|
||
|
|
c691f4f4aa |
runbook: reconciledAt is not a heartbeat
ci / lint (push) Successful in 33s
ci / types (push) Successful in 50s
ci / unit (push) Successful in 31s
ci / security (push) Successful in 1m17s
ci / dockerfile (push) Successful in 10s
ci / chart (push) Successful in 8s
ci / integration (push) Successful in 1m6s
ci / image (api) (push) Successful in 4m44s
ci / image (reconciler) (push) Successful in 4m5s
ci / image (worker) (push) Successful in 2m55s
ci / bump (push) Successful in 1m12s
The stuck-deploy section printed `.status.reconciledAt` as a diagnostic without saying how to read it, which invites exactly the wrong conclusion: ArgoCD only writes that field when the computed status changes, so on an idle cluster it stops advancing while the controller is healthy. Measured 2026-07-22 -- six samples 50s apart against a 120s reconciliation timeout, zero movement. It only means something alongside a sync.revision that is behind master. Replaces the "is it doing anything" step with two metrics that answer it directly: the controller's Redis cache reads, which happen every refresh cycle regardless of change, and the reconcile queue's unfinished-work seconds, which is precisely what the 2026-07-21 webhook wedge looked like. Log-based check kept as the no-Prometheus fallback. |
||
|
|
a0215ef63f |
runbook: ArgoCD synced-but-stale, and how to tell a wedged controller from a slow poll
ci / lint (push) Successful in 27s
ci / types (push) Successful in 49s
ci / unit (push) Successful in 1m51s
ci / chart (push) Successful in 10s
ci / integration (push) Successful in 1m15s
ci / dockerfile (push) Failing after 10m28s
ci / security (push) Failing after 11m1s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
Cost 11 hours yesterday and looked like nothing was wrong, because the Application reported Synced the whole time — at a revision five commits behind. Records the diagnosis that actually works (log-lines-per-hour, and a flat goroutine count meaning blocked rather than idle), the Kyverno failurePolicy root cause, and the two escalating fixes. Folds in the hook-finalizer deadlock, which has now happened four times and was only written down in chat. |
||
|
|
7079d6340f |
docs: bring RUNBOOK and ARCHITECTURE up to date
ci / lint (push) Successful in 23s
ci / types (push) Successful in 32s
ci / unit (push) Successful in 27s
ci / security (push) Successful in 37s
ci / dockerfile (push) Successful in 6s
ci / chart (push) Successful in 7s
ci / integration (push) Successful in 47s
ci / image (api) (push) Successful in 1m0s
ci / image (reconciler) (push) Successful in 2m14s
ci / image (worker) (push) Successful in 2m16s
ci / bump (push) Successful in 26s
The runner moved to node0 and the drift check reads release secrets via the Kubernetes API in-cluster; the docs still described node2 and `helm list`. - RUNBOOK: the durable image-cache fix is now the node0 hostPath store, not a node2 pin; /data is NFS RWX, so the Multi-Attach wait is gone. Cross-references entries 9 and 10. - ARCHITECTURE: the reconciler's edge to the cluster is "list releases", not "helm list" (helm is the out-of-cluster fallback). - ARCHITECTURE: the state diagram and its prose described fail() moving every dead-lettered instance to `failed`. Corrected to the per-kind behaviour — only provision fails the instance; deprovision stays `deleting` for retry, upgrade and verify stay `ready` — matching the fix in tasks.py. |
||
|
|
a843494627 |
runbook: the runner's three ephemeral caches and the boot cascade
ci / lint (push) Successful in 2m25s
ci / types (push) Successful in 38s
ci / unit (push) Successful in 28s
ci / security (push) Successful in 48s
ci / dockerfile (push) Successful in 18s
ci / chart (push) Successful in 7s
ci / integration (push) Successful in 1m11s
ci / image (api) (push) Successful in 4m57s
ci / image (reconciler) (push) Successful in 2m10s
ci / image (worker) (push) Successful in 2m21s
ci / bump (push) Successful in 15s
Three caches on this runner were container-layer only, each found because something was slow: dind's image store, the trivy vuln DB, and act's action clones. All three now have real storage. act clones actions with full history, not shallow — 66.7MB/538 commits for setup-uv, 24.4MB/222 for actions/checkout, and this workflow uses five. After a restart that made `Set up job` an 11-minute step with the job container sat idle running `sleep` while the runner cloned GitHub. Includes the command to tell those two apart. Also records the cascade the dind fix creates. A persistent image store makes dockerd scan on boot — 38s, 2m13s, or over 5 minutes depending on node load — and two timeouts then fire: dind's startup probe, and the runner image's own hardcoded `Docker wait timeout of 5m0s`, which the chart cannot configure. The runner exits 1, restarts, and kills the running job, which looks like every step failing at once with no error after a green `Set up job`. It self-heals in about ten minutes at the cost of one CI run, so runner restarts are now something to do deliberately rather than casually. |
||
|
|
c2a27952d1 |
runbook: fix section ordering and a duplicate number
The dind entry landed ahead of the postgres one, and the postgres entry I added earlier was numbered 7 while 'Verify the whole loop' already was. Now 7 postgres, 8 dind, 9 verify. |
||
|
|
0dbb5af1d3 |
runbook: record the dind liveness probe as an open issue
dind is killed by a probe whose only job is to check a socket exists, with a 1s timeout the nodes cannot always meet: 27 failures over 156 minutes. Worth recording because it probably explains build failures already attributed to something else. `DeadlineExceeded: no active session` was blamed on CPU starvation and addressed by dropping runner capacity to 1; the likelier mechanism is kubelet killing dind mid-build and taking the buildkit session with it. Capacity reduced the load that trips the probe, which fits #17 passing and #18-#21 failing anyway. Written as a hypothesis with the command to confirm it, not as a conclusion. The chart exposes no probe knobs, so the candidate fix is a Kyverno mutation in Ansible, following the existing force-best-effort-cpu precedent. |
||
|
|
d4ac3801a3 |
runbook: salvaging a faulted Longhorn volume
The documented scale 0/1 recovery did not work this time: the volume returned detached/faulted and refused to attach, so the pod sat in ContainerCreating. auto-salvage: true cannot rescue it, because salvage happens during attach and a faulted volume never gets that far. The replica's failedAt timestamp is the only thing holding it faulted. Clearing it brought the volume to attached/healthy and postgres to 1/1 in under a minute, with repo data and CI history intact. Adds the backup check first, which matters more now that Longhorn runs at one replica and there is no second copy to fall back on. |
||
|
|
5f18f9eeeb |
runbook: correct the claim that runner restarts orphan jobs
The previous entry stated that restarting act_runner leaves every in-flight job orphaned. That is wrong: run #16 had its remaining jobs re-dispatched to the new pod and finished normally, while run #14 really was left stuck. Both outcomes happen, so in_progress after a restart is ambiguous and the runner logs are what settle it. Also corrects the recovery advice. Gitea 1.26 has no cancel endpoint anywhere in its swagger, and DELETE on a run returned 204 against a live run without stopping it. The UI button is the only way to cancel. |
||
|
|
5d7f46483e |
runbook: recovering the queue after a runner restart
Restarting act_runner leaves its in-flight jobs in_progress with nothing behind
them, and at capacity 1 one orphan blocks every later run. Gitea 1.26 has no
cancel endpoint; DELETE .../actions/runs/{index} does the job and takes the run
index, not the database id.
|
||
|
|
f87d8d4d78 |
runbook: the 14-minute 'Set up job' failure and its cause
ci / lint (push) Waiting to run
ci / types (push) Blocked by required conditions
ci / unit (push) Blocked by required conditions
ci / integration (push) Blocked by required conditions
ci / security (push) Blocked by required conditions
ci / dockerfile (push) Blocked by required conditions
ci / chart (push) Blocked by required conditions
ci / image (api) (push) Blocked by required conditions
ci / image (reconciler) (push) Blocked by required conditions
ci / image (worker) (push) Blocked by required conditions
ci / bump (push) Blocked by required conditions
Every job was failing at Set up job after ~14 minutes, then the first real step died at 0s. It read like a broken action; it was an empty dind image cache re-pulling the 1.6GB act job image on every run. dind has no volume for /var/lib/docker, so the cache lives in its writable layer and dies with every pod restart. A restart-looping runner therefore never keeps one. Warming it by hand took lint from failure to success with no code change. |
||
|
|
78a0a6d500 |
runbook: record how the image gate was taken to zero findings
ci / lint (push) Successful in 21s
ci / unit (push) Failing after 41s
ci / integration (push) Has been skipped
ci / types (push) Successful in 1m18s
ci / dockerfile (push) Successful in 6s
ci / chart (push) Successful in 7s
ci / security (push) Successful in 54s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
trivy on the worker image: 39 findings (2 CRITICAL) -> 18 -> 5 -> 0. Verified in CI
run #7 on
|
||
|
|
c76154aeaa |
review: fix 26 findings from a 4-agent audit
ci / lint (push) Successful in 34s
ci / unit (push) Successful in 1m41s
ci / types (push) Successful in 1m41s
ci / dockerfile (push) Successful in 18s
ci / security (push) Successful in 1m27s
ci / chart (push) Failing after 1m11s
ci / integration (push) Successful in 1m10s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
CORRECTNESS - lost-lease race: complete()/fail() did not check ownership, so a worker whose lease expired could mark a task done while another worker was running it, or requeue a task someone else owned. Reproduced, fixed with a CAS on (state, locked_by), pinned by two regression tests. - worker died on report failure: _run_one's docstring claimed no exception escapes the TaskGroup; fail()/complete() were outside the guarded block, so a DB blip cancelled every sibling provision on the pod. - claim query used an INNER join, which could strand a just-claimed task and report 'queue empty'. LEFT join. - InstanceRepo.set_error bypassed the state machine and had no callers. Deleted. - handle_deprovision ignored its CAS result, so a wrong-state instance kept a dangling endpoint and got re-provisioned by the drift check 60s later. - handle_verify re-notified on every retry: five pages for one halt. DEPLOY-BREAKING - the migration Job could never succeed: no Dockerfile copied migrations/, and migrate.py resolved the path relative to the source tree, which only works for an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR. - ServiceMonitor selector did not match the Service: API metrics never scraped. - SvcforgeReconcilerStale fired permanently from every pod, because the gauge is module-level and every service exports it as 0. Scoped to the reconciler job. - SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m]. - the digest guard accepted the all-zeros placeholder. - worker terminationGracePeriodSeconds was 60s against a 600s helm timeout. DEAD CODE THAT SHOULD NOT HAVE BEEN - adapters/k8s.py was never called, so tenant namespaces were never created and the first provision for a new team would fail. Wired into handle_provision. - adapters/redis.py was never imported by any service. Rate limiting is now wired into the API, failing open. - Settings.check_production() had no callers. Given an explicit environment and called from every entrypoint. OBSERVABILITY - the API never called obs.setup(): no JSON logs, no trace correlation, log_json silently inert. - LogNotifier's structured fields were discarded by the stdlib->structlog bridge. - bind_task_context cleared the 'service' binding for the life of every task. - split tasks_failed into task_attempts_failed and tasks_dead_lettered. SECURITY - trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl 1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3, which also closes a four-minor skew against the v1.35.3 cluster. TESTS THAT COULD NOT FAIL - the concurrency cap test passed on a fully serial worker. - the alert/metric cross-check asserted a hardcoded list instead of reading the chart, so it could not catch a rename on the chart side. - fixed OTel tracer-provider pollution between test files. DOCS - ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD guarantee (verified with --dry-run=server). - AGENTS.md + CLAUDE.md. - prose sweep for back-and-forth phrasing across 19 files. |
||
|
|
c9d0176bb3 |
ci: run trivy directly; document CI/CD setup in RUNBOOK
trivy-action@v0.29.0 internally uses setup-trivy@v0.2.2, a tag removed upstream (earliest published is now v0.2.6), so it cannot resolve on any runner. Run trivy from a digest-pinned image instead, as gitleaks already is. RUNBOOK gains a 'Setting up CI/CD from scratch' section with the traps that actually cost time: GITEA_TOKEN is 401 at the package registry, the runner's cache fails soft, service containers resolve by name not localhost. Not pushed: pushing triggers a run, and the runner is being restarted by the ansible change that enables its cache. |
||
|
|
50c2fe2a1e |
svcforge: reference implementation
ci / lint (push) Successful in 1m19s
ci / unit (push) Failing after 1m2s
ci / integration (push) Has been skipped
ci / types (push) Successful in 1m37s
ci / security (push) Failing after 38s
ci / dockerfile (push) Successful in 14s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
Complete working build of the system learn-python/ teaches. 164 tests, mypy --strict clean, domain coverage 99%. |