The Job ran as the api ServiceAccount, which is an ordinary chart resource.
Hooks are created before the release's ordinary manifests, so on a first
install that account does not exist and the Job never starts:
Error creating: pods "svcforge-migrate-" is forbidden: error looking up
service account svcforge/svcforge-api: serviceaccount "svcforge-api" not
found
Not an ArgoCD quirk — helm orders hooks the same way, so both deploy paths
failed identically. It survived review because the chart was only ever checked
with `helm template` and `helm install --dry-run=server`, and neither creates a
Job. The pod is what fails, so only a real install can catch it.
The new account is a hook at weight -10, ahead of the Job at -5, and is bound
to no Role: the migration talks to Postgres and wants nothing from the
Kubernetes API. automountServiceAccountToken is off for the same reason.
Verified with a real `helm install` into a scratch namespace: STATUS deployed,
job Complete 1/1, four migrations applied, and the hook ServiceAccount was 27s
old against 10s for the ordinary ones — the ordering the bug turned on.
The block describes a ClusterSecretStore named `vault` with HashiCorp-style
key/property refs. This cluster has `oci-vault` — OCI Vault via
InstancePrincipal — whose provider addresses a secret by name and takes a JSON
property, so those remoteRefs do not translate as written. There are no
ExternalSecrets anywhere on the cluster, so the path had never been exercised.
svcforge.secretName still resolves through targetName, so the deployments and
the migrate hook read a Secret named svcforge-secrets, created out of band from
~/.config/svcforge/secrets.env.
This is a deviation and the comment says so. ArgoCD does not manage that
Secret, so prune and selfHeal cannot touch it, and it is the one part of the
deployment that cannot be read from git. Restoring the intended design means
adding oci_vault_secret resources to oci-k8s/infra/vault.tf and repointing
secretStoreRef at oci-vault.
CORRECTNESS
- lost-lease race: complete()/fail() did not check ownership, so a worker whose
lease expired could mark a task done while another worker was running it, or
requeue a task someone else owned. Reproduced, fixed with a CAS on
(state, locked_by), pinned by two regression tests.
- worker died on report failure: _run_one's docstring claimed no exception
escapes the TaskGroup; fail()/complete() were outside the guarded block, so a
DB blip cancelled every sibling provision on the pod.
- claim query used an INNER join, which could strand a just-claimed task and
report 'queue empty'. LEFT join.
- InstanceRepo.set_error bypassed the state machine and had no callers. Deleted.
- handle_deprovision ignored its CAS result, so a wrong-state instance kept a
dangling endpoint and got re-provisioned by the drift check 60s later.
- handle_verify re-notified on every retry: five pages for one halt.
DEPLOY-BREAKING
- the migration Job could never succeed: no Dockerfile copied migrations/, and
migrate.py resolved the path relative to the source tree, which only works for
an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR.
- ServiceMonitor selector did not match the Service: API metrics never scraped.
- SvcforgeReconcilerStale fired permanently from every pod, because the gauge is
module-level and every service exports it as 0. Scoped to the reconciler job.
- SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m].
- the digest guard accepted the all-zeros placeholder.
- worker terminationGracePeriodSeconds was 60s against a 600s helm timeout.
DEAD CODE THAT SHOULD NOT HAVE BEEN
- adapters/k8s.py was never called, so tenant namespaces were never created and
the first provision for a new team would fail. Wired into handle_provision.
- adapters/redis.py was never imported by any service. Rate limiting is now wired
into the API, failing open.
- Settings.check_production() had no callers. Given an explicit environment and
called from every entrypoint.
OBSERVABILITY
- the API never called obs.setup(): no JSON logs, no trace correlation, log_json
silently inert.
- LogNotifier's structured fields were discarded by the stdlib->structlog bridge.
- bind_task_context cleared the 'service' binding for the life of every task.
- split tasks_failed into task_attempts_failed and tasks_dead_lettered.
SECURITY
- trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl
1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3,
which also closes a four-minor skew against the v1.35.3 cluster.
TESTS THAT COULD NOT FAIL
- the concurrency cap test passed on a fully serial worker.
- the alert/metric cross-check asserted a hardcoded list instead of reading the
chart, so it could not catch a rename on the chart side.
- fixed OTel tracer-provider pollution between test files.
DOCS
- ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD
guarantee (verified with --dry-run=server).
- AGENTS.md + CLAUDE.md.
- prose sweep for back-and-forth phrasing across 19 files.