review: fix 26 findings from a 4-agent audit
ci / lint (push) Successful in 34s
ci / unit (push) Successful in 1m41s
ci / types (push) Successful in 1m41s
ci / dockerfile (push) Successful in 18s
ci / security (push) Successful in 1m27s
ci / chart (push) Failing after 1m11s
ci / integration (push) Successful in 1m10s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped

CORRECTNESS
- lost-lease race: complete()/fail() did not check ownership, so a worker whose
  lease expired could mark a task done while another worker was running it, or
  requeue a task someone else owned. Reproduced, fixed with a CAS on
  (state, locked_by), pinned by two regression tests.
- worker died on report failure: _run_one's docstring claimed no exception
  escapes the TaskGroup; fail()/complete() were outside the guarded block, so a
  DB blip cancelled every sibling provision on the pod.
- claim query used an INNER join, which could strand a just-claimed task and
  report 'queue empty'. LEFT join.
- InstanceRepo.set_error bypassed the state machine and had no callers. Deleted.
- handle_deprovision ignored its CAS result, so a wrong-state instance kept a
  dangling endpoint and got re-provisioned by the drift check 60s later.
- handle_verify re-notified on every retry: five pages for one halt.

DEPLOY-BREAKING
- the migration Job could never succeed: no Dockerfile copied migrations/, and
  migrate.py resolved the path relative to the source tree, which only works for
  an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR.
- ServiceMonitor selector did not match the Service: API metrics never scraped.
- SvcforgeReconcilerStale fired permanently from every pod, because the gauge is
  module-level and every service exports it as 0. Scoped to the reconciler job.
- SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m].
- the digest guard accepted the all-zeros placeholder.
- worker terminationGracePeriodSeconds was 60s against a 600s helm timeout.

DEAD CODE THAT SHOULD NOT HAVE BEEN
- adapters/k8s.py was never called, so tenant namespaces were never created and
  the first provision for a new team would fail. Wired into handle_provision.
- adapters/redis.py was never imported by any service. Rate limiting is now wired
  into the API, failing open.
- Settings.check_production() had no callers. Given an explicit environment and
  called from every entrypoint.

OBSERVABILITY
- the API never called obs.setup(): no JSON logs, no trace correlation, log_json
  silently inert.
- LogNotifier's structured fields were discarded by the stdlib->structlog bridge.
- bind_task_context cleared the 'service' binding for the life of every task.
- split tasks_failed into task_attempts_failed and tasks_dead_lettered.

SECURITY
- trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl
  1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3,
  which also closes a four-minor skew against the v1.35.3 cluster.

TESTS THAT COULD NOT FAIL
- the concurrency cap test passed on a fully serial worker.
- the alert/metric cross-check asserted a hardcoded list instead of reading the
  chart, so it could not catch a rename on the chart side.
- fixed OTel tracer-provider pollution between test files.

DOCS
- ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD
  guarantee (verified with --dry-run=server).
- AGENTS.md + CLAUDE.md.
- prose sweep for back-and-forth phrasing across 19 files.
This commit is contained in:
Nguyen Minh Phuc
2026-07-18 12:13:49 +00:00
parent 77d560ddae
commit c76154aeaa
45 changed files with 1520 additions and 216 deletions
+55 -4
View File
@@ -190,12 +190,63 @@ jobs:
dockerfile: "services/*/Dockerfile"
failure-threshold: warning
# --- stage 9b: the chart must render --------------------------------------------------
# Without this, a chart that does not template reaches ArgoCD and fails in the cluster,
# where the error surfaces as a sync failure with no PR attached to it. `helm template`
# is the real gate: it is what ArgoCD does, and _helpers.tpl's image helper calls `fail`
# on anything that is not a full sha256 digest.
chart:
runs-on: ubuntu-latest
needs: [lint]
env:
# Pinned by digest like gitleaks and trivy, and the same helm the worker and
# reconciler images carry — CI renders with the version that ships.
HELM: alpine/helm:3.21.3@sha256:35da09ba0716fc7c3cd63b6b31ee380a9c7662e95f29ab0e4ae962420afd315b
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
- name: helm lint
run: |
docker run --rm -v "$PWD:/repo" -w /repo "$HELM" lint deploy/chart
- name: helm template (rejects unbumped digests)
# values.yaml ships all-zeros placeholders, so a bare `helm template` MUST fail.
# That is the guard working, not a broken chart — asserting the failure here is
# what stops the guard silently regressing into a prefix check again.
run: |
set -euo pipefail
if docker run --rm -v "$PWD:/repo" -w /repo "$HELM" \
template svcforge deploy/chart >/dev/null 2>&1; then
echo "FAIL: chart rendered against the placeholder digests in values.yaml." >&2
echo "The digest guard in _helpers.tpl is not guarding." >&2
exit 1
fi
echo "ok: placeholder digests rejected"
- name: helm template (renders with real digests)
# Dummy but well-formed digests: this checks the templates themselves render, with
# the two values-gated monitoring blocks explicitly on so they are covered too.
run: |
set -euo pipefail
A="sha256:$(printf 'a%.0s' $(seq 64))"
B="sha256:$(printf 'b%.0s' $(seq 64))"
C="sha256:$(printf 'c%.0s' $(seq 64))"
docker run --rm -v "$PWD:/repo" -w /repo "$HELM" \
template svcforge deploy/chart \
--set image.api.digest="$A" \
--set image.worker.digest="$B" \
--set image.reconciler.digest="$C" \
--set serviceMonitor.enabled=true \
--set prometheusRule.enabled=true \
>/dev/null
echo "ok: chart renders"
# --- stage 10: build -> trivy -> push by digest --------------------------------------
image:
runs-on: ubuntu-latest
# Every gate above is required. An image is not built until all of them are green,
# which is what makes "the digest CI pushed is a digest that passed everything" true.
needs: [types, unit, integration, security, dockerfile]
needs: [types, unit, integration, security, dockerfile, chart]
permissions:
contents: read
strategy:
@@ -314,9 +365,9 @@ jobs:
steps:
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
# A bot token with contents:write on this repo and nothing else. Note what is
# absent: no kubeconfig, no cluster credential, no ArgoCD API token. CI's maximum
# blast radius is a bad commit, which is revertable.
# A bot token with contents:write on this repo and nothing else: no kubeconfig,
# no cluster credential, no ArgoCD API token. CI's maximum blast radius is a bad
# commit, which is revertable.
token: ${{ secrets.CI_BOT_TOKEN }}
ref: master
- uses: docker/login-action@9780b0c442fbb1117ed29e0efdff1e18412f7567 # v3.3.0
+150
View File
@@ -0,0 +1,150 @@
# AGENTS.md — working in svcforge-reference
Rules for any AI agent (Claude Code, Cursor, Codex, Copilot, Gemini CLI) editing this repo.
Read this before your first edit.
---
## What this repo is, and the one rule that follows from it
This is a **reference implementation**: a complete, verified build of the system that
`../learn-python/` teaches Phuc to build himself. It exists to be *compared against*, not
copied from.
**THE RULE: never propose copying `domain/`, `repo/`, `services/worker/`, or the claim
query into `learn-python/`.** Those four are the course. If asked to "help with Module 4",
help him find his own bug; do not paste `repo/tasks.py`. Scaffolding — `pyproject.toml`,
Dockerfiles, CI, the chart — is fair game to copy, because struggling with hatchling
teaches nothing.
`learn-python/AGENTS.md` has its own, stricter rules. They still apply over there.
---
## Before you change anything
Read [ARCHITECTURE.md](ARCHITECTURE.md). Every design choice here was paid for by a
specific failure mode, and the comments say which. A change that removes a comment's
subject usually removes a guard.
Four things that look like over-engineering and are not. Do not "simplify" them:
1. **The claim query is one statement.** `repo/tasks.py::_CLAIM_SQL`. Splitting the
`UPDATE ... WHERE id = (SELECT ... FOR UPDATE SKIP LOCKED)` into select-then-update
creates a double-provision race that testing will not catch.
2. **`complete()` and `fail()` check `locked_by`.** That check is a repro-confirmed bug
fix, with regression tests in `tests/integration/test_tasks.py`. Removing it lets a
worker whose lease expired clobber a task another worker now owns.
3. **`domain/` has no I/O and no `async`.** CI greps for both.
4. **Migrations never run at app startup.** They are a Helm `pre-upgrade,pre-install` hook.
---
## Conventions
**Layers.** `transport → domain ← repo/adapters`. `domain/` imports nothing from the other
three. `repo/` knows SQL and not HTTP. `adapters/` owns every subprocess and network call.
**Python 3.12.** Full annotations everywhere including tests. `X | None`, `StrEnum`,
`Final`, PEP 695 `type` aliases.
**Errors.** `SvcforgeError` in `svcforge_core/errors.py` is the base. New exception types
inherit from it, plus a stdlib base when the behaviour matters (`HelmError(SvcforgeError,
RuntimeError)`).
**Logging.** `obs.get_logger(__name__)`, structlog kwargs. Never stdlib `logging` directly,
and never `extra={...}` — the ProcessorFormatter bridge silently discards it. Never use
`message` or `event` as a field name; both are reserved and raise.
**SQL.** Raw psycopg3, no ORM. Parameters, never f-strings. Rows come back as dicts
(`DictPool` in `repo/db.py`); annotate with that alias rather than a bare
`AsyncConnectionPool`, which mypy resolves to tuple rows.
**Subprocess.** `create_subprocess_exec` with `start_new_session=True`. Never `shell=True`
`team` is tenant input that reaches a helm release name. Timeouts kill the process *group*.
**Protocols.** A `Protocol` needs two implementations (a real one and a fake) or it should
be a concrete class.
**Datetimes.** Always aware, always UTC. Domain functions take `now` as a parameter;
only adapters hold a `Clock`.
**Pinning.** Everything by digest: base images, tool images, GitHub Actions by SHA.
A SHA-pinned third-party action still resolves *its own* dependencies by mutable tag —
that is how `trivy-action` broke. Prefer running tools directly from a digest-pinned image.
---
## Writing style
Plain declaratives. **No back-and-forth phrasing**, which means none of:
- Rhetorical setup-knockdowns: "That is not a nicety. It is the whole design."
- not-X-but-Y pivots used for rhythm: "It is not a lock, but a lease."
- Throat-clearing: "It is worth noting that", "The key insight here", "Here's the thing".
- Rhetorical questions you then answer.
- Dramatic fragments for rhythm: "Not slow. Not fast. Meaningless."
Write the claim directly. Keep the *why* — it is the most valuable content in this repo —
and drop the scaffolding around it. Factual contrasts that carry information stay
("LEFT join, not inner"; "202, never 201").
---
## Testing
- **Never mock the database.** Integration tests run against real Postgres. Each pytest
process clones its own database (`tests/integration/conftest.py`).
- **Fakes, not mocks**, for helm/kubectl/Redis/clock. `tests/fakes.py`.
- **A test that cannot fail is worse than no test.** Before adding one, ask what change
would make it red. Bound assertions on both sides where it matters: `max_concurrent
<= cap` alone passes on a fully serial worker.
- Markers: `slow` (>1s or hits real Upstash), `e2e` (needs a cluster). `make test` runs
neither.
```bash
export SVCFORGE_TEST_DSN="postgresql://svcforge:svcforge@127.0.0.1:5432/svcforge_test"
make lint # ruff + ruff format + mypy --strict
make test # unit
make integration # needs the DSN above
make ci # every gate CI runs, same commands
```
---
## Environment facts that will bite you
- **This host has no Docker daemon, and must not get one.** Its `containerd` belongs to a
Kubernetes kubelet; installing `docker.io` would evict the runtime and take down every
pod. You cannot `docker build` here. Verify Dockerfiles by inspection and `hadolint`.
- **Postgres is local** (apt, v16) for tests. testcontainers is the CI path.
- **Redis is real Upstash on a free tier: 500K commands/month.** Redis-hitting tests are
marked `slow`. Do not write loops that spend thousands of commands; use the fakes.
- **Secrets live in `~/.config/svcforge/secrets.env`**, chmod 600, outside every repo.
Never echo, log, commit, or paste a DSN or token. `.env` is gitignored.
- **Infra changes go through Terraform/Ansible**, never `kubectl edit`. The Gitea runner's
config lives in `oci-k8s/k8s/roles/addons/tasks/main.yml`.
---
## CI/CD
`.gitea/workflows/ci.yaml`. Gates in order: ruff → mypy → unit+coverage → migrate+integration
→ bandit/gitleaks/pip-audit → hadolint → helm lint/template → build → trivy → push by digest
→ bump chart digests → ArgoCD syncs.
- **CI holds no kubeconfig and must never hold one.** Its last act is a git commit.
- **Build once, promote by digest.** The image is scanned before it is pushed.
- **Every gate blocks.** If a gate is failing, fix the cause. Do not add an ignore, lower a
threshold, or mark it advisory. A CVE finding from trivy is real: bump the pinned version.
`RUNBOOK.md` has the setup steps and the traps that cost real time.
---
## When you are unsure
Say so, and check. This repo's value is that its claims were executed, not asserted — the
race conditions have reproductions, the timeout kill has a negative control, the coverage
gate was caught measuring 0%. Guessing and sounding confident is the one failure mode that
makes it worthless.
+473
View File
@@ -0,0 +1,473 @@
# svcforge — how it works
A control plane for internal service provisioning. A team asks for an Elasticsearch; a
worker installs one; a control loop keeps reality matching the database.
This document explains the design and **why each piece is the way it is**. Every "why" here
was paid for by a specific failure mode.
---
## The one-paragraph version
`POST /v1/instances` writes two rows in one transaction: the instance, and a task to build
it. It returns `202` immediately. A worker claims that task with a single `SELECT ... FOR
UPDATE SKIP LOCKED` statement, runs `helm upgrade --install`, and marks the instance ready.
A reconciler sweeps every 60 seconds for the things that go wrong when a process dies at
the wrong moment. The queue is a Postgres table, not Redis, and that single decision
determines most of the rest of the design.
---
## The shape
```mermaid
flowchart LR
subgraph tenant["Tenant"]
CLI["svcforge CLI<br/><i>HTTP only, never the DB</i>"]
end
subgraph plane["svcforge control plane"]
API["<b>api</b><br/>FastAPI, N replicas<br/>validates, authenticates,<br/>enqueues"]
WORKER["<b>worker</b><br/>N replicas<br/>claims tasks, runs helm"]
RECON["<b>reconciler</b><br/><i>exactly 1</i><br/>4 checks every 60s"]
end
subgraph state["State"]
PG[("<b>Postgres</b><br/>instances + tasks<br/><i>the truth</i>")]
REDIS[("<b>Redis</b><br/>rate limit, cache<br/><i>derived only</i>")]
end
K8S["<b>Kubernetes</b><br/>helm releases<br/><i>the real world</i>"]
CLI -->|"POST /v1/instances"| API
API -->|"one transaction:<br/>instance + task"| PG
API -.->|"best effort"| REDIS
WORKER -->|"claim<br/>SKIP LOCKED"| PG
WORKER -->|"helm upgrade --install"| K8S
RECON -->|"drift, leases,<br/>TTL, versions"| PG
RECON -->|"helm list"| K8S
classDef truth fill:#2d4a22,stroke:#5a8f3d,color:#fff
classDef derived fill:#4a3222,stroke:#8f6a3d,color:#fff
class PG truth
class REDIS derived
```
Three services, four layers, two stores. The layer rule is one line:
```
transport/ HTTP handlers, CLI entrypoints. Knows FastAPI. Knows nothing about SQL.
domain/ Pure logic: state machine, catalog, backoff, windows. No I/O. No async.
repo/ SQL. Rows in, domain objects out. Knows psycopg. Knows nothing about HTTP.
adapters/ The outside world: helm, kubectl, webhooks, Redis, the clock.
```
Dependencies point inward: `transport → domain ← repo/adapters`. `domain/` imports nothing
from the other three, which is why its tests need no mocks and run in milliseconds.
---
## The decision everything else follows from
**The queue is a Postgres table.**
A task and the instance state it describes must commit atomically. Split them across two
stores and you own a distributed commit problem with no winning move: the process can die
between the two writes, and whichever you write first is the one that lies.
- Task first, then instance → an orphan task pointing at an instance that never existed.
- Instance first, then task → an instance nobody will ever build.
In one table, in one transaction, neither is possible:
```sql
BEGIN;
INSERT INTO instances (...); -- state='requested'
INSERT INTO tasks (...); -- kind='provision'
COMMIT; -- both, or neither
```
Redis cannot do this, so Redis is not the queue. It holds derived state only — things that
can be recomputed and whose loss is an inconvenience, never a corruption.
| | Postgres | Redis |
|---|---|---|
| Holds | instances, tasks | rate limits, idempotency keys, cache |
| If it is down | the platform is down | the platform is fine |
| If it disagrees with reality | reality wins, reconciler fixes it | discard it |
| Lives on the | control loop **and** request path | request path only |
That last row is a budget constraint. Upstash's free tier is 500K
commands/month = **0.19 commands/second sustained**. One worker polling Redis every 5
seconds is 518,400/month — the entire budget, spent by one pod doing nothing.
---
## Provisioning, end to end
```mermaid
sequenceDiagram
autonumber
participant T as Tenant
participant A as api
participant P as Postgres
participant W as worker
participant K as Kubernetes
T->>A: POST /v1/instances {elasticsearch, small}
A->>A: verify JWT, look up catalog
rect rgb(45, 74, 34)
A->>P: BEGIN
A->>P: INSERT instance (state=requested)
A->>P: INSERT task (kind=provision, traceparent)
A->>P: COMMIT
end
A-->>T: 202 Accepted + Location
Note over W,P: every 5s, while a semaphore slot is free
W->>P: UPDATE ... FOR UPDATE SKIP LOCKED
P-->>W: task (attempts now 1, locked_by=me)
W->>K: kubectl apply namespace (idempotent)
W->>K: helm upgrade --install --wait
K-->>W: release ready
W->>P: CAS provisioning -> ready, set endpoint
W->>P: complete(task, worker_id)
T->>A: GET /v1/instances/{id}
A-->>T: {state: ready, endpoint: ...}
```
**Why 202 and not 201.** A provision is `helm --wait` on a StatefulSet: minutes. Holding
an HTTP connection open for that is a request that dies to any proxy timeout, and a client
that cannot tell "still working" from "lost". The queue absorbs the API's output, which is
the property that lets 50 simultaneous POSTs all return instantly.
---
## The claim query
This is the heart of the system. It is one statement, and it must stay one statement.
```sql
WITH claimed AS (
UPDATE tasks SET state='running', attempts=attempts+1,
locked_by=%(worker)s, locked_at=now()
WHERE id = (
SELECT id FROM tasks
WHERE state='queued' AND run_after <= now()
ORDER BY run_after
FOR UPDATE SKIP LOCKED -- step over rows other workers hold, do not queue behind them
LIMIT 1
)
RETURNING *
)
SELECT claimed.*, instances.team
FROM claimed LEFT JOIN instances ON instances.id = claimed.instance_id;
```
- **`FOR UPDATE SKIP LOCKED`** is what makes N workers scale. Without `SKIP LOCKED` they
queue single-file behind whoever holds the oldest row.
- **The subquery exists because Postgres has no `UPDATE ... LIMIT`.**
- **Select-then-update as two statements is the bug this prevents.** Between the `SELECT`
and the `UPDATE`, a second worker reads the same id and both provision. The window is
small, which means you will not hit it in testing and you will hit it in production.
- **`LEFT` join, not inner.** The `UPDATE` has already taken effect when the outer select
runs. An inner join matching nothing would return no row, so `claim()` would report
"queue empty" for a task it had just marked `running` — stranding it until the lease
expires, having silently burned an attempt.
- **`attempts` increments at claim time, not on failure.** A worker that dies without
reporting has still burned an attempt, so a task that reliably kills its worker cannot
retry forever.
---
## The instance lifecycle
```mermaid
stateDiagram-v2
[*] --> requested: POST /v1/instances
requested --> provisioning: worker claims
provisioning --> ready: helm --wait succeeded
ready --> deleting: DELETE, or TTL expired
deleting --> deleted: helm uninstall succeeded
requested --> failed: attempts exhausted
provisioning --> failed: attempts exhausted
ready --> failed: drift — the release vanished
failed --> provisioning: retry
failed --> deleting: give up, tear it down
deleting --> failed: attempts exhausted
deleted --> [*]: terminal
```
`LEGAL` is a `dict[InstanceState, frozenset[InstanceState]]` in `domain/states.py`, not a
chain of `if`s. `deleted` maps to an **empty frozenset** rather than being absent, so
"terminal" is stated rather than implied by a missing key.
**The state machine is enforced in SQL too.** `TaskRepo.fail` writes `instances.state`
directly, so it derives its guard from the same `LEGAL` table:
```python
_CAN_FAIL = tuple(s.value for s, allowed in LEGAL.items() if InstanceState.FAILED in allowed)
...
UPDATE instances SET state='failed' WHERE id=%s AND state = ANY(%s)
```
Without that, a deprovision exhausting its retries against an already-`deleted` instance
would resurrect it into `failed` — a transition `transition()` explicitly forbids,
performed by raw SQL that never asked it. A state machine only one layer respects is
decoration.
---
## What happens when things die
This is the part that matters. Every guarantee below has a test.
```mermaid
flowchart TD
START["worker claims task<br/>state=running, locked_by=me"] --> WORK["helm upgrade --install"]
WORK -->|success| REPORT["complete(task, worker_id)"]
WORK -->|"raises"| FAIL["fail(task, err, worker_id)"]
WORK -->|"pod SIGKILLed"| DEAD["nothing reported<br/>row stuck at 'running'"]
REPORT --> CAS{"still locked_by me?"}
CAS -->|yes| DONE["state=done"]
CAS -->|"no — lease was stolen"| DROP["log and drop.<br/>the new owner reports"]
FAIL --> ATT{"attempts < max?"}
ATT -->|yes| REQUEUE["state=queued<br/>run_after += backoff+jitter"]
ATT -->|no| DEADLETTER["state=failed<br/>error copied to instance"]
DEAD --> LEASE["reconciler: locked_at older<br/>than lease_seconds"]
LEASE --> REQUEUE
REQUEUE --> START
classDef bad fill:#4a2222,stroke:#8f3d3d,color:#fff
classDef good fill:#2d4a22,stroke:#5a8f3d,color:#fff
class DEAD,DEADLETTER bad
class DONE,DROP good
```
**Recovery is by lease.** No distributed lock survives a power cut. A SIGKILLed worker leaves
`state='running'` with `locked_by` set and nobody running it; that row would sit there
forever. The lease is the only thing that recovers it, which is why `locked_at` exists.
**Ownership is checked on report.** A worker that hangs past its lease has its task
requeued and re-claimed by someone else. When it finally returns, `complete()` and `fail()`
both check `AND state='running' AND locked_by=%s`. Without that check, the stale worker
marks the task done while the new owner is still running it — and if the new owner then
fails, a *third* worker provisions the same instance. That is the double-provision the
claim query exists to prevent, arriving through the back door.
**Idempotency is what makes all of this safe.** It is bought in two places:
1. A deterministic release name: `f"{team}-{service_type}-{id[:8]}"`, `UNIQUE` in the schema.
2. Adapters that state desired state — `helm upgrade --install`, `kubectl apply` — instead
of issuing imperative commands.
Running a handler twice equals running it once, so redelivery is boring.
---
## The reconciler
One replica. Four checks. Every 60 seconds.
```mermaid
flowchart LR
TICK(("tick<br/>every 60s")) --> D["<b>drift</b><br/>helm list vs DB"]
TICK --> L["<b>lease expiry</b><br/>running + locked_at old"]
TICK --> T["<b>TTL</b><br/>ready + expires_at passed"]
TICK --> V["<b>version drift</b><br/>chart_version ≠ catalog"]
D --> D1["release missing → re-enqueue provision"]
D --> D2["release unknown → <b>log only</b>"]
L --> L1["→ queued, locked_by=null"]
T --> T1["→ deleting + deprovision task"]
V --> V1["→ upgrade task, inside<br/>the maintenance window"]
```
**The drift check never auto-deletes.** A bug in a delete path is unrecoverable; a bug in a
report path is a Tuesday. Unknown releases are logged for a human.
**It is a singleton because it sweeps.** Two reconcilers double-enqueue and race on TTL.
`replicas: 1` and `strategy: Recreate` in the chart, plus per-check idempotency guards that
re-verify under `FOR UPDATE`.
---
## Day 2: upgrading the fleet
The team that runs this also patches it. That is the whole module, and it is **two columns
and one query**.
```mermaid
flowchart TD
EDIT["edit catalog.yaml<br/>21.3.15 → 21.3.16"] --> Q
Q["<b>the work list</b><br/>state=ready AND chart_version ≠ catalog<br/>AND NOT halted<br/>ORDER BY own_team DESC<br/>LIMIT max_in_flight"]
Q --> U["upgrade task<br/>run_after = next maintenance window"]
U --> H["helm upgrade --install"]
H --> W["write instances.chart_version<br/><i>only after helm succeeds</i>"]
W --> VER["verify task"]
VER -->|healthy| NEXT["next instance"]
VER -->|"failed"| HALT["catalog_versions.rollout_state='halted'<br/><i>work list goes empty</i>"]
NEXT --> Q
HALT --> HUMAN["cleared by hand, with SQL"]
classDef stop fill:#4a2222,stroke:#8f3d3d,color:#fff
class HALT,HUMAN stop
```
- **`instances.chart_version` is written only after helm succeeds.** Write it optimistically
and the fleet looks upgraded while it is not.
- **`ORDER BY team = own_team DESC`** puts your own instances first, so you are the tenant
who discovers the chart is broken. Eating your own dog food is enforced by an `ORDER BY`
rather than left to policy.
- **`max_in_flight` starts at 1** — a config value, not a scheduler. One at a time is what
makes the halt meaningful: the fleet stops after the first casualty, not after all of them.
- **The halt is one column, cleared by hand.** An automatic un-halt would just resume
breaking things.
Deliberately **not** built: `resize`, `backup`/`restore`, `helm rollback` automation,
deprecation timers, a rollouts table with history, a pause/resume CLI. Backup is a whole
subsystem and an untested restore is a rumour.
---
## Observability: one trace across the queue
```mermaid
sequenceDiagram
participant A as api
participant P as tasks table
participant W as worker
A->>A: span "POST /v1/instances" (trace abc123)
A->>P: INSERT ... traceparent='00-abc123-...'
Note over P: minutes pass. different pod.
W->>P: claim → row carries traceparent
W->>W: span "task.provision", parent=abc123
Note over A,W: one trace: POST → queue → helm
```
Trace context does **not** survive a queue on its own. The worker picks the row up in
another process with no ambient context. So the traceparent rides in the table. Skip this
and Tempo shows two unrelated traces for one provision, which is worse than no tracing
because it looks like it works.
Logs carry `instance_id` / `task_id` / `team` on every line via contextvars, bound once at
claim. Metrics are deliberately few, and two of them are named to prevent a specific
mistake: `svcforge_task_attempts_failed_total` counts *attempts* that raised, while
`svcforge_tasks_dead_lettered_total` counts tasks that gave up. Alerting on the first
pages you for ordinary retries that later succeed.
---
## Delivery: CI never touches the cluster
```mermaid
flowchart LR
PUSH["push to master"] --> GATES
subgraph GATES["gates — all required, none advisory"]
direction TB
G1["ruff"] --> G2["mypy --strict"] --> G3["pytest unit + coverage"]
G3 --> G4["migrate + pytest integration"]
G4 --> G5["bandit / gitleaks / pip-audit"]
G5 --> G6["hadolint + helm lint/template"]
end
GATES --> BUILD["build image (--load)"]
BUILD --> SCAN["trivy HIGH,CRITICAL"]
SCAN -->|clean| PUSHIMG["push by digest"]
SCAN -->|"CVE"| STOP["pipeline fails"]
PUSHIMG --> BUMP["commit digest to values.yaml<br/>[skip ci]"]
BUMP --> ARGO["ArgoCD notices the commit"]
ARGO --> CLUSTER["cluster"]
classDef stop fill:#4a2222,stroke:#8f3d3d,color:#fff
class STOP stop
```
**CI holds no kubeconfig, and must never hold one.** Its last act is a git commit; ArgoCD
pulls. The maximum blast radius of a compromised pipeline is a bad commit, which is
revertable.
**Build once, promote the artifact.** The image is loaded locally, scanned, and only then
pushed — push-then-scan means a CRITICAL sits in the registry behind a green checkmark.
Deploys are **by digest**, never a mutable tag.
**Migrations never run on app startup.** N replicas would race. They run as a Helm
`pre-upgrade,pre-install` hook Job, once, before any new pod serves traffic. Forward-only,
expand/contract: a rename is three deploys.
### The chart works with or without ArgoCD
The chart is **pure Helm**. It contains no `argocd.argoproj.io/*` annotations, no sync
waves, and no ArgoCD-specific ordering. Both paths are supported and both are verified:
```bash
# Path 1: plain Helm, no ArgoCD anywhere
helm upgrade --install svcforge deploy/chart -n svcforge \
--set image.api.digest=sha256:... --set image.worker.digest=sha256:... \
--set image.reconciler.digest=sha256:...
# Path 2: GitOps. ArgoCD watches master and applies the same chart.
kubectl apply -f deploy/argocd/app.yaml
```
Ordering survives both because **ArgoCD translates Helm hooks into its own sync phases**
rather than ignoring them:
| Annotation | Plain Helm | ArgoCD |
|---|---|---|
| `helm.sh/hook: pre-install,pre-upgrade` | runs before the release, aborts it on failure | mapped to the **PreSync** phase |
| `helm.sh/hook-weight: "-5"` | orders hooks within the phase | mapped to hook ordering |
| `helm.sh/hook-delete-policy: before-hook-creation` | deletes the previous Job first | mapped to `BeforeHookCreation` |
Using `argocd.argoproj.io/hook` instead would have been the trap: plain `helm install`
does not understand that annotation, so it would create the migration Job as an ordinary
resource with no ordering guarantee — the migration and the new pods would start together,
and the failure would appear only in whichever path nobody tested.
Verified with `helm install --dry-run=server` against a real cluster, which validates every
manifest through the API server rather than only rendering the templates locally.
---
## User stories, and what each one exercises
| As a… | I want… | So that… | Exercised by |
|---|---|---|---|
| tenant team | to request an Elasticsearch without filing a ticket | I am unblocked in minutes | `POST /v1/instances` → 202 |
| tenant team | to see why my instance failed | I can fix my own request | `instances.error`, `svcforge status` |
| tenant team | a throwaway instance to clean itself up | I do not pay for what I forgot | `ttl_days` → reconciler TTL sweep |
| platform team | a worker pod to be killable at any instant | a rolling deploy is not an outage | SIGTERM drain + lease recovery |
| platform team | to patch a CVE across every tenant | one edit, not N | `catalog.yaml` bump → work list |
| platform team | a bad chart to stop after the first casualty | I do not break 40 tenants | `verify``rollout_state='halted'` |
| platform team | to know the queue is stuck before a tenant tells me | I look competent | `SvcforgeQueueDepthRising` → RUNBOOK |
| on-call | a copy-pasteable diagnosis at 3am | I do not have to think | [RUNBOOK.md](RUNBOOK.md) |
---
## Where to read the code
| To understand | Read | Then |
|---|---|---|
| the data model | `domain/models.py`, `migrations/001_init.sql` | `domain/states.py` |
| **the queue** | `repo/tasks.py` — read `_CLAIM_SQL` twice | `services/worker/main.py` |
| crash safety | `services/worker/handlers.py` | `tests/integration/test_worker.py` |
| the API contract | `services/api/routes/instances.py` | `tests/integration/test_api.py` |
| subprocess discipline | `adapters/helm.py::_run` | `tests/integration/test_helm_timeout.py` |
| day 2 | `domain/windows.py`, `InstanceRepo.list_upgradable` | `services/reconciler/main.py` |
| how it ships | `.gitea/workflows/ci.yaml`, `deploy/chart/` | [RUNBOOK.md](RUNBOOK.md) |
Related: [README.md](README.md) for how to read this repo without spoiling the course,
[RUNBOOK.md](RUNBOOK.md) for operating it.
+10
View File
@@ -0,0 +1,10 @@
# CLAUDE.md
The rules for this repo live in [AGENTS.md](AGENTS.md) — one file, so every tool reads the
same thing. Read it before your first edit.
The two that matter most:
1. This is a **reference implementation**. Never propose copying `domain/`, `repo/`,
`services/worker/`, or the claim query into `../learn-python/`. Those are the course.
2. **No back-and-forth writing style** in prose, comments, or docstrings. Plain declaratives.
+35 -11
View File
@@ -27,21 +27,45 @@ run-reconciler:
# Every gate CI runs, except the image build — that needs a docker daemon and three
# minutes, and the point of this target is to find out you failed before pushing.
# Same commands as .gitea/workflows/ci.yaml, deliberately: a gate that only exists in CI
# is a gate you debug through a web UI.
# is a gate you debug through a web UI. Each recipe line below is the command from the
# corresponding job, character for character. If you change one, change both.
ci:
uv run ruff check .
uv run ruff format --check .
uv run ruff check . && uv run ruff format --check .
uv run mypy --strict .
uv run pytest tests/unit --cov=libs/svcforge_core/domain --cov-fail-under=90
uv run pytest tests/unit --cov=svcforge_core.domain --cov-fail-under=90
uv run python -m svcforge_core.migrate
uv run pytest tests/integration
uv run --with bandit bandit -r libs services -ll
uv run --with pip-audit pip-audit --strict
uv run --with 'bandit[toml]' bandit -c pyproject.toml -r libs services -ll
uv export --frozen --no-dev \
--no-emit-project --no-emit-package svcforge-core \
-o /tmp/requirements-audit.txt
uv run --with pip-audit pip-audit --strict -r /tmp/requirements-audit.txt
hadolint services/*/Dockerfile
helm lint deploy/chart
helm template svcforge deploy/chart >/dev/null
$(MAKE) chart
# The chart's own gates. `helm template` is not a formality here: the image helper calls
# `fail` on any digest that is not a sha256, so this catches an unbumped values.yaml.
# The chart's own gates, mirroring the `chart` job in ci.yaml.
#
# `helm template` against plain values.yaml MUST FAIL: the digests there are all-zeros
# placeholders and _helpers.tpl refuses to build an image reference from one. So this
# asserts the failure rather than running the command bare — a bare `helm template` here
# would report the guard working as a broken build, and the previous version of this
# target did exactly that.
#
# The second render uses well-formed dummy digests to prove the templates themselves are
# valid, with both values-gated monitoring blocks turned on.
chart:
helm lint deploy/chart
helm template svcforge deploy/chart
@if helm template svcforge deploy/chart >/dev/null 2>&1; then \
echo "FAIL: chart rendered against the placeholder digests in values.yaml."; \
echo "The digest guard in _helpers.tpl is not guarding."; \
exit 1; \
fi
@echo "ok: placeholder digests rejected"
helm template svcforge deploy/chart \
--set image.api.digest=sha256:$(shell printf 'a%.0s' $$(seq 64)) \
--set image.worker.digest=sha256:$(shell printf 'b%.0s' $$(seq 64)) \
--set image.reconciler.digest=sha256:$(shell printf 'c%.0s' $$(seq 64)) \
--set serviceMonitor.enabled=true \
--set prometheusRule.enabled=true \
>/dev/null
@echo "ok: chart renders"
+5 -5
View File
@@ -11,11 +11,11 @@ A complete, working, verified build of the system `../learn-python/` teaches you
exactly that code. You asked for it deliberately, and the rule allows you to overrule it —
but the reason for the rule did not go away when you did.
Module 4 is five evenings of getting the claim query wrong. That is not a tax on the way to
the answer; it **is** the answer. The evening you spend watching two workers grab the same
task is the evening `FOR UPDATE SKIP LOCKED` stops being a phrase and starts being a thing
you understand. Reading `repo/tasks.py` here takes ninety seconds and teaches you close to
nothing, while feeling exactly like learning. That feeling is the trap.
Module 4 is five evenings of getting the claim query wrong, and those five evenings are the
learning. The evening you spend watching two workers grab the same task is the evening
`FOR UPDATE SKIP LOCKED` stops being a phrase and becomes something you understand.
Reading `repo/tasks.py` here takes ninety seconds, teaches you close to nothing, and feels
exactly like learning. That feeling is the trap.
So:
+1 -1
View File
@@ -39,7 +39,7 @@ Then set three repo Actions secrets (`Settings → Actions → Secrets`, or the
|---|---|---|
| `REGISTRY_USER` | `gitea_admin` | |
| `REGISTRY_TOKEN` | the PAT | **The auto-injected `GITEA_TOKEN` is rejected by the package registry with a 401.** This is the single most common reason a first pipeline fails at `docker push`. |
| `CI_BOT_TOKEN` | the PAT | Used only by the `bump` job to push the digest commit. Note what it is *not*: a kubeconfig. CI's maximum blast radius is a bad commit. |
| `CI_BOT_TOKEN` | the PAT | Used only by the `bump` job to push the digest commit. It is deliberately not a kubeconfig, so CI's maximum blast radius is a bad commit. |
### 2. The runner needs a cache server, and it fails SOFT without one
+13 -1
View File
@@ -42,7 +42,19 @@ spec:
# non-zero exit fails the sync instead of rolling out pods onto an unmigrated schema.
- ApplyOutOfSyncOnly=true
retry:
limit: 3
# 0, not 3, and for the same reason migrate-job.yaml sets backoffLimit: 0.
#
# An ArgoCD retry re-runs the WHOLE sync including the PreSync phase, and the migrate
# Job is a helm pre-install/pre-upgrade hook that ArgoCD maps onto PreSync. So
# `limit: 3` quietly reinstated the retry-a-failed-DDL behaviour that backoffLimit: 0
# exists to forbid — three attempts at the same failed migration, one readable error
# turned into three, against a schema that may now be half-applied.
#
# There is no recovery path here that a retry helps with. Recovery from a failed
# migration is a REVERT COMMIT: fix the SQL forward, push, and let ArgoCD sync the
# new revision. Rolling back the app image does not roll back DDL that already
# committed.
limit: 0
backoff:
duration: 20s
factor: 2
+12 -2
View File
@@ -53,8 +53,18 @@ comments — the acceptance gate greps for it and does not know what a comment i
{{- if not $img -}}
{{- fail (printf "no image config for component %q" .component) -}}
{{- end -}}
{{- if not (hasPrefix "sha256:" ($img.digest | default "")) -}}
{{- fail (printf "image.%s.digest must be a sha256 digest, not a tag — CI bumps it; got %q" .component ($img.digest | default "<empty>")) -}}
{{- $digest := $img.digest | default "" -}}
{{/*
Full-shape match, not `hasPrefix "sha256:"`. A prefix check accepts the all-zeros
placeholder in values.yaml, so a fresh clone rendered clean and the guard guarded nothing.
Two conditions, both required: the digest must be sha256: plus exactly 64 lowercase hex
characters, AND it must not be the placeholder literal.
*/}}
{{- if not (regexMatch "^sha256:[0-9a-f]{64}$" $digest) -}}
{{- fail (printf "image.%s.digest must be a sha256 digest (sha256: + 64 hex chars), not a tag — CI bumps it; got %q" .component ($digest | default "<empty>")) -}}
{{- end -}}
{{- if eq $digest "sha256:0000000000000000000000000000000000000000000000000000000000000000" -}}
{{- fail (printf "image.%s.digest is still the all-zeros placeholder from values.yaml — this chart has never been bumped by CI and must not be deployed" .component) -}}
{{- end -}}
{{- printf "%s@%s" $img.repo $img.digest -}}
{{- end -}}
+6
View File
@@ -59,6 +59,12 @@ spec:
{{- include "svcforge.env" . | nindent 12 }}
- name: OTEL_SERVICE_NAME
value: svcforge-migrate
# Where services/api/Dockerfile copies migrations/ to. The Dockerfile sets the
# same value as an ENV; this states it in the manifest as well so the path is
# visible to anyone reading the Job rather than only to whoever opens the
# image. The two MUST agree — if one moves, move both.
- name: SVCFORGE_MIGRATIONS_DIR
value: /app/migrations
resources:
{{- toYaml .Values.migrate.resources | nindent 12 }}
volumeMounts:
+11 -1
View File
@@ -6,6 +6,14 @@ is banned: it drifts from the chart, survives a `helm uninstall`, and nothing ow
The api is scraped through its Service. The worker and reconciler have no Service — they
are scraped by pod, which is why their metrics port is named and their pods carry the
component label.
A ServiceMonitor's selector matches the SERVICE OBJECT's own metadata labels, not the
Service's pod selector. api-service.yaml labels itself with `svcforge.labels` plus
`app.kubernetes.io/component: api` — it does NOT carry the `app: api` that
`svcforge.selectorLabels` adds, because that label exists for the chaos experiments'
`kubectl delete pod -l app=worker` and belongs on pods. So this must not reuse
selectorLabels: doing so matched nothing and api metrics were never scraped. These three
keys are exactly the ones the Service metadata carries, and are enough to be unambiguous.
*/}}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
@@ -16,7 +24,9 @@ metadata:
spec:
selector:
matchLabels:
{{- include "svcforge.selectorLabels" (dict "ctx" $ "component" "api") | nindent 6 }}
app.kubernetes.io/name: {{ include "svcforge.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: api
endpoints:
- port: http
path: /metrics
+13 -3
View File
@@ -25,9 +25,19 @@ spec:
{{- toYaml . | nindent 8 }}
{{- end }}
# The load-bearing one. On SIGTERM the loop stops claiming and finishes the task in
# hand; 60s is the budget for that. Chaos experiment 2 asserts the pod exits 0 inside
# it. Lower this and a rolling deploy starts orphaning tasks to lease expiry.
terminationGracePeriodSeconds: 60
# hand. Chaos experiment 2 asserts the pod exits 0 inside it. Lower this and a
# rolling deploy starts orphaning tasks to lease expiry.
#
# 660s, and the number is derived, not chosen. The task being drained is a
# `helm upgrade --install`, and this MUST exceed the longest that call can run before
# it returns:
# - settings.py helm_timeout_s = 300 (default)
# - HelmProvisioner timeout_s = 600 (default, the ceiling)
# 600s is the value to beat; 660 leaves a 60s margin for the pool to close and the
# process to exit cleanly. At the old 60s the kubelet SIGKILLed the worker mid-helm —
# precisely the orphaning this setting exists to prevent. If either timeout above is
# raised, raise this past it first.
terminationGracePeriodSeconds: 660
securityContext:
{{- include "svcforge.podSecurityContext" . | nindent 8 }}
containers:
+17 -3
View File
@@ -148,15 +148,29 @@ prometheusRule:
annotations:
summary: svcforge p95 provision time is over 5 minutes
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#provision-failing
- alert: SvcforgeTaskFailed
expr: sum(svcforge_tasks_failed_total) > 0
# `increase(...[15m])`, not the raw counter. `sum(counter) > 0` on a monotonic counter
# latches: one dead-lettered task at any point keeps this firing until the pod restarts,
# and a restart silently clears it — so it can never distinguish "failing now" from
# "failed last Tuesday". The dead-letter counter is also the right one: the attempts
# counter increments on ordinary transient retries that later succeed.
- alert: SvcforgeTaskDeadLettered
expr: sum(increase(svcforge_tasks_dead_lettered_total[15m])) > 0
for: 5m
labels:
severity: warning
annotations:
summary: a svcforge task exhausted its retries
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#provision-failing
# Scoped to the reconciler job, and aggregated with max().
#
# RECONCILER_LAST_TICK is a module-level Gauge in obs.py, so EVERY service that imports
# svcforge_core.obs registers and exports it — api and worker included, permanently at
# 0. Unscoped, `time() - 0` is ~1.7e9, so this critical alert fires from the moment the
# chart is installed, from pods that have no reconciler in them. Scope by job, then
# max() so a rolling restart of the single reconciler does not flap it.
- alert: SvcforgeReconcilerStale
expr: time() - svcforge_reconciler_last_tick_timestamp_seconds > 300
expr: time() - max(svcforge_reconciler_last_tick_timestamp_seconds{job=~".*reconciler.*"}) > 300
for: 5m
labels:
severity: critical
annotations:
@@ -32,6 +32,7 @@ import yaml
from pydantic import BaseModel, ConfigDict, Field
from svcforge_core.domain.models import CatalogEntry
from svcforge_core.errors import SvcforgeError
# How long the process group gets to honour SIGTERM before SIGKILL. Helm traps SIGTERM
# and tries to leave the release in a coherent state; give it a moment to do so.
@@ -45,8 +46,12 @@ _RUN_TIMEOUT_MARGIN_S = 30
_STDERR_TAIL_BYTES = 2048
class HelmError(RuntimeError):
"""Non-zero exit. str(self) is the stderr tail that lands in instances.error."""
class HelmError(SvcforgeError, RuntimeError):
"""Non-zero exit. str(self) is the stderr tail that lands in instances.error.
RuntimeError stays in the MRO so callers written against it keep catching; SvcforgeError
comes first so `except SvcforgeError` can separate a modelled failure from a stray bug.
"""
class ReleaseInfo(BaseModel):
@@ -166,6 +171,19 @@ class HelmProvisioner:
argv += ["--kubeconfig", str(self._kubeconfig)]
return argv
async def _run_helm(self, argv: Sequence[str]) -> str:
"""`_run`, with the timeout path translated to this adapter's declared error type.
`_run` raises a bare TimeoutError so that the process-group test can assert on it
directly, but every public method here is documented as raising HelmError; a wedged
helm arriving as TimeoutError sails straight past a caller's `except HelmError` and
fails the task as an unmodelled crash. Translate once, at the public boundary.
"""
try:
return await _run(argv, timeout_s=self._run_timeout_s)
except TimeoutError as exc:
raise HelmError(f"{argv[0]} timed out after {self._run_timeout_s}s and was killed") from exc
async def install(self, release: str, ns: str, entry: CatalogEntry, values: dict[str, Any]) -> None:
"""helm upgrade --install --wait --timeout. Idempotent by construction."""
# `upgrade --install` is why this is idempotent: a retried task after a crash mid-provision
@@ -190,7 +208,7 @@ class HelmProvisioner:
"--timeout",
f"{self._timeout_s}s",
)
await _run(argv, timeout_s=self._run_timeout_s)
await self._run_helm(argv)
async def uninstall(self, release: str, ns: str) -> None:
"""helm uninstall --wait. `--ignore-not-found` makes the retry of a half-done delete a no-op."""
@@ -204,12 +222,12 @@ class HelmProvisioner:
"--timeout",
f"{self._timeout_s}s",
)
await _run(argv, timeout_s=self._run_timeout_s)
await self._run_helm(argv)
async def list_releases(self) -> list[ReleaseInfo]:
"""Every release helm knows about, in every namespace. The reconciler's view of reality."""
argv = self._base_argv("list", "--all-namespaces", "--output", "json")
raw = await _run(argv, timeout_s=self._run_timeout_s)
raw = await self._run_helm(argv)
try:
parsed: Any = json.loads(raw or "[]")
except json.JSONDecodeError as exc:
@@ -21,12 +21,17 @@ from typing import Any
import yaml
from svcforge_core.adapters.helm import HelmError, _run
from svcforge_core.errors import SvcforgeError
_KUBECTL_TIMEOUT_S = 60
class K8sError(RuntimeError):
"""kubectl failed. str(self) is the stderr tail, already truncated by `_run`."""
class K8sError(SvcforgeError, RuntimeError):
"""kubectl failed. str(self) is the stderr tail, already truncated by `_run`.
RuntimeError stays in the MRO so existing `except RuntimeError` callers keep catching;
SvcforgeError comes first so a modelled cluster failure is distinguishable from a bug.
"""
class SecretNotFound(K8sError):
@@ -112,6 +117,11 @@ class KubectlClient:
return await _run(self._base_argv(*args), timeout_s=self._timeout_s)
except HelmError as exc:
raise K8sError(str(exc)) from exc
except TimeoutError as exc:
# `_run` re-raises the bare TimeoutError after killing the process group, so the
# timeout path does not come through HelmError. Without this clause a wedged
# kubectl surfaces as TimeoutError past a caller written to `except K8sError`.
raise K8sError(f"kubectl {args[0] if args else ''} timed out after {self._timeout_s}s") from exc
class _ManifestFile:
@@ -11,15 +11,29 @@ tests and local dev get) and WebhookNotifier (the one that leaves the process).
from __future__ import annotations
import logging
from typing import Protocol
import httpx
_log = logging.getLogger(__name__)
from svcforge_core import obs
# obs imports nothing from adapters (it is settings + structlog + otel only), so this is a
# plain top-level import and not a lazy one — the cycle it would otherwise create does not
# exist. Keep it that way: obs must stay importable before any adapter is.
_DEFAULT_TIMEOUT_S = 5.0
# Keys a caller's `fields` must not be allowed to occupy. `event` is structlog's first
# positional parameter, so passing it as a kwarg is a TypeError, not a shadowed value; the
# rest are stdlib LogRecord attributes that the ProcessorFormatter bridge refuses to
# overwrite. A tenant-supplied field named `event` must not be able to crash the notifier.
_RESERVED_KEYS = frozenset({"event", "msg", "args", "name", "levelname", "exc_info", "stack_info"})
def _safe_fields(fields: dict[str, str] | None) -> dict[str, str]:
"""Fields with reserved names prefixed rather than dropped — the value still gets logged."""
return {(f"field_{k}" if k in _RESERVED_KEYS else k): v for k, v in (fields or {}).items()}
class Notifier(Protocol):
"""Implemented by LogNotifier, WebhookNotifier, and FakeNotifier (tests/fakes.py)."""
@@ -33,11 +47,20 @@ class LogNotifier:
"""Writes the event to the log. The default: structured logs are already shipped somewhere."""
async def send(self, event: str, message: str, fields: dict[str, str] | None = None) -> None:
# NOT extra={"message": ...}. `message` is a reserved LogRecord attribute, and
# logging raises KeyError("Attempt to overwrite 'message' in LogRecord") at call
# time — so the notifier would crash the caller it was meant to inform. Same trap
# for `msg`, `args`, `name`, `levelname`, `exc_info`.
_log.info("notify", extra={"event": event, "detail": message, "fields": fields or {}})
# structlog kwargs, NOT logging's `extra=`. obs bridges stdlib records through
# ProcessorFormatter, which builds the event dict from `record.msg` alone — every
# key passed via `extra=` is dropped on the floor, so the default notifier used to
# emit a bare {"event": "notify"} with the payload gone.
#
# Fields are splatted rather than nested under "fields" so each one is its own
# queryable key in Loki. `detail`, not `message`: `message` is a reserved LogRecord
# attribute and the stdlib bridge raises KeyError on it.
#
# `notify_event`, not `event`: structlog's first positional parameter IS named
# `event` (it becomes the rendered line's "event" key, here the literal "notify"),
# so passing event= alongside it is a TypeError at the call, not a rename.
log = obs.get_logger(__name__)
log.info("notify", notify_event=event, detail=message, **_safe_fields(fields))
class WebhookNotifier:
@@ -52,26 +75,33 @@ class WebhookNotifier:
"""
self._url = url
self._timeout_s = timeout_s
self._client = client
self._owns_client = client is None
async def _get_client(self) -> httpx.AsyncClient:
if self._client is None:
self._client = httpx.AsyncClient(timeout=self._timeout_s)
return self._client
# Eager, not lazy. `AsyncClient()` does no I/O, so laziness bought nothing and cost a
# race: two concurrent `send`s could both see None, both construct a client, and the
# loser's connection pool would leak because only one of them survived the assignment.
self._client = client if client is not None else httpx.AsyncClient(timeout=self._timeout_s)
async def send(self, event: str, message: str, fields: dict[str, str] | None = None) -> None:
payload = {"event": event, "message": message, "fields": fields or {}}
try:
client = await self._get_client()
resp = await client.post(self._url, json=payload, timeout=self._timeout_s)
resp = await self._client.post(self._url, json=payload, timeout=self._timeout_s)
resp.raise_for_status()
except httpx.HTTPError as exc:
# Deliberately swallowed. See the module docstring: the work already succeeded.
_log.warning("notify webhook failed", extra={"event": event, "error": str(exc)})
except Exception as exc: # the bare `except Exception` IS the specification here
# `send` must not raise; that is the contract in the module docstring, and it is
# not satisfiable by catching httpx.HTTPError alone. `httpx.InvalidURL` is not an
# HTTPError subclass, and posting on an already-aclose()d client raises
# RuntimeError — so a typo'd webhook URL would fail a task whose helm work has
# already succeeded. exc_info so the traceback survives the swallowing.
obs.get_logger(__name__).warning(
"notify webhook failed", notify_event=event, error=str(exc), exc_info=exc
)
async def aclose(self) -> None:
"""Close the client, if we made it."""
if self._client is not None and self._owns_client:
"""Close the client, if we made it. Call at process shutdown, next to the pool's close.
The client reference is kept rather than cleared: a `send` that races shutdown now
raises RuntimeError on a closed client, and `send` swallows and logs that like any
other delivery failure instead of resurrecting a pool nobody will close.
"""
if self._owns_client:
await self._client.aclose()
self._client = None
@@ -5,9 +5,8 @@ holds the instances, the tasks, the leases and the `release_name` UNIQUE constra
holds a counter, a claim marker and a copy — all of it rebuildable by doing nothing and
waiting for a TTL.
That framing is not philosophy, it decides the error handling, and the error handling is
the module. Each class below catches `RedisError` and returns a *safe* answer rather than
raising:
That framing decides the error handling, and the error handling is the module. Each class
below catches `RedisError` and returns a *safe* answer rather than raising:
| Path | Redis is down | Why |
|-------------|--------------------------|--------------------------------------------------|
@@ -22,7 +21,7 @@ Consequently nothing here raises out to a caller, and `/readyz` stays Postgres-o
Redis outage must not make a single pod unready — that would convert "the cache is down"
into "the platform is down", which is the exact inversion this module exists to prevent.
**The budget is a design constraint, not a footnote.** Upstash free tier:
**The budget is a design constraint.** Upstash free tier:
500,000 commands / month = 16,129 / day = 11 / minute = 0.19 / second, sustained
@@ -0,0 +1,21 @@
"""The one base class every svcforge-raised exception shares.
Without it, a caller that wants "the cluster failed" has to write `except Exception`, which
also swallows the `AttributeError` from a typo three frames down. The two are not the same
incident: one is retried, the other is a bug that must reach the dead-letter loudly. A single
root makes that distinction expressible in one clause.
Subclasses keep their existing stdlib base as well (`HelmError(SvcforgeError, RuntimeError)`),
so code already written against `except RuntimeError` keeps working. The MRO order matters:
`SvcforgeError` first, so the svcforge-specific class is the more derived one.
"""
from __future__ import annotations
class SvcforgeError(Exception):
"""Root of every exception svcforge raises on purpose.
Catch this to mean "an operation svcforge models failed". Anything not deriving from it
that escapes a handler is, by definition, a programming error rather than a modelled one.
"""
+25 -1
View File
@@ -13,6 +13,7 @@ Run: python -m svcforge_core.migrate
from __future__ import annotations
import os
import sys
from pathlib import Path
@@ -20,7 +21,30 @@ import psycopg
from svcforge_core.settings import load_settings
MIGRATIONS_DIR = Path(__file__).resolve().parents[3] / "migrations"
def _migrations_dir() -> Path:
"""Where the .sql files live, in a way that survives being installed as a wheel.
`Path(__file__).parents[3] / "migrations"` works only for the editable dev install,
where the package really does sit three levels under the repo root. The images install
a built wheel into site-packages (deliberately — an editable install in an image ships
a path, not a package), so the same expression resolves to
`/app/.venv/lib/migrations`, which does not exist. The migration Job then finds zero
files and exits 1, and the Helm pre-upgrade hook fails on every single sync.
That failure is invisible in dev and in CI, because both run editable. It only appears
the first time you deploy — which is the worst time to discover it.
So: the image sets SVCFORGE_MIGRATIONS_DIR and copies the directory in; the repo
checkout falls back to the path relative to this file.
"""
env = os.getenv("SVCFORGE_MIGRATIONS_DIR")
if env:
return Path(env)
return Path(__file__).resolve().parents[3] / "migrations"
MIGRATIONS_DIR = _migrations_dir()
# Advisory lock: two Jobs racing (a retried hook, a hand-run) must not both apply DDL.
# Session-scoped, so this needs the session pooler (5432), not pgbouncer (6543).
+23 -4
View File
@@ -62,9 +62,17 @@ TASKS_CLAIMED = Counter(
["kind"],
)
TASKS_FAILED = Counter(
"svcforge_tasks_failed_total",
"Tasks that exhausted their attempts and went to 'failed'.",
TASK_ATTEMPTS_FAILED = Counter(
"svcforge_task_attempts_failed_total",
"Task ATTEMPTS that raised. Not the same as tasks that dead-lettered: one task that "
"succeeds on its third try increments this twice. Alert on increase(), not on the raw "
"total, or ordinary transient retries page you.",
["kind"],
)
TASKS_DEAD_LETTERED = Counter(
"svcforge_tasks_dead_lettered_total",
"Tasks that exhausted max_attempts and went to 'failed'. These need a human.",
["kind"],
)
@@ -73,7 +81,7 @@ PROVISION_TIME = Histogram(
"Wall time of a provision task, claim to terminal report.",
# NOT the defaults. See the module docstring: the defaults end at 10s and a provision
# takes minutes. The top finite bucket is 1800 because helm's own --timeout is 600 and
# a provision past thirty minutes is not slow, it is broken and belongs in +Inf.
# a provision past thirty minutes is broken and belongs in +Inf.
buckets=(10, 30, 60, 120, 300, 600, 1800, float("inf")),
)
@@ -97,6 +105,9 @@ RECONCILER_LAST_TICK = Gauge(
_TRACER_NAME = "svcforge"
# Set by setup(); re-bound by bind_task_context after it clears the context.
_service_name: str = "svcforge"
# setup() is idempotent because it is called from three entrypoints and from tests, and
# because configuring structlog twice silently discards the first configuration while
# adding a second stdout handler to the root logger — every line then prints twice.
@@ -194,6 +205,11 @@ def _setup_logging(service_name: str, settings: Settings) -> None:
root.handlers = [handler]
root.setLevel(level)
# Remembered so bind_task_context can restore it after clearing. Without this, every
# log line emitted inside a task loses `service`, and those are exactly the lines you
# filter on when you are trying to tell worker output from reconciler output.
global _service_name
_service_name = service_name
structlog.contextvars.bind_contextvars(service=service_name)
@@ -278,6 +294,9 @@ def bind_task_context(instance_id: UUID, task_id: int, team: str) -> None:
"""
structlog.contextvars.clear_contextvars()
structlog.contextvars.bind_contextvars(
# `service` is re-bound because the clear above took it with it. It is set once in
# setup() and is not per-task, but clear_contextvars() is indiscriminate.
service=_service_name,
instance_id=str(instance_id),
task_id=task_id,
team=team,
@@ -129,14 +129,6 @@ class InstanceRepo:
)
return cur.rowcount == 1
async def set_error(self, id: UUID, error: str) -> None:
"""Record a terminal failure message. Used when a task exhausts its attempts."""
async with self._pool.connection() as conn, conn.cursor() as cur:
await cur.execute(
"update instances set error = %s, state = %s, updated_at = now() where id = %s",
(error[-2000:], InstanceState.FAILED.value, id),
)
async def list_upgradable(
self,
service_type: str,
@@ -153,8 +145,8 @@ class InstanceRepo:
rollouts table, no state machine, no pause/resume CLI. You clear it with SQL.
`order by team = %(own_team)s desc` — your own instances upgrade first, so you are
the tenant who finds out the chart is broken. Eating your own dog food is a
`order by`, not a policy document.
the tenant who finds out the chart is broken. Eating your own dog food is enforced
by an `order by` rather than a policy document.
`limit %(max_in_flight)s` — a config value, not a scheduler. It stays at 1 until 1
is too slow, and 1 is what makes the halt meaningful: the fleet stops after the
@@ -93,7 +93,7 @@ class ReconcileRepo:
None if the row moved, or if a provision is already outstanding.
The two-hop state change is the interesting part, and it is not busywork:
The two-hop state change is the interesting part:
* `LEGAL` has no `ready -> provisioning` edge. The tenant-visible lifecycle only
leaves `ready` through `deleting` or `failed`, and drift is a failure — the
+47 -19
View File
@@ -19,7 +19,7 @@ from psycopg import AsyncConnection
from svcforge_core.domain.backoff import next_attempt_at
from svcforge_core.domain.models import Task, TaskKind
from svcforge_core.domain.states import LEGAL, InstanceState
from svcforge_core.obs import inject_traceparent
from svcforge_core.obs import TASKS_DEAD_LETTERED, inject_traceparent
from svcforge_core.repo.db import DictPool
# Which states may legally become `failed`, derived from the domain's own table rather
@@ -49,6 +49,13 @@ _CAN_FAIL: Final[tuple[str, ...]] = tuple(
# before it has loaded anything. A data-modifying CTE runs exactly once and cannot claim
# twice. The alternative — a second SELECT for the team — would be a second round trip per
# task to fetch a column the database already had in hand.
#
# LEFT join, not inner. The UPDATE inside the CTE has already taken effect by the time the
# outer select runs, so an inner join that matches nothing would return no row — and
# `claim()` would report "queue empty" for a task it had just marked `running`, stranding
# it until the lease expires and silently burning an attempt. The FK cascade makes that
# nearly impossible in practice; "nearly" is not a reason to leave a silent failure in the
# one query the whole system depends on. `Task.team` is already `str | None`.
_CLAIM_SQL = """
with claimed as (
update tasks set state='running', attempts=attempts+1, locked_by=%(worker)s, locked_at=now()
@@ -62,7 +69,7 @@ with claimed as (
returning *
)
select claimed.*, instances.team
from claimed join instances on instances.id = claimed.instance_id;
from claimed left join instances on instances.id = claimed.instance_id;
"""
@@ -135,24 +142,33 @@ class TaskRepo:
async def complete(
self,
task_id: int,
worker_id: str,
conn: AsyncConnection[dict[str, Any]] | None = None,
) -> None:
"""Mark done. Pass `conn` to commit alongside the caller's instance update.
) -> bool:
"""Mark done, but only if this worker still owns the task. False if it does not.
Completing the task and recording what it accomplished belong in one transaction:
commit them separately and a crash in between leaves a task marked done whose
work never landed, or work that landed and will be redone.
Pass `conn` to commit alongside the caller's instance update: completing the task
and recording what it accomplished belong in one transaction, or a crash between
them leaves a task marked done whose work never landed.
`and state='running' and locked_by=%s` is not defensive padding — without it this
is a lost-update bug with a real trigger. A worker that hangs past `lease_seconds`
has its task requeued by the reconciler and re-claimed by someone else. When the
hung worker finally returns, an unconditional UPDATE here marks the task `done`
while the new owner is still running it, and its work goes unaccounted for. The
loser gets False and must treat it as "someone else owns this now", not an error.
"""
sql = "update tasks set state='done', locked_by=null where id = %s"
sql = "update tasks set state='done', locked_by=null where id=%s and state='running' and locked_by=%s"
if conn is not None:
async with conn.cursor() as cur:
await cur.execute(sql, (task_id,))
return
await cur.execute(sql, (task_id, worker_id))
return cur.rowcount == 1
async with self._pool.connection() as own, own.cursor() as cur:
await cur.execute(sql, (task_id,))
await cur.execute(sql, (task_id, worker_id))
return cur.rowcount == 1
async def fail(self, task_id: int, err: str, max_attempts: int = 5) -> None:
"""Retry with backoff, or give up.
async def fail(self, task_id: int, err: str, worker_id: str, max_attempts: int = 5) -> bool:
"""Retry with backoff, or give up. False if this worker no longer owns the task.
Under max_attempts: back to 'queued' with run_after pushed out by exponential
backoff with full jitter. Jitter matters — a cluster-wide outage fails every task
@@ -161,17 +177,25 @@ class TaskRepo:
At max_attempts: 'failed', and the error is copied onto the instance so the tenant
can see it. A dead-letter state, not an infinite retry: a task that cannot succeed
must stop and become someone's problem.
The ownership check in the SELECT is the same lost-lease guard as `complete`, and
it matters more here: a stale worker reporting failure would push a task the new
owner is actively running back to `queued`, letting a *third* worker claim it.
"""
now = datetime.now(UTC)
async with self._pool.connection() as conn:
async with conn.transaction(), conn.cursor() as cur:
await cur.execute(
"select attempts, instance_id from tasks where id = %s for update",
(task_id,),
"""select attempts, instance_id, kind from tasks
where id=%s and state='running' and locked_by=%s
for update""",
(task_id, worker_id),
)
row = await cur.fetchone()
if row is None:
return
# Either the task is gone, or the lease was stolen. Both mean: not ours
# to report on. Writing anything here would corrupt the new owner's run.
return False
attempts = int(row["attempts"])
instance_id = row["instance_id"]
@@ -183,7 +207,7 @@ class TaskRepo:
where id = %s""",
(err[-2000:], next_attempt_at(attempts - 1, now=now), task_id),
)
return
return True
await cur.execute(
"""update tasks
@@ -199,14 +223,18 @@ class TaskRepo:
where id=%s and state = any(%s)""",
(err[-2000:], InstanceState.FAILED.value, instance_id, list(_CAN_FAIL)),
)
# Counted here, not in the worker: this is the only place that knows the
# difference between "attempt 2 of 5 failed" and "this task is done trying".
TASKS_DEAD_LETTERED.labels(kind=str(row["kind"])).inc()
return True
async def reset_expired_leases(self, lease_seconds: int) -> int:
"""Return tasks whose worker died back to the queue. Called by the reconciler.
No distributed lock survives a power cut. A worker that is SIGKILLed leaves
`state='running'` and `locked_by` set with nobody running it, and that row would
sit there forever. The lease is the only thing that recovers it: not a lock, a
timeout. This is why `locked_at` exists.
sit there forever. The lease is the only thing that recovers it, which is why
`locked_at` exists.
"""
async with self._pool.connection() as conn, conn.cursor() as cur:
await cur.execute(
+23 -2
View File
@@ -23,6 +23,11 @@ class Settings(BaseSettings):
frozen=True,
)
# Anything other than "local" makes check_production() enforce. The chart sets it;
# a laptop does not. Defaulting to "local" means a forgotten env var costs you a
# refused dev shortcut, never an unauthenticated production API.
environment: str = "local"
# --- Postgres -----------------------------------------------------------------
# Transaction pooler (6543 on Supabase). Everything the services do at runtime.
pg_dsn: PostgresDsn
@@ -35,6 +40,9 @@ class Settings(BaseSettings):
# --- Redis (derived state only; never the source of truth) ---------------------
redis_dsn: RedisDsn | None = None
# Per-team request budget. Generous on purpose: this exists to stop one team's runaway
# script from starving the others, not to meter usage.
rate_limit_per_minute: int = Field(default=60, ge=1)
# --- API ----------------------------------------------------------------------
jwks_url: str | None = None
@@ -81,9 +89,22 @@ class Settings(BaseSettings):
return str(self.pg_dsn_session or self.pg_dsn)
def check_production(self) -> None:
"""Refuse the dev escape hatches when they would matter."""
"""Refuse the dev escape hatches outside local development. Call at startup.
This is a no-op unless `SVCFORGE_ENVIRONMENT` says otherwise, which is what makes
it safe to call unconditionally from every entrypoint — and calling it
unconditionally is the point. The previous version could only be invoked from a
branch that already knew it was production, so no such branch was ever written and
the check never ran: `SVCFORGE_AUTH_DISABLED=true` in prod would have started the
API with JWT verification off, serving every unauthenticated request as team
`platform`, silently.
"""
if self.environment == "local":
return
if self.auth_disabled:
raise ValueError("SVCFORGE_AUTH_DISABLED=true is refused outside local development")
raise ValueError(
f"SVCFORGE_AUTH_DISABLED=true is refused when SVCFORGE_ENVIRONMENT={self.environment!r}"
)
def load_settings() -> Settings:
+22
View File
@@ -0,0 +1,22 @@
-- 005_task_indexes.sql — index the lookups the control loops actually do.
--
-- Postgres does NOT auto-index the referencing side of a foreign key. `tasks.instance_id`
-- had no index, and three hot paths look rows up by it:
--
-- * ReconcileRepo._has_unfinished — the idempotency guard on every enqueue, so once per
-- candidate per 60s tick.
-- * due_for_deprovision's `not exists` correlated subquery — once per instance per tick.
-- * the ON DELETE CASCADE itself, on every instance delete.
--
-- `tasks` is append-only in practice (done/failed rows are never pruned by the app; the
-- runbook prunes them by hand), so a sequential scan there gets slower every day the
-- system runs. This is the index that stops a 60-second control loop degrading into a
-- full table scan of all history.
--
-- Composite on (instance_id, kind, state) because _has_unfinished filters all three.
create index tasks_by_instance on tasks (instance_id, kind, state);
-- `instances.state` is filtered by ready_instances, instance_counts, due_for_deprovision
-- and list_upgradable — every reconciler check. Partial would not help: the sweeps look
-- for different states, so the whole column earns its index.
create index instances_by_state on instances (state);
+56 -29
View File
@@ -1,9 +1,9 @@
"""Throwaway load generator. Enqueue N instances, watch the queue drain, print three numbers.
The point is not a benchmark. It is to find the ceiling on purpose, in a place where finding
it is free, so that the number in RUNBOOK.md comes from an observation instead of a guess.
The point is to find the ceiling on purpose, in a place where finding it is free, so that
the number in RUNBOOK.md comes from an observation instead of a guess.
The ceiling you are looking for is arithmetic, not mysterious:
The ceiling you are looking for is arithmetic:
total connections = (api_replicas + worker_replicas) x pool_max_size
@@ -23,20 +23,28 @@ import argparse
import asyncio
import time
from datetime import UTC, datetime
from typing import Any
from uuid import uuid4
import psycopg
from psycopg.rows import dict_row
from svcforge_core.domain.catalog import load_catalog
from svcforge_core.settings import load_settings
_SERVICE_TYPE = "elasticsearch"
async def _seed_direct(dsn: str, count: int) -> float:
async def _seed_direct(dsn: str, count: int, chart_version: str) -> float:
"""Insert `count` instances + provision tasks. Returns seconds taken.
--direct exists to separate two questions that a single POST run conflates: "how fast can
the API accept work" and "how fast can workers drain it". Measure them apart or you will
tune the wrong one.
`chart_version` comes from the catalog rather than a literal. Hardcoding it meant the
seeded rows carried a version the catalog could not resolve, so every task failed fast
and the drain measurement — the whole point of the script — timed the failure path.
"""
started = time.monotonic()
async with await psycopg.AsyncConnection.connect(dsn, row_factory=dict_row) as conn:
@@ -46,9 +54,9 @@ async def _seed_direct(dsn: str, count: int) -> float:
await cur.execute(
"""insert into instances (id, team, service_type, size, state, namespace,
release_name, chart_version)
values (%s, 'loadtest', 'elasticsearch', 'small', 'requested',
'tenant-loadtest', %s, '21.3.19')""",
(iid, f"loadtest-elasticsearch-{str(iid)[:8]}"),
values (%s, 'loadtest', %s, 'small', 'requested',
'tenant-loadtest', %s, %s)""",
(iid, _SERVICE_TYPE, f"loadtest-{_SERVICE_TYPE}-{str(iid)[:8]}", chart_version),
)
await cur.execute(
"insert into tasks (instance_id, kind) values (%s, 'provision')",
@@ -57,38 +65,49 @@ async def _seed_direct(dsn: str, count: int) -> float:
return time.monotonic() - started
async def _depth(dsn: str) -> dict[str, int]:
async with await psycopg.AsyncConnection.connect(dsn, row_factory=dict_row) as conn:
cur = await conn.execute("select state, count(*) as n from tasks group by 1")
return {str(r["state"]): int(r["n"]) for r in await cur.fetchall()}
async def _depth(conn: psycopg.AsyncConnection[dict[str, Any]]) -> dict[str, int]:
"""Task counts by state, on a connection the caller owns."""
cur = await conn.execute("select state, count(*) as n from tasks group by 1")
return {str(r["state"]): int(r["n"]) for r in await cur.fetchall()}
async def _watch(dsn: str, timeout_s: float) -> None:
"""Print queue depth once a second until it drains. The slope is the number you want."""
"""Print queue depth once a second until it drains. The slope is the number you want.
One connection for the whole loop, held open. Reconnecting every second added a
connection to the pooler on every tick of a script whose entire purpose is finding the
connection ceiling — the measurement was perturbing the thing being measured.
"""
started = time.monotonic()
peak = 0
print(f"{'t(s)':>6} {'queued':>7} {'running':>8} {'done':>6} {'failed':>7} slope/s")
prev_done, prev_t = 0, started
while time.monotonic() - started < timeout_s:
d = await _depth(dsn)
queued, running = d.get("queued", 0), d.get("running", 0)
done, failed = d.get("done", 0), d.get("failed", 0)
peak = max(peak, queued + running)
# autocommit: a held connection without it sits idle-in-transaction between polls, which
# pins a snapshot on the pooler and is exactly the pathology this script hunts for.
async with await psycopg.AsyncConnection.connect(dsn, row_factory=dict_row, autocommit=True) as conn:
while time.monotonic() - started < timeout_s:
d = await _depth(conn)
queued, running = d.get("queued", 0), d.get("running", 0)
done, failed = d.get("done", 0), d.get("failed", 0)
peak = max(peak, queued + running)
now = time.monotonic()
slope = (done - prev_done) / max(now - prev_t, 1e-9)
prev_done, prev_t = done, now
now = time.monotonic()
slope = (done - prev_done) / max(now - prev_t, 1e-9)
prev_done, prev_t = done, now
print(f"{now - started:6.1f} {queued:7d} {running:8d} {done:6d} {failed:7d} {slope:7.1f}")
print(f"{now - started:6.1f} {queued:7d} {running:8d} {done:6d} {failed:7d} {slope:7.1f}")
if queued == 0 and running == 0:
elapsed = now - started
print(f"\ndrained in {elapsed:.1f}s peak depth {peak} throughput {done / elapsed:.1f} task/s")
if failed:
print(f"WARNING: {failed} tasks failed — the number above is not a clean drain")
return
await asyncio.sleep(1.0)
if queued == 0 and running == 0:
elapsed = now - started
print(
f"\ndrained in {elapsed:.1f}s peak depth {peak} "
f"throughput {done / elapsed:.1f} task/s"
)
if failed:
print(f"WARNING: {failed} tasks failed — the number above is not a clean drain")
return
await asyncio.sleep(1.0)
print(f"\nstill draining after {timeout_s}s — that IS the result. Record it.")
@@ -117,8 +136,16 @@ async def _amain() -> None:
"with k6 (one dependency, not two — do not add locust for this)."
)
# The version the workers will actually resolve. Read it rather than restate it: a seeded
# row whose chart_version disagrees with the catalog drains through the failure path.
catalog = load_catalog(settings.catalog_path)
entry = catalog.get(_SERVICE_TYPE)
if entry is None:
raise SystemExit(f"{settings.catalog_path} has no '{_SERVICE_TYPE}' entry to load-test with")
print(f"seeding {args.count} instances at {datetime.now(UTC).isoformat()} ...")
took = await _seed_direct(dsn, args.count)
print(f" service_type={_SERVICE_TYPE} chart_version={entry.chart_version} (from catalog)")
took = await _seed_direct(dsn, args.count, entry.chart_version)
print(f"enqueued {args.count} in {took:.2f}s ({args.count / took:.0f}/s)\n")
if args.watch:
+61 -13
View File
@@ -3,8 +3,17 @@
$ python3 scripts/redis_budget.py
$ python3 scripts/redis_budget.py --url http://localhost:8000/metrics --budget 500000
$ python3 scripts/redis_budget.py --url http://api:8000/metrics \
--url http://worker:9100/metrics \
--url http://reconciler:9100/metrics
Upstash's free tier is 500,000 commands/month, which sounds enormous and is not:
The budget is per Upstash database; the counters are per process. api, worker and
reconciler each keep their own registry (see obs.py — one process per pod, no multiproc
directory), so scraping one endpoint measures one third of the burn. Repeat `--url` to sum
them; a run that covers fewer than three sources says so in its output rather than printing
a reassuring number derived from one process.
Upstash's free tier is 500,000 commands/month, which is far smaller than it sounds:
500,000 / month = 16,129 / day = 11 / minute = 0.19 / second, sustained
@@ -14,7 +23,7 @@ Redis is only ever on the request path here, and why this script exists: the rul
to state and invisible to violate. A `cache.get()` added inside the reconciler's
per-instance loop is one line in review and 2,160,000 commands/month in production.
**Why a projection and not an alarm on the counter.** Exhausting the budget is a slow,
**This projects the burn rather than alarming on the counter.** Exhausting the budget is a slow,
silent failure with a cliff at the end: nothing degrades, nothing pages, every call
succeeds, and then the month rolls over and every Redis call starts erroring at once. By
then the fix is a bill or an outage. A burn rate extrapolated from the counter is visible
@@ -113,7 +122,23 @@ def collect(text: str) -> tuple[dict[str, float], float]:
return per_op, started_at
def report(per_op: dict[str, float], started_at: float, budget: int, now: float) -> int:
def merge(scrapes: list[tuple[dict[str, float], float]]) -> tuple[dict[str, float], float]:
"""Fold several processes' scrapes into one budget view.
The budget is per-Upstash-database, but the counters are per-process: api, worker and
reconciler each hold their own prometheus_client registry, so scraping one of them
projects a third of the truth. Command totals sum across processes; the window is the
EARLIEST start time, because a counter that has been running longest bounds how far back
the summed total can be attributed — using the latest would inflate the rate.
"""
per_op: dict[str, float] = {}
for scraped, _ in scrapes:
for op, value in scraped.items():
per_op[op] = per_op.get(op, 0.0) + value
return per_op, min(started for _, started in scrapes)
def report(per_op: dict[str, float], started_at: float, budget: int, now: float, sources: int = 1) -> int:
"""Print the projection. Returns the process exit code."""
elapsed_s = max(1.0, now - started_at)
total = sum(per_op.values())
@@ -128,6 +153,17 @@ def report(per_op: dict[str, float], started_at: float, budget: int, now: float)
print(f"rate {rate:.4f} /s (budget allows {budget / _MONTH_S:.4f} /s sustained)")
print(f"projected {projected:,.0f} / month")
print(f"budget {budget:,} / month")
print(f"sources {sources} process(es) scraped")
if sources < 3:
# Be honest about what was measured. api/worker/reconciler each keep their own
# in-process registry (obs.py: one process per pod, no multiproc dir), so a
# single-endpoint run undercounts the shared Upstash budget by however many
# processes were left out. An optimistic verdict here is worse than no verdict.
print(
" UNDERCOUNT: svcforge runs api + worker + reconciler, each with "
"its own\n registry. Pass --url once per process for the real "
"total; the numbers\n above cover only what was scraped."
)
if total < 100:
# Extrapolating a month from a handful of commands is astrology. Say so rather than
@@ -147,20 +183,32 @@ def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(
description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter
)
parser.add_argument("--url", default="http://localhost:8000/metrics", help="Prometheus endpoint")
parser.add_argument(
"--url",
action="append",
dest="urls",
metavar="URL",
help="Prometheus endpoint; repeat once per process (api, worker, reconciler)",
)
parser.add_argument("--budget", type=int, default=_FREE_TIER_BUDGET, help="commands per month")
args = parser.parse_args(argv)
try:
per_op, started_at = collect(scrape(args.url))
except BudgetError as exc:
print(f"error: {exc}", file=sys.stderr)
return 2
except OSError as exc:
print(f"error: cannot scrape {args.url}: {exc}", file=sys.stderr)
return 2
# action="append" cannot carry a default (argparse appends to it), so apply it here.
urls: list[str] = args.urls or ["http://localhost:8000/metrics"]
return report(per_op, started_at, args.budget, time.time())
scrapes: list[tuple[dict[str, float], float]] = []
for url in urls:
try:
scrapes.append(collect(scrape(url)))
except BudgetError as exc:
print(f"error: {url}: {exc}", file=sys.stderr)
return 2
except OSError as exc:
print(f"error: cannot scrape {url}: {exc}", file=sys.stderr)
return 2
per_op, started_at = merge(scrapes)
return report(per_op, started_at, args.budget, time.time(), sources=len(scrapes))
if __name__ == "__main__":
+8 -1
View File
@@ -28,6 +28,12 @@ RUN --mount=type=cache,target=/root/.cache/uv \
COPY libs/ libs/
COPY services/api/ services/api/
COPY catalog.yaml ./
# The migrate Job runs from THIS image (migrate-job.yaml pins the api digest), so the SQL
# has to be in it. svcforge_core.migrate's fallback resolves MIGRATIONS_DIR relative to
# its own __file__, which lands under site-packages here — a directory that does not and
# should not contain SQL — so main() returned 1 and the pre-install/pre-upgrade hook
# failed every sync. SVCFORGE_MIGRATIONS_DIR below points it at this copy instead.
COPY migrations/ migrations/
# --no-editable is what turns svcforge-core into a real wheel in site-packages.
# pyproject.toml declares it `editable = true` for local dev; an editable install in an
# image points at /app/libs, which is a source tree that need not survive the final stage.
@@ -50,6 +56,7 @@ WORKDIR /app
COPY --from=builder --chown=10001:10001 /app /app
ENV PATH="/app/.venv/bin:$PATH" \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1
PYTHONDONTWRITEBYTECODE=1 \
SVCFORGE_MIGRATIONS_DIR=/app/migrations
USER 10001
ENTRYPOINT ["python", "-m", "services.api"]
+28 -5
View File
@@ -15,6 +15,7 @@ from fastapi import Depends, HTTPException, Request, status
from fastapi.security import HTTPAuthorizationCredentials, HTTPBearer
from jwt import PyJWKClient
from svcforge_core.adapters.redis import RateLimiterProto
from svcforge_core.domain.models import CatalogEntry
from svcforge_core.repo.db import DictPool
from svcforge_core.repo.instances import InstanceRepo
@@ -146,13 +147,35 @@ async def _signing_key(client: PyJWKClient, token: str) -> jwt.PyJWK:
return await asyncio.to_thread(client.get_signing_key_from_jwt, token)
async def rate_limit(team: Annotated[str, Depends(get_current_team)]) -> None:
"""Per-team rate limiting. Seam only — Module 10 fills this in (Redis, Lua, token bucket).
def get_rate_limiter(request: Request) -> RateLimiterProto | None:
"""The limiter lifespan built, or None when no Redis is configured."""
limiter: RateLimiterProto | None = getattr(request.app.state, "rate_limiter", None)
return limiter
It exists now, wired into the routes, so that turning it on is an edit to one function
body rather than a change to every handler signature.
async def rate_limit(
request: Request,
team: Annotated[str, Depends(get_current_team)],
) -> None:
"""Per-team rate limiting. One Redis command per check, and it fails OPEN.
Failing open is the entire policy. Redis holds derived state; losing it must degrade
the platform, never stop it. A limiter that fails closed converts a cache outage into
a total outage, which is a strictly worse incident than the burst it was protecting
against — so `RateLimiter.check` swallows its own errors and returns `allowed=True`.
The 429 below therefore only ever comes from a real, counted overage.
"""
return None
limiter = get_rate_limiter(request)
if limiter is None:
return
result = await limiter.check(team)
if not result.allowed:
raise HTTPException(
status_code=status.HTTP_429_TOO_MANY_REQUESTS,
detail={"code": "rate_limited", "message": "too many requests"},
headers={"Retry-After": str(result.retry_after_s)},
)
async def idempotency_key(request: Request) -> str | None:
+31 -6
View File
@@ -8,7 +8,6 @@ environment at import time — before any fixture can say otherwise.
from __future__ import annotations
import asyncio
import logging
from collections.abc import AsyncIterator
from contextlib import asynccontextmanager
@@ -18,11 +17,13 @@ from jwt import PyJWKClient
from services.api.models import ErrorBody
from services.api.routes import health, instances
from svcforge_core import obs
from svcforge_core.adapters.redis import RateLimiter, make_redis
from svcforge_core.domain.catalog import load_catalog
from svcforge_core.repo.db import make_pool
from svcforge_core.settings import Settings, load_settings
log = logging.getLogger(__name__)
log = obs.get_logger("svcforge.api")
@asynccontextmanager
@@ -42,6 +43,16 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
app.state.catalog = load_catalog(settings.catalog_path)
# Redis is optional by construction. `make_redis` returns None when no DSN is set, and
# every consumer treats None as "skip" — so a deployment without Redis loses rate
# limiting and keeps everything else. Built here rather than per request because a
# connection pool per request is a connection pool per request.
redis = make_redis(settings)
app.state.redis = redis
app.state.rate_limiter = (
RateLimiter(redis, limit=settings.rate_limit_per_minute, window_s=60) if redis is not None else None
)
pool = make_pool(str(settings.pg_dsn), settings.pool_min_size, settings.pool_max_size)
# wait=True fails NOW, loudly, if the DSN is wrong — instead of at the first request,
# as a PoolTimeout, in front of a user.
@@ -68,6 +79,8 @@ async def lifespan(app: FastAPI) -> AsyncIterator[None]:
yield
finally:
await pool.close()
if redis is not None:
await redis.aclose()
async def _http_exception_handler(request: Request, exc: Exception) -> JSONResponse:
@@ -93,6 +106,18 @@ def create_app(settings: Settings | None = None) -> FastAPI:
"""App factory: lifespan, routers, exception handler, /metrics."""
settings = settings or load_settings()
# FIRST, before any router is built and before any logger is bound. Without this the
# API is the one service of three that never configures structlog: its lines go out
# through logging.lastResort as bare text on stderr with no service, no trace_id and
# no JSON envelope — a parse failure in the collector, and unattributable in Loki.
# `settings.log_json` was silently inert here for the same reason.
obs.setup("svcforge-api", settings)
# Refuse the dev escape hatches when SVCFORGE_ENVIRONMENT says this is not a laptop.
# Called unconditionally and early: a check that only runs from a branch someone
# remembered to write is a check that does not run.
settings.check_production()
app = FastAPI(
title="svcforge",
version="0.1.0",
@@ -115,7 +140,7 @@ def app() -> FastAPI:
return create_app()
if __name__ == "__main__": # pragma: no cover
import uvicorn
uvicorn.run("services.api.main:app", factory=True, host="0.0.0.0", port=8000) # noqa: S104
# There is deliberately no `if __name__ == "__main__"` here. `services/api/__main__.py` is
# the single entrypoint, and the image's ENTRYPOINT uses it. A second one in this module
# drifted from it — different log_level, different access_log — so `python -m services.api`
# and `python services/api/main.py` started the same app two different ways.
+2 -2
View File
@@ -1,7 +1,7 @@
"""Liveness, readiness, metrics.
The distinction between the first two is not pedantry, it is the difference between a
30-second blip and a fleet-wide outage:
The distinction between the first two is the difference between a 30-second blip and a
fleet-wide outage:
* `/healthz` (liveness) answers "is this process wedged?" A failure here gets the
container KILLED. It must therefore touch NOTHING external. Wire it to the DB and a
+4 -1
View File
@@ -41,7 +41,10 @@ RUN useradd -u 10001 -m -s /usr/sbin/nologin svcforge
WORKDIR /app
COPY --from=builder --chown=10001:10001 /app /app
COPY --from=alpine/helm:3.16.2@sha256:a19a2968fd672336d39771f6c899781424d725229148656dbc2a1e305003cdec /usr/bin/helm /usr/local/bin/helm
# helm 3.21.3, not 3.16.2 — see services/worker/Dockerfile. 3.16.2 is a Go 1.22.9 build
# carrying CRITICAL CVE-2025-68121 (crypto/tls) and CVE-2026-33186 (grpc) and HIGH
# CVE-2026-35469 (spdystream). Kept on 3.x on purpose: helm 4 is a breaking change.
COPY --from=alpine/helm:3.21.3@sha256:35da09ba0716fc7c3cd63b6b31ee380a9c7662e95f29ab0e4ae962420afd315b /usr/bin/helm /usr/local/bin/helm
ENV PATH="/app/.venv/bin:$PATH" \
PYTHONUNBUFFERED=1 \
+7 -7
View File
@@ -81,8 +81,8 @@ class ReconcilerDeps:
clock: Clock
catalog: dict[str, CatalogEntry]
settings: Settings
# Whose instances go first in the day-2 work list. Eating your own dog food is an
# `order by`, not a policy document — see InstanceRepo.list_upgradable.
# Whose instances go first in the day-2 work list. Eating your own dog food is enforced
# by an `order by` rather than a policy document — see InstanceRepo.list_upgradable.
own_team: str
# A config value, not a scheduler. Leave it at 1 until 1 is too slow.
max_in_flight: int
@@ -152,15 +152,15 @@ async def check_drift(deps: ReconcilerDeps) -> None:
async def check_lease_expiry(deps: ReconcilerDeps) -> None:
"""Tasks whose worker died -> back to `queued`.
A lease, not a lock. No lock survives a power cut: a worker SIGKILLed mid-provision
A lease. No lock survives a power cut: a worker SIGKILLed mid-provision
leaves `state='running'` with `locked_by` set and nobody running it, and no amount of
cleanup code in the worker helps, because the worker is the part that died. `locked_at`
plus a timeout is the only thing that recovers the row, which is why `locked_at` exists.
The 5-minute default is not arbitrary: it must exceed the longest a healthy task can
hold a lease, or the reconciler hands a still-running provision to a second worker.
Handlers are idempotent, so that is survivable rather than fatal — but survivable is
not free, and `lease_seconds` sits above helm's `--timeout` for that reason.
The 5-minute default must exceed the longest a healthy task can hold a lease, or the
reconciler hands a still-running provision to a second worker. Handlers are idempotent,
so that is survivable, though it still costs a duplicated helm run — which is why
`lease_seconds` sits above helm's `--timeout`.
"""
freed = await deps.tasks.reset_expired_leases(deps.settings.lease_seconds)
if freed:
+15 -2
View File
@@ -41,8 +41,21 @@ RUN useradd -u 10001 -m -s /usr/sbin/nologin svcforge
WORKDIR /app
COPY --from=builder --chown=10001:10001 /app /app
COPY --from=alpine/helm:3.16.2@sha256:a19a2968fd672336d39771f6c899781424d725229148656dbc2a1e305003cdec /usr/bin/helm /usr/local/bin/helm
COPY --from=bitnamilegacy/kubectl:1.31.2@sha256:0eab9ec8f5e0f75271277467ebb7513b36dea0122bc68615d18c47fece4fd82c /opt/bitnami/kubectl/bin/kubectl /usr/local/bin/kubectl
# helm 3.21.3, not 3.16.2. 3.16.2 was built with Go 1.22.9 and carries CRITICAL
# CVE-2025-68121 (crypto/tls) and CVE-2026-33186 (grpc), plus HIGH CVE-2026-35469
# (spdystream, fixed in 0.5.1) — trivy fails the build on them and is right to.
# Deliberately 3.x: helm 4 is a breaking change and is not a CVE fix.
COPY --from=alpine/helm:3.21.3@sha256:35da09ba0716fc7c3cd63b6b31ee380a9c7662e95f29ab0e4ae962420afd315b /usr/bin/helm /usr/local/bin/helm
# kubectl 1.35.3 from the OFFICIAL registry.k8s.io image, replacing
# bitnamilegacy/kubectl:1.31.2. Two reasons, either sufficient:
# - CVEs: the bitnami image is Go 1.22.9 and ships containerd < 1.7.29 (HIGH
# CVE-2024-25621) along with the same crypto/tls and grpc CRITICALs as helm above.
# - Skew: the target cluster runs v1.35.3. 1.31.2 is four minors behind, well outside
# kubectl's supported +/-1 window, so it was unsupported against this cluster.
# bitnamilegacy publishes no 1.35 tag. registry.k8s.io/kubectl is the upstream-maintained
# image, is a manifest list with linux/arm64 (this cluster is Ampere), and puts the binary
# at /bin/kubectl rather than bitnami's /opt/bitnami path.
COPY --from=registry.k8s.io/kubectl:v1.35.3@sha256:8dad99b604a2c0bafe17f53cadf78482d6f667a6da687f385508f5f4e4696d37 /bin/kubectl /usr/local/bin/kubectl
ENV PATH="/app/.venv/bin:$PATH" \
PYTHONUNBUFFERED=1 \
+19
View File
@@ -8,6 +8,7 @@ run in milliseconds against a FakeProvisioner instead of needing a cluster.
from __future__ import annotations
from dataclasses import dataclass
from typing import Protocol
from svcforge_core.adapters.clock import Clock
from svcforge_core.adapters.helm import Provisioner
@@ -19,6 +20,20 @@ from svcforge_core.repo.tasks import TaskRepo
from svcforge_core.settings import Settings
class NamespaceEnsurer(Protocol):
"""The one thing the worker needs from kubectl.
Narrower than `KubectlClient` on purpose: the handler creates namespaces and does
nothing else with the cluster, so that is the whole interface. Satisfied by
`KubectlClient` in production and by a fake in tests, which is what keeps the worker
suite running in milliseconds with no cluster.
"""
async def ensure_namespace(self, ns: str, labels: dict[str, str] | None = None) -> None:
"""Create the namespace if absent. Idempotent — safe on every retry."""
...
@dataclass(frozen=True)
class WorkerDeps:
"""Everything a handler is allowed to touch."""
@@ -27,6 +42,10 @@ class WorkerDeps:
instances: InstanceRepo
tasks: TaskRepo
provisioner: Provisioner
# Creates the tenant namespace before helm is pointed at it. `helm --namespace X`
# does NOT create X, so without this the very first provision for a new team fails
# with "namespace not found" — the one path that is guaranteed untested by a fake.
namespaces: NamespaceEnsurer
notifier: Notifier
clock: Clock
catalog: dict[str, CatalogEntry]
+34 -8
View File
@@ -2,7 +2,7 @@
Every handler here obeys one rule: running it twice must equal running it once.
That is not a nicety. A worker can be SIGKILLed after helm has installed the release but
A worker can be SIGKILLed after helm has installed the release but
before the DB row says so; the lease expires; another worker claims the same task and runs
this function again. If the handler is not idempotent, the tenant gets two Elasticsearches
and you get a bill. Idempotency is what makes the crash safe, and it is bought in two
@@ -63,6 +63,15 @@ async def handle_provision(task: Task, deps: WorkerDeps) -> None:
# below is idempotent either way, so this is bookkeeping, not a lock.
await deps.instances.update_state(inst.id, InstanceState.REQUESTED, InstanceState.PROVISIONING)
# `helm --namespace X` does not create X. Every tenant's first provision lands in a
# namespace that does not exist yet, so this has to happen before helm runs or the
# install fails with "namespaces not found". Idempotent (kubectl apply of a Namespace
# manifest), so it costs one no-op API call on every subsequent provision.
await deps.namespaces.ensure_namespace(
inst.namespace,
labels={"svcforge.io/team": inst.team},
)
await deps.provisioner.install(
release=inst.release_name,
ns=inst.namespace,
@@ -92,7 +101,15 @@ async def handle_deprovision(task: Task, deps: WorkerDeps) -> None:
# `helm uninstall` of an already-gone release is not an error to us: the adapter
# swallows not-found, because the desired state — no release — is already true.
await deps.provisioner.uninstall(release=inst.release_name, ns=inst.namespace)
await deps.instances.update_state(inst.id, InstanceState.DELETING, InstanceState.DELETED)
# Raise rather than ignore the CAS result. Swallowing it means: the release is gone,
# the row keeps `state=ready` and its now-dangling endpoint, the task is marked done,
# and 60 seconds later the reconciler's drift check re-provisions the thing the tenant
# asked to delete. Failing loudly turns a silent ping-pong into one visible error.
if not await deps.instances.update_state(inst.id, InstanceState.DELETING, InstanceState.DELETED):
raise HandlerError(
f"instance {inst.id} was {inst.state.value}, expected {InstanceState.DELETING.value}"
)
async def handle_upgrade(task: Task, deps: WorkerDeps) -> None:
@@ -138,18 +155,27 @@ async def handle_verify(task: Task, deps: WorkerDeps) -> None:
if inst.release_name in releases:
return
# `returning` + a `where` on the update half tells us whether THIS call was the one
# that halted the rollout. The halt itself is idempotent; the page is not. Without the
# distinction, a verify that fails its full retry budget sends five identical
# notifications for one incident, spread across the backoff curve.
async with deps.pool.connection() as conn, conn.cursor() as cur:
await cur.execute(
"""insert into catalog_versions (service_type, rollout_state)
values (%s, 'halted')
on conflict (service_type) do update set rollout_state = 'halted'""",
on conflict (service_type) do update set rollout_state = 'halted'
where catalog_versions.rollout_state <> 'halted'
returning service_type""",
(inst.service_type,),
)
await deps.notifier.send(
"rollout.halted",
f"rollout halted for {inst.service_type}: {inst.release_name} failed verify",
{"instance_id": str(inst.id), "team": inst.team, "service_type": inst.service_type},
)
newly_halted = await cur.fetchone() is not None
if newly_halted:
await deps.notifier.send(
"rollout.halted",
f"rollout halted for {inst.service_type}: {inst.release_name} failed verify",
{"instance_id": str(inst.id), "team": inst.team, "service_type": inst.service_type},
)
raise HandlerError(f"verify failed for {inst.release_name}; rollout halted")
+48 -7
View File
@@ -13,6 +13,7 @@ import asyncio
import contextlib
import signal
import time
from collections.abc import Awaitable
from opentelemetry import trace
@@ -21,9 +22,10 @@ from services.worker.handlers import HANDLERS
from svcforge_core import obs
from svcforge_core.adapters.clock import SystemClock
from svcforge_core.adapters.helm import HelmProvisioner
from svcforge_core.adapters.k8s import KubectlClient
from svcforge_core.adapters.notify import LogNotifier
from svcforge_core.domain.catalog import load_catalog
from svcforge_core.domain.models import Task
from svcforge_core.domain.models import Task, TaskKind
from svcforge_core.repo.db import make_pool
from svcforge_core.repo.instances import InstanceRepo
from svcforge_core.repo.tasks import TaskRepo
@@ -42,8 +44,28 @@ async def _sleep_or_stop(stop: asyncio.Event, seconds: float) -> None:
await asyncio.wait_for(stop.wait(), timeout=seconds)
async def _report(coro: Awaitable[bool], task_id: int, what: str) -> None:
"""Run a terminal report, and never let its failure escape.
Reporting is the one thing that must not kill the worker. `_run_one` runs inside a
TaskGroup, and a TaskGroup cancels every sibling the moment one child raises — so a
DB blip during `tasks.fail()` would abort every other in-flight provision on this pod,
not just this one. The task itself is safe either way: it stays `running` and the
reconciler's lease sweep returns it to the queue. Losing the report costs one lease
interval; losing the siblings costs their work.
"""
try:
if not await coro:
# The lease was stolen while we were working: another worker owns this task
# now and is mid-run. Reporting is theirs to do, not ours.
log.warning("lease lost before report; another worker owns this task", task_id=task_id)
except Exception:
log.exception("could not report task %s (%s); lease will expire", task_id, what)
async def _run_one(deps: WorkerDeps, task: Task, sem: asyncio.Semaphore) -> None:
"""Run one task to a terminal report. Never lets an exception escape the TaskGroup."""
worker_id = deps.settings.worker_id
try:
# Every log line from here carries instance_id/task_id/team. Bound once, at claim,
# rather than passed down: the alternative is threading three arguments through
@@ -55,7 +77,11 @@ async def _run_one(deps: WorkerDeps, task: Task, sem: asyncio.Semaphore) -> None
handler = HANDLERS.get(task.kind)
if handler is None:
await deps.tasks.fail(task.id, f"no handler for {task.kind}", max_attempts=1)
await _report(
deps.tasks.fail(task.id, f"no handler for {task.kind}", worker_id, max_attempts=1),
task.id,
"no-handler",
)
return
# Re-parent to the span that enqueued this task. Without the stored traceparent
@@ -77,17 +103,30 @@ async def _run_one(deps: WorkerDeps, task: Task, sem: asyncio.Semaphore) -> None
# CancelledError inherits from BaseException, so `except Exception` below
# would never see it. Catch it only to release the claim, then ALWAYS
# re-raise: swallowing it breaks cancellation for everyone above us.
await deps.tasks.fail(task.id, "cancelled", max_attempts=deps.settings.max_attempts)
await _report(
deps.tasks.fail(task.id, "cancelled", worker_id, deps.settings.max_attempts),
task.id,
"cancelled",
)
raise
except Exception as exc:
log.exception("task failed", kind=task.kind.value, error=str(exc))
span.record_exception(exc)
span.set_status(trace.Status(trace.StatusCode.ERROR, str(exc)))
obs.TASKS_FAILED.labels(kind=task.kind.value).inc()
await deps.tasks.fail(task.id, str(exc), max_attempts=deps.settings.max_attempts)
obs.TASK_ATTEMPTS_FAILED.labels(kind=task.kind.value).inc()
await _report(
deps.tasks.fail(task.id, str(exc), worker_id, deps.settings.max_attempts),
task.id,
"fail",
)
else:
obs.PROVISION_TIME.observe(time.monotonic() - started)
await deps.tasks.complete(task.id)
# Only provisions go in the provision histogram. The buckets run 10s..1800s
# because they were sized for helm installs; a sub-second `verify` dropped
# into the same series drags the p95 down and quietly stops
# SvcforgeProvisionSlow from ever firing.
if task.kind is TaskKind.PROVISION:
obs.PROVISION_TIME.observe(time.monotonic() - started)
await _report(deps.tasks.complete(task.id, worker_id), task.id, "complete")
finally:
sem.release()
@@ -135,6 +174,7 @@ async def _amain() -> None:
# the SvcforgeTaskFailed / SvcforgeProvisionSlow alerts query do not exist until the
# registry is up.
obs.setup("svcforge-worker", settings)
settings.check_production()
obs.start_metrics_server(settings.metrics_port)
pool = make_pool(settings.pg_dsn.unicode_string(), settings.pool_min_size, settings.pool_max_size)
@@ -145,6 +185,7 @@ async def _amain() -> None:
instances=InstanceRepo(pool),
tasks=TaskRepo(pool),
provisioner=HelmProvisioner(helm_bin=settings.helm_bin, timeout_s=int(settings.helm_timeout_s)),
namespaces=KubectlClient(kubectl_bin=settings.kubectl_bin),
notifier=LogNotifier(),
clock=SystemClock(),
catalog=load_catalog(settings.catalog_path),
+1 -1
View File
@@ -98,7 +98,7 @@ def pg_dsn() -> Iterator[str]:
async def pool(pg_dsn: str) -> AsyncIterator[DictPool]:
"""A clean database and an open pool, per test.
max_size=60 is not a performance choice. test_skip_locked_claims_each_task_exactly_once
max_size=60 is a correctness requirement. test_skip_locked_claims_each_task_exactly_once
races 50 concurrent claims; a pool smaller than that serialises them at the pool
instead of at the database, and the test passes while proving nothing.
"""
+19 -5
View File
@@ -545,11 +545,25 @@ async def test_202_is_on_the_decorator_not_the_response_object(client: httpx.Asy
assert "202" in schema["paths"]["/v1/instances/{instance_id}"]["delete"]["responses"]
def test_auth_disabled_short_circuits(settings: Settings) -> None:
"""The dev escape hatch exists, and Settings.check_production() refuses it in prod."""
dev = settings.model_copy(update={"auth_disabled": True})
with pytest.raises(ValueError, match="refused outside local development"):
dev.check_production()
def test_auth_disabled_is_allowed_locally_and_refused_everywhere_else(settings: Settings) -> None:
"""The dev escape hatch exists, and check_production() refuses it outside `local`.
Both halves matter. The permissive half is why every entrypoint can call this
unconditionally at startup; the refusing half is the actual safety property. An
earlier version raised unconditionally, which meant it could only be called from a
branch that already knew it was production — so nobody ever wrote that branch and the
check never ran at all.
"""
dev_locally = settings.model_copy(update={"auth_disabled": True, "environment": "local"})
dev_locally.check_production() # must not raise
for env in ("prod", "staging"):
dev_in_prod = settings.model_copy(update={"auth_disabled": True, "environment": env})
with pytest.raises(ValueError, match="refused"):
dev_in_prod.check_production()
# And a correctly-configured production app starts fine.
settings.model_copy(update={"auth_disabled": False, "environment": "prod"}).check_production()
@pytest.mark.slow
+17 -5
View File
@@ -267,8 +267,8 @@ async def test_lease_expiry_returns_a_dead_workers_task_to_the_queue(pool: DictP
async def test_lease_expiry_leaves_a_live_worker_alone(pool: DictPool) -> None:
"""A task claimed a second ago is not a dead worker. Reclaiming it would double-provision.
Handlers are idempotent, so a wrongly-freed lease is survivable — but survivable is not
free, and this is why lease_seconds sits above helm's own --timeout.
Handlers are idempotent, so a wrongly-freed lease is survivable, though it still costs a
duplicated helm run — which is why lease_seconds sits above helm's own --timeout.
"""
instance_id, _, _ = await _seed(pool)
tasks = TaskRepo(pool)
@@ -419,8 +419,9 @@ async def test_version_drift_parks_the_upgrade_until_the_maintenance_window(
) -> None:
"""The queue does the waiting, in `where run_after <= now()`.
Not a scheduler and not an in-memory timer: a task parked in Postgres until 03:00 Sunday
survives a reconciler restart. That is the whole reason `run_after` exists.
The waiting is done by the queue rather than by a scheduler or an in-memory timer: a
task parked in Postgres until 03:00 Sunday survives a reconciler restart. That is the
whole reason `run_after` exists.
"""
instance_id, _, _ = await _seed(pool, maintenance_window=SUNDAY_0300_HCM)
@@ -520,8 +521,19 @@ def tracing() -> InMemorySpanExporter:
Global because OTEL's is: `trace.set_tracer_provider` takes once per process, and
`obs.tracer()` resolves it at call time. Session-scoped so the second call never
happens.
happens from HERE.
The internals reset is load-bearing rather than cosmetic. `set_tracer_provider` is
one-shot: a second call logs "Overriding of current TracerProvider is not allowed" at
WARNING and is otherwise ignored. Any earlier test that builds a FastAPI app calls
`obs.setup()` and burns that one shot, after which this fixture silently installs
nothing, `get_finished_spans()` returns `[]`, and the failure reads as "the reconciler
stopped writing traceparents" rather than "another test got there first". Clearing the
module globals is the only way to take the shot back.
"""
trace._TRACER_PROVIDER = None
trace._TRACER_PROVIDER_SET_ONCE._done = False
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
+1 -1
View File
@@ -206,7 +206,7 @@ async def test_real_lua_refuses_the_eleventh(upstash: Redis) -> None:
async def test_a_check_costs_exactly_one_redis_command(upstash: Redis) -> None:
"""One EVALSHA. Not GET+INCR+EXPIRE, which is three billed commands and a race.
At 500K/month the difference is not academic: three commands per request caps the
At 500K/month the difference is material: three commands per request caps the
platform at 166K requests/month instead of 500K, for a limiter that is also wrong.
The first call is excluded from the count on purpose. redis-py sends EVALSHA, Upstash
+62 -5
View File
@@ -35,7 +35,7 @@ async def test_complete_marks_done_and_releases_lock(pool: DictPool) -> None:
tid = await repo.enqueue_standalone(iid, TaskKind.PROVISION)
claimed = await repo.claim("w1")
assert claimed is not None
await repo.complete(claimed.id)
await repo.complete(claimed.id, "w1")
row = await _task_row(pool, tid)
assert row["state"] == "done"
assert row["locked_by"] is None
@@ -49,7 +49,7 @@ async def test_fail_under_max_attempts_requeues_with_future_run_after(pool: Dict
assert claimed is not None
assert claimed.attempts == 1 # attempts increments at CLAIM time, not on failure
await repo.fail(tid, "helm exploded", max_attempts=5)
await repo.fail(tid, "helm exploded", "w1", max_attempts=5)
row = await _task_row(pool, tid)
assert row["state"] == "queued"
assert row["last_error"] == "helm exploded"
@@ -66,10 +66,12 @@ async def test_fail_at_max_attempts_dead_letters_and_marks_instance(pool: DictPo
tasks, instances = TaskRepo(pool), InstanceRepo(pool)
tid = await tasks.enqueue_standalone(iid, TaskKind.PROVISION)
claimed = await tasks.claim("w1")
assert claimed is not None
async with pool.connection() as conn, conn.cursor() as cur:
await cur.execute("update tasks set attempts = 5 where id = %s", (tid,))
await tasks.fail(tid, "chart not found", max_attempts=5)
assert await tasks.fail(tid, "chart not found", "w1", max_attempts=5) is True
row = await _task_row(pool, tid)
assert row["state"] == "failed"
@@ -90,10 +92,12 @@ async def test_fail_does_not_resurrect_a_deleted_instance(pool: DictPool) -> Non
tasks, instances = TaskRepo(pool), InstanceRepo(pool)
tid = await tasks.enqueue_standalone(iid, TaskKind.DEPROVISION)
claimed = await tasks.claim("w1")
assert claimed is not None
async with pool.connection() as conn, conn.cursor() as cur:
await cur.execute("update tasks set attempts = 5 where id = %s", (tid,))
await tasks.fail(tid, "helm uninstall kept failing", max_attempts=5)
assert await tasks.fail(tid, "helm uninstall kept failing", "w1", max_attempts=5) is True
# The task still dead-letters — that part is unconditional.
row = await _task_row(pool, tid)
@@ -111,7 +115,7 @@ async def test_fail_truncates_error_to_2kb(pool: DictPool) -> None:
repo = TaskRepo(pool)
tid = await repo.enqueue_standalone(iid, TaskKind.PROVISION)
await repo.claim("w1")
await repo.fail(tid, "x" * 9000, max_attempts=5)
await repo.fail(tid, "x" * 9000, "w1", max_attempts=5)
row = await _task_row(pool, tid)
assert isinstance(row["last_error"], str)
assert len(row["last_error"]) == 2000
@@ -153,3 +157,56 @@ async def test_fresh_lease_is_not_reset(pool: DictPool) -> None:
await repo.enqueue_standalone(iid, TaskKind.PROVISION)
await repo.claim("w1")
assert await repo.reset_expired_leases(lease_seconds=300) == 0
async def test_stale_worker_cannot_complete_a_task_another_worker_now_owns(pool: DictPool) -> None:
"""The lost-lease race, as a regression test.
Worker A hangs past its lease. The reconciler requeues the task. Worker B claims it and
starts working. Worker A finally returns and reports success. Without the ownership
check in `complete`, A marks the task done while B is still running it: B's helm
install is unaccounted for, and if B then fails, the task is resurrected and a THIRD
worker provisions the same instance. That is the double-provision the claim query's
whole design exists to prevent, arriving through the back door.
"""
iid = await make_instance(pool)
repo = TaskRepo(pool)
tid = await repo.enqueue_standalone(iid, TaskKind.PROVISION)
a = await repo.claim("worker-A")
assert a is not None
# A's lease expires and the reconciler hands the task back to the queue.
async with pool.connection() as conn, conn.cursor() as cur:
await cur.execute("update tasks set locked_at = now() - interval '10 minutes' where id=%s", (tid,))
assert await repo.reset_expired_leases(lease_seconds=300) == 1
b = await repo.claim("worker-B")
assert b is not None, "worker-B should have been able to claim the requeued task"
# A reports. It must lose.
assert await repo.complete(tid, "worker-A") is False, "stale worker completed a task it no longer owns"
row = await _task_row(pool, tid)
assert row["state"] == "running", "the task must still belong to worker-B"
assert row["locked_by"] == "worker-B"
async def test_stale_worker_cannot_fail_a_task_another_worker_now_owns(pool: DictPool) -> None:
"""The mirror case: a stale failure report must not requeue someone else's task."""
iid = await make_instance(pool)
repo = TaskRepo(pool)
tid = await repo.enqueue_standalone(iid, TaskKind.PROVISION)
assert await repo.claim("worker-A") is not None
async with pool.connection() as conn, conn.cursor() as cur:
await cur.execute("update tasks set locked_at = now() - interval '10 minutes' where id=%s", (tid,))
await repo.reset_expired_leases(lease_seconds=300)
assert await repo.claim("worker-B") is not None
assert await repo.fail(tid, "stale report", "worker-A", max_attempts=5) is False
row = await _task_row(pool, tid)
assert row["state"] == "running"
assert row["locked_by"] == "worker-B"
assert row["last_error"] is None, "a stale worker wrote its error onto another worker's task"
+20
View File
@@ -41,10 +41,21 @@ def _settings(**over: object) -> Settings:
return Settings(**base) # type: ignore[arg-type]
class FakeNamespaceEnsurer:
"""Records the namespaces it was asked to create. Satisfies NamespaceEnsurer."""
def __init__(self) -> None:
self.ensured: list[str] = []
async def ensure_namespace(self, ns: str, labels: dict[str, str] | None = None) -> None:
self.ensured.append(ns)
def _deps(
pool: DictPool,
prov: FakeProvisioner,
notifier: FakeNotifier | None = None,
namespaces: FakeNamespaceEnsurer | None = None,
**over: object,
) -> WorkerDeps:
return WorkerDeps(
@@ -52,6 +63,7 @@ def _deps(
instances=InstanceRepo(pool),
tasks=TaskRepo(pool),
provisioner=prov,
namespaces=namespaces or FakeNamespaceEnsurer(),
notifier=notifier or FakeNotifier(),
clock=SystemClock(),
catalog=CATALOG,
@@ -200,4 +212,12 @@ async def test_concurrency_cap_is_respected(pool: DictPool, concurrency: int) ->
stop.set()
await asyncio.wait_for(worker, timeout=10)
# BOTH bounds. `<= concurrency` alone passes on a worker whose semaphore is broken to 1,
# or that lost concurrency entirely — it only proves the cap is not exceeded, never that
# concurrency exists. With 6 tasks seeded and a 0.2s handler, a working worker reaches
# its cap exactly.
assert prov.max_concurrent <= concurrency, f"ran {prov.max_concurrent} at once, cap was {concurrency}"
assert prov.max_concurrent == concurrency, (
f"only reached {prov.max_concurrent} concurrent with a cap of {concurrency}: "
"the worker is not actually running tasks in parallel"
)
+31 -16
View File
@@ -13,11 +13,13 @@ import json
import logging
import re
from collections.abc import Iterator
from pathlib import Path
from typing import Any
from uuid import uuid4
import pytest
import structlog
import yaml
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from prometheus_client import REGISTRY
@@ -89,8 +91,8 @@ def test_provision_histogram_has_a_bucket_for_a_thirty_minute_provision() -> Non
With the library defaults the top finite bucket is 10s, every provision lands in +Inf,
and histogram_quantile interpolates across a bucket spanning 10s to infinity. The p95
it returns is not slow or fast, it is meaningless and it is meaningless silently,
which is why this is asserted rather than eyeballed on a dashboard.
it returns is meaningless, and it is meaningless silently, which is why this is
asserted rather than eyeballed on a dashboard.
"""
bounds = obs.PROVISION_TIME._upper_bounds
assert 1800.0 in bounds
@@ -105,21 +107,34 @@ def test_provision_histogram_exports_the_1800_bucket() -> None:
assert bucket is not None
def test_metric_names_are_the_ones_the_alerts_query() -> None:
"""The four alerts are PromQL strings in values.yaml; nothing type-checks them.
def test_every_metric_the_alerts_query_actually_exists() -> None:
"""Parse the alerts out of values.yaml and check each metric is really registered.
A rename here is a silent alert that never fires again. This test is the link between
the chart's rules and the code.
Reading the chart is the whole point. An earlier version of this test asserted against
a hardcoded list, which meant it could only catch a rename on the *code* side rename
a metric in values.yaml and it stayed green while the alert silently never fired again.
A test that duplicates the thing it is meant to cross-check is not a cross-check.
"""
for name in (
"svcforge_tasks_claimed_total",
"svcforge_tasks_failed_total",
"svcforge_provision_duration_seconds",
"svcforge_queue_depth",
"svcforge_instances",
"svcforge_reconciler_last_tick_timestamp_seconds",
):
assert REGISTRY._names_to_collectors.get(name) is not None, f"{name} is queried by an alert"
values = yaml.safe_load(
(Path(__file__).resolve().parents[2] / "deploy" / "chart" / "values.yaml").read_text()
)
rules = (values.get("prometheusRule") or {}).get("rules") or []
assert rules, "no alert rules found under prometheusRule.rules — this test would silently pass"
# Prometheus appends _bucket/_sum/_count to a histogram's series; the collector is
# registered under the base name.
suffixes = ("_bucket", "_sum", "_count")
registered = set(REGISTRY._names_to_collectors)
checked = 0
for rule in rules:
for token in re.findall(r"\bsvcforge_[a-z0-9_]+\b", str(rule.get("expr", ""))):
base = next((token[: -len(s)] for s in suffixes if token.endswith(s)), token)
assert base in registered, (
f"alert {rule.get('alert')!r} queries {token!r}, but no such metric is registered"
)
checked += 1
assert checked >= len(rules), "expected at least one metric reference per alert"
# --- Trace context across the queue -------------------------------------------------------
@@ -232,7 +247,7 @@ def test_json_renderer_is_last_and_output_is_one_object_per_line(logs: io.String
def test_setup_is_idempotent(logs: io.StringIO) -> None:
"""Called twice, one handler, one line. Not a nicety: two handlers is two of every line.
"""Called twice, one handler, one line. Two handlers would be two of every line.
Each service calls setup() from its entrypoint, and an entrypoint that imports another
entrypoint (the CLI shelling into the reconciler) calls it twice.
+2 -2
View File
@@ -79,8 +79,8 @@ def take_cache(c: InstanceCacheProto) -> None: ...
def test_redis_implementations_conform() -> None:
"""A client at a closed port is still a Redis. Nothing here connects — no commands billed.
That is not a trick to keep the test fast; it is the module's thesis restated as a
fixture. Every one of these classes has to be constructible and callable with Redis
That is the module's thesis restated as a fixture, rather than a trick to keep the
test fast. Every one of these classes has to be constructible and callable with Redis
unreachable, because that is the state they are designed for.
"""
from datetime import UTC, datetime