4a426dbe50
ci / lint (push) Successful in 25s
ci / unit (push) Successful in 59s
ci / types (push) Successful in 1m8s
ci / dockerfile (push) Successful in 13s
ci / chart (push) Successful in 8s
ci / security (push) Successful in 1m2s
ci / integration (push) Successful in 54s
ci / image (api) (push) Successful in 1m6s
ci / image (reconciler) (push) Successful in 3m6s
ci / image (worker) (push) Successful in 2m28s
ci / bump (push) Has been cancelled
trivy took the worker image 39 -> 18 -> 5 findings across two version bumps, and the last 5 (4x golang.org/x/net, 1x Go stdlib) live in kubectl v1.36.2 — the newest kubectl that exists. No release clears them; upstream has not rebuilt against the patched Go yet. Chasing the version further has no end. kubectl was in that image for exactly one call: ensure_namespace before helm. `helm upgrade --install --create-namespace` does the same thing, idempotently, as part of the install it already runs. So the binary goes, and its vendored CVEs go with it. read_secret had no callers. Tradeoff recorded: the namespace no longer gets an svcforge.io/team label, since --create-namespace makes a bare one. Nothing reads that label today.
473 lines
19 KiB
Markdown
473 lines
19 KiB
Markdown
# svcforge — how it works
|
|
|
|
A control plane for internal service provisioning. A team asks for an Elasticsearch; a
|
|
worker installs one; a control loop keeps reality matching the database.
|
|
|
|
This document explains the design and **why each piece is the way it is**. Every "why" here
|
|
was paid for by a specific failure mode.
|
|
|
|
---
|
|
|
|
## The one-paragraph version
|
|
|
|
`POST /v1/instances` writes two rows in one transaction: the instance, and a task to build
|
|
it. It returns `202` immediately. A worker claims that task with a single `SELECT ... FOR
|
|
UPDATE SKIP LOCKED` statement, runs `helm upgrade --install`, and marks the instance ready.
|
|
A reconciler sweeps every 60 seconds for the things that go wrong when a process dies at
|
|
the wrong moment. The queue is a Postgres table, not Redis, and that single decision
|
|
determines most of the rest of the design.
|
|
|
|
---
|
|
|
|
## The shape
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
subgraph tenant["Tenant"]
|
|
CLI["svcforge CLI<br/><i>HTTP only, never the DB</i>"]
|
|
end
|
|
|
|
subgraph plane["svcforge control plane"]
|
|
API["<b>api</b><br/>FastAPI, N replicas<br/>validates, authenticates,<br/>enqueues"]
|
|
WORKER["<b>worker</b><br/>N replicas<br/>claims tasks, runs helm"]
|
|
RECON["<b>reconciler</b><br/><i>exactly 1</i><br/>4 checks every 60s"]
|
|
end
|
|
|
|
subgraph state["State"]
|
|
PG[("<b>Postgres</b><br/>instances + tasks<br/><i>the truth</i>")]
|
|
REDIS[("<b>Redis</b><br/>rate limit, cache<br/><i>derived only</i>")]
|
|
end
|
|
|
|
K8S["<b>Kubernetes</b><br/>helm releases<br/><i>the real world</i>"]
|
|
|
|
CLI -->|"POST /v1/instances"| API
|
|
API -->|"one transaction:<br/>instance + task"| PG
|
|
API -.->|"best effort"| REDIS
|
|
WORKER -->|"claim<br/>SKIP LOCKED"| PG
|
|
WORKER -->|"helm upgrade --install"| K8S
|
|
RECON -->|"drift, leases,<br/>TTL, versions"| PG
|
|
RECON -->|"helm list"| K8S
|
|
|
|
classDef truth fill:#2d4a22,stroke:#5a8f3d,color:#fff
|
|
classDef derived fill:#4a3222,stroke:#8f6a3d,color:#fff
|
|
class PG truth
|
|
class REDIS derived
|
|
```
|
|
|
|
Three services, four layers, two stores. The layer rule is one line:
|
|
|
|
```
|
|
transport/ HTTP handlers, CLI entrypoints. Knows FastAPI. Knows nothing about SQL.
|
|
domain/ Pure logic: state machine, catalog, backoff, windows. No I/O. No async.
|
|
repo/ SQL. Rows in, domain objects out. Knows psycopg. Knows nothing about HTTP.
|
|
adapters/ The outside world: helm, kubectl, webhooks, Redis, the clock.
|
|
```
|
|
|
|
Dependencies point inward: `transport → domain ← repo/adapters`. `domain/` imports nothing
|
|
from the other three, which is why its tests need no mocks and run in milliseconds.
|
|
|
|
---
|
|
|
|
## The decision everything else follows from
|
|
|
|
**The queue is a Postgres table.**
|
|
|
|
A task and the instance state it describes must commit atomically. Split them across two
|
|
stores and you own a distributed commit problem with no winning move: the process can die
|
|
between the two writes, and whichever you write first is the one that lies.
|
|
|
|
- Task first, then instance → an orphan task pointing at an instance that never existed.
|
|
- Instance first, then task → an instance nobody will ever build.
|
|
|
|
In one table, in one transaction, neither is possible:
|
|
|
|
```sql
|
|
BEGIN;
|
|
INSERT INTO instances (...); -- state='requested'
|
|
INSERT INTO tasks (...); -- kind='provision'
|
|
COMMIT; -- both, or neither
|
|
```
|
|
|
|
Redis cannot do this, so Redis is not the queue. It holds derived state only — things that
|
|
can be recomputed and whose loss is an inconvenience, never a corruption.
|
|
|
|
| | Postgres | Redis |
|
|
|---|---|---|
|
|
| Holds | instances, tasks | rate limits, idempotency keys, cache |
|
|
| If it is down | the platform is down | the platform is fine |
|
|
| If it disagrees with reality | reality wins, reconciler fixes it | discard it |
|
|
| Lives on the | control loop **and** request path | request path only |
|
|
|
|
That last row is a budget constraint. Upstash's free tier is 500K
|
|
commands/month = **0.19 commands/second sustained**. One worker polling Redis every 5
|
|
seconds is 518,400/month — the entire budget, spent by one pod doing nothing.
|
|
|
|
---
|
|
|
|
## Provisioning, end to end
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
autonumber
|
|
participant T as Tenant
|
|
participant A as api
|
|
participant P as Postgres
|
|
participant W as worker
|
|
participant K as Kubernetes
|
|
|
|
T->>A: POST /v1/instances {elasticsearch, small}
|
|
A->>A: verify JWT, look up catalog
|
|
rect rgb(45, 74, 34)
|
|
A->>P: BEGIN
|
|
A->>P: INSERT instance (state=requested)
|
|
A->>P: INSERT task (kind=provision, traceparent)
|
|
A->>P: COMMIT
|
|
end
|
|
A-->>T: 202 Accepted + Location
|
|
|
|
Note over W,P: every 5s, while a semaphore slot is free
|
|
W->>P: UPDATE ... FOR UPDATE SKIP LOCKED
|
|
P-->>W: task (attempts now 1, locked_by=me)
|
|
W->>K: helm upgrade --install --create-namespace --wait
|
|
K-->>W: release ready
|
|
W->>P: CAS provisioning -> ready, set endpoint
|
|
W->>P: complete(task, worker_id)
|
|
|
|
T->>A: GET /v1/instances/{id}
|
|
A-->>T: {state: ready, endpoint: ...}
|
|
```
|
|
|
|
**Why 202 and not 201.** A provision is `helm --wait` on a StatefulSet: minutes. Holding
|
|
an HTTP connection open for that is a request that dies to any proxy timeout, and a client
|
|
that cannot tell "still working" from "lost". The queue absorbs the API's output, which is
|
|
the property that lets 50 simultaneous POSTs all return instantly.
|
|
|
|
---
|
|
|
|
## The claim query
|
|
|
|
This is the heart of the system. It is one statement, and it must stay one statement.
|
|
|
|
```sql
|
|
WITH claimed AS (
|
|
UPDATE tasks SET state='running', attempts=attempts+1,
|
|
locked_by=%(worker)s, locked_at=now()
|
|
WHERE id = (
|
|
SELECT id FROM tasks
|
|
WHERE state='queued' AND run_after <= now()
|
|
ORDER BY run_after
|
|
FOR UPDATE SKIP LOCKED -- step over rows other workers hold, do not queue behind them
|
|
LIMIT 1
|
|
)
|
|
RETURNING *
|
|
)
|
|
SELECT claimed.*, instances.team
|
|
FROM claimed LEFT JOIN instances ON instances.id = claimed.instance_id;
|
|
```
|
|
|
|
- **`FOR UPDATE SKIP LOCKED`** is what makes N workers scale. Without `SKIP LOCKED` they
|
|
queue single-file behind whoever holds the oldest row.
|
|
- **The subquery exists because Postgres has no `UPDATE ... LIMIT`.**
|
|
- **Select-then-update as two statements is the bug this prevents.** Between the `SELECT`
|
|
and the `UPDATE`, a second worker reads the same id and both provision. The window is
|
|
small, which means you will not hit it in testing and you will hit it in production.
|
|
- **`LEFT` join, not inner.** The `UPDATE` has already taken effect when the outer select
|
|
runs. An inner join matching nothing would return no row, so `claim()` would report
|
|
"queue empty" for a task it had just marked `running` — stranding it until the lease
|
|
expires, having silently burned an attempt.
|
|
- **`attempts` increments at claim time, not on failure.** A worker that dies without
|
|
reporting has still burned an attempt, so a task that reliably kills its worker cannot
|
|
retry forever.
|
|
|
|
---
|
|
|
|
## The instance lifecycle
|
|
|
|
```mermaid
|
|
stateDiagram-v2
|
|
[*] --> requested: POST /v1/instances
|
|
requested --> provisioning: worker claims
|
|
provisioning --> ready: helm --wait succeeded
|
|
ready --> deleting: DELETE, or TTL expired
|
|
deleting --> deleted: helm uninstall succeeded
|
|
|
|
requested --> failed: attempts exhausted
|
|
provisioning --> failed: attempts exhausted
|
|
ready --> failed: drift — the release vanished
|
|
failed --> provisioning: retry
|
|
failed --> deleting: give up, tear it down
|
|
deleting --> failed: attempts exhausted
|
|
|
|
deleted --> [*]: terminal
|
|
```
|
|
|
|
`LEGAL` is a `dict[InstanceState, frozenset[InstanceState]]` in `domain/states.py`, not a
|
|
chain of `if`s. `deleted` maps to an **empty frozenset** rather than being absent, so
|
|
"terminal" is stated rather than implied by a missing key.
|
|
|
|
**The state machine is enforced in SQL too.** `TaskRepo.fail` writes `instances.state`
|
|
directly, so it derives its guard from the same `LEGAL` table:
|
|
|
|
```python
|
|
_CAN_FAIL = tuple(s.value for s, allowed in LEGAL.items() if InstanceState.FAILED in allowed)
|
|
...
|
|
UPDATE instances SET state='failed' WHERE id=%s AND state = ANY(%s)
|
|
```
|
|
|
|
Without that, a deprovision exhausting its retries against an already-`deleted` instance
|
|
would resurrect it into `failed` — a transition `transition()` explicitly forbids,
|
|
performed by raw SQL that never asked it. A state machine only one layer respects is
|
|
decoration.
|
|
|
|
---
|
|
|
|
## What happens when things die
|
|
|
|
This is the part that matters. Every guarantee below has a test.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
START["worker claims task<br/>state=running, locked_by=me"] --> WORK["helm upgrade --install"]
|
|
WORK -->|success| REPORT["complete(task, worker_id)"]
|
|
WORK -->|"raises"| FAIL["fail(task, err, worker_id)"]
|
|
WORK -->|"pod SIGKILLed"| DEAD["nothing reported<br/>row stuck at 'running'"]
|
|
|
|
REPORT --> CAS{"still locked_by me?"}
|
|
CAS -->|yes| DONE["state=done"]
|
|
CAS -->|"no — lease was stolen"| DROP["log and drop.<br/>the new owner reports"]
|
|
|
|
FAIL --> ATT{"attempts < max?"}
|
|
ATT -->|yes| REQUEUE["state=queued<br/>run_after += backoff+jitter"]
|
|
ATT -->|no| DEADLETTER["state=failed<br/>error copied to instance"]
|
|
|
|
DEAD --> LEASE["reconciler: locked_at older<br/>than lease_seconds"]
|
|
LEASE --> REQUEUE
|
|
|
|
REQUEUE --> START
|
|
|
|
classDef bad fill:#4a2222,stroke:#8f3d3d,color:#fff
|
|
classDef good fill:#2d4a22,stroke:#5a8f3d,color:#fff
|
|
class DEAD,DEADLETTER bad
|
|
class DONE,DROP good
|
|
```
|
|
|
|
**Recovery is by lease.** No distributed lock survives a power cut. A SIGKILLed worker leaves
|
|
`state='running'` with `locked_by` set and nobody running it; that row would sit there
|
|
forever. The lease is the only thing that recovers it, which is why `locked_at` exists.
|
|
|
|
**Ownership is checked on report.** A worker that hangs past its lease has its task
|
|
requeued and re-claimed by someone else. When it finally returns, `complete()` and `fail()`
|
|
both check `AND state='running' AND locked_by=%s`. Without that check, the stale worker
|
|
marks the task done while the new owner is still running it — and if the new owner then
|
|
fails, a *third* worker provisions the same instance. That is the double-provision the
|
|
claim query exists to prevent, arriving through the back door.
|
|
|
|
**Idempotency is what makes all of this safe.** It is bought in two places:
|
|
|
|
1. A deterministic release name: `f"{team}-{service_type}-{id[:8]}"`, `UNIQUE` in the schema.
|
|
2. Adapters that state desired state — `helm upgrade --install`, `kubectl apply` — instead
|
|
of issuing imperative commands.
|
|
|
|
Running a handler twice equals running it once, so redelivery is boring.
|
|
|
|
---
|
|
|
|
## The reconciler
|
|
|
|
One replica. Four checks. Every 60 seconds.
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
TICK(("tick<br/>every 60s")) --> D["<b>drift</b><br/>helm list vs DB"]
|
|
TICK --> L["<b>lease expiry</b><br/>running + locked_at old"]
|
|
TICK --> T["<b>TTL</b><br/>ready + expires_at passed"]
|
|
TICK --> V["<b>version drift</b><br/>chart_version ≠ catalog"]
|
|
|
|
D --> D1["release missing → re-enqueue provision"]
|
|
D --> D2["release unknown → <b>log only</b>"]
|
|
L --> L1["→ queued, locked_by=null"]
|
|
T --> T1["→ deleting + deprovision task"]
|
|
V --> V1["→ upgrade task, inside<br/>the maintenance window"]
|
|
```
|
|
|
|
**The drift check never auto-deletes.** A bug in a delete path is unrecoverable; a bug in a
|
|
report path is a Tuesday. Unknown releases are logged for a human.
|
|
|
|
**It is a singleton because it sweeps.** Two reconcilers double-enqueue and race on TTL.
|
|
`replicas: 1` and `strategy: Recreate` in the chart, plus per-check idempotency guards that
|
|
re-verify under `FOR UPDATE`.
|
|
|
|
---
|
|
|
|
## Day 2: upgrading the fleet
|
|
|
|
The team that runs this also patches it. That is the whole module, and it is **two columns
|
|
and one query**.
|
|
|
|
```mermaid
|
|
flowchart TD
|
|
EDIT["edit catalog.yaml<br/>21.3.15 → 21.3.16"] --> Q
|
|
|
|
Q["<b>the work list</b><br/>state=ready AND chart_version ≠ catalog<br/>AND NOT halted<br/>ORDER BY own_team DESC<br/>LIMIT max_in_flight"]
|
|
|
|
Q --> U["upgrade task<br/>run_after = next maintenance window"]
|
|
U --> H["helm upgrade --install"]
|
|
H --> W["write instances.chart_version<br/><i>only after helm succeeds</i>"]
|
|
W --> VER["verify task"]
|
|
VER -->|healthy| NEXT["next instance"]
|
|
VER -->|"failed"| HALT["catalog_versions.rollout_state='halted'<br/><i>work list goes empty</i>"]
|
|
|
|
NEXT --> Q
|
|
HALT --> HUMAN["cleared by hand, with SQL"]
|
|
|
|
classDef stop fill:#4a2222,stroke:#8f3d3d,color:#fff
|
|
class HALT,HUMAN stop
|
|
```
|
|
|
|
- **`instances.chart_version` is written only after helm succeeds.** Write it optimistically
|
|
and the fleet looks upgraded while it is not.
|
|
- **`ORDER BY team = own_team DESC`** puts your own instances first, so you are the tenant
|
|
who discovers the chart is broken. Eating your own dog food is enforced by an `ORDER BY`
|
|
rather than left to policy.
|
|
- **`max_in_flight` starts at 1** — a config value, not a scheduler. One at a time is what
|
|
makes the halt meaningful: the fleet stops after the first casualty, not after all of them.
|
|
- **The halt is one column, cleared by hand.** An automatic un-halt would just resume
|
|
breaking things.
|
|
|
|
Deliberately **not** built: `resize`, `backup`/`restore`, `helm rollback` automation,
|
|
deprecation timers, a rollouts table with history, a pause/resume CLI. Backup is a whole
|
|
subsystem and an untested restore is a rumour.
|
|
|
|
---
|
|
|
|
## Observability: one trace across the queue
|
|
|
|
```mermaid
|
|
sequenceDiagram
|
|
participant A as api
|
|
participant P as tasks table
|
|
participant W as worker
|
|
|
|
A->>A: span "POST /v1/instances" (trace abc123)
|
|
A->>P: INSERT ... traceparent='00-abc123-...'
|
|
Note over P: minutes pass. different pod.
|
|
W->>P: claim → row carries traceparent
|
|
W->>W: span "task.provision", parent=abc123
|
|
Note over A,W: one trace: POST → queue → helm
|
|
```
|
|
|
|
Trace context does **not** survive a queue on its own. The worker picks the row up in
|
|
another process with no ambient context. So the traceparent rides in the table. Skip this
|
|
and Tempo shows two unrelated traces for one provision, which is worse than no tracing
|
|
because it looks like it works.
|
|
|
|
Logs carry `instance_id` / `task_id` / `team` on every line via contextvars, bound once at
|
|
claim. Metrics are deliberately few, and two of them are named to prevent a specific
|
|
mistake: `svcforge_task_attempts_failed_total` counts *attempts* that raised, while
|
|
`svcforge_tasks_dead_lettered_total` counts tasks that gave up. Alerting on the first
|
|
pages you for ordinary retries that later succeed.
|
|
|
|
---
|
|
|
|
## Delivery: CI never touches the cluster
|
|
|
|
```mermaid
|
|
flowchart LR
|
|
PUSH["push to master"] --> GATES
|
|
|
|
subgraph GATES["gates — all required, none advisory"]
|
|
direction TB
|
|
G1["ruff"] --> G2["mypy --strict"] --> G3["pytest unit + coverage"]
|
|
G3 --> G4["migrate + pytest integration"]
|
|
G4 --> G5["bandit / gitleaks / pip-audit"]
|
|
G5 --> G6["hadolint + helm lint/template"]
|
|
end
|
|
|
|
GATES --> BUILD["build image (--load)"]
|
|
BUILD --> SCAN["trivy HIGH,CRITICAL"]
|
|
SCAN -->|clean| PUSHIMG["push by digest"]
|
|
SCAN -->|"CVE"| STOP["pipeline fails"]
|
|
PUSHIMG --> BUMP["commit digest to values.yaml<br/>[skip ci]"]
|
|
BUMP --> ARGO["ArgoCD notices the commit"]
|
|
ARGO --> CLUSTER["cluster"]
|
|
|
|
classDef stop fill:#4a2222,stroke:#8f3d3d,color:#fff
|
|
class STOP stop
|
|
```
|
|
|
|
**CI holds no kubeconfig, and must never hold one.** Its last act is a git commit; ArgoCD
|
|
pulls. The maximum blast radius of a compromised pipeline is a bad commit, which is
|
|
revertable.
|
|
|
|
**Build once, promote the artifact.** The image is loaded locally, scanned, and only then
|
|
pushed — push-then-scan means a CRITICAL sits in the registry behind a green checkmark.
|
|
Deploys are **by digest**, never a mutable tag.
|
|
|
|
**Migrations never run on app startup.** N replicas would race. They run as a Helm
|
|
`pre-upgrade,pre-install` hook Job, once, before any new pod serves traffic. Forward-only,
|
|
expand/contract: a rename is three deploys.
|
|
|
|
### The chart works with or without ArgoCD
|
|
|
|
The chart is **pure Helm**. It contains no `argocd.argoproj.io/*` annotations, no sync
|
|
waves, and no ArgoCD-specific ordering. Both paths are supported and both are verified:
|
|
|
|
```bash
|
|
# Path 1: plain Helm, no ArgoCD anywhere
|
|
helm upgrade --install svcforge deploy/chart -n svcforge \
|
|
--set image.api.digest=sha256:... --set image.worker.digest=sha256:... \
|
|
--set image.reconciler.digest=sha256:...
|
|
|
|
# Path 2: GitOps. ArgoCD watches master and applies the same chart.
|
|
kubectl apply -f deploy/argocd/app.yaml
|
|
```
|
|
|
|
Ordering survives both because **ArgoCD translates Helm hooks into its own sync phases**
|
|
rather than ignoring them:
|
|
|
|
| Annotation | Plain Helm | ArgoCD |
|
|
|---|---|---|
|
|
| `helm.sh/hook: pre-install,pre-upgrade` | runs before the release, aborts it on failure | mapped to the **PreSync** phase |
|
|
| `helm.sh/hook-weight: "-5"` | orders hooks within the phase | mapped to hook ordering |
|
|
| `helm.sh/hook-delete-policy: before-hook-creation` | deletes the previous Job first | mapped to `BeforeHookCreation` |
|
|
|
|
Using `argocd.argoproj.io/hook` instead would have been the trap: plain `helm install`
|
|
does not understand that annotation, so it would create the migration Job as an ordinary
|
|
resource with no ordering guarantee — the migration and the new pods would start together,
|
|
and the failure would appear only in whichever path nobody tested.
|
|
|
|
Verified with `helm install --dry-run=server` against a real cluster, which validates every
|
|
manifest through the API server rather than only rendering the templates locally.
|
|
|
|
---
|
|
|
|
## User stories, and what each one exercises
|
|
|
|
| As a… | I want… | So that… | Exercised by |
|
|
|---|---|---|---|
|
|
| tenant team | to request an Elasticsearch without filing a ticket | I am unblocked in minutes | `POST /v1/instances` → 202 |
|
|
| tenant team | to see why my instance failed | I can fix my own request | `instances.error`, `svcforge status` |
|
|
| tenant team | a throwaway instance to clean itself up | I do not pay for what I forgot | `ttl_days` → reconciler TTL sweep |
|
|
| platform team | a worker pod to be killable at any instant | a rolling deploy is not an outage | SIGTERM drain + lease recovery |
|
|
| platform team | to patch a CVE across every tenant | one edit, not N | `catalog.yaml` bump → work list |
|
|
| platform team | a bad chart to stop after the first casualty | I do not break 40 tenants | `verify` → `rollout_state='halted'` |
|
|
| platform team | to know the queue is stuck before a tenant tells me | I look competent | `SvcforgeQueueDepthRising` → RUNBOOK |
|
|
| on-call | a copy-pasteable diagnosis at 3am | I do not have to think | [RUNBOOK.md](RUNBOOK.md) |
|
|
|
|
---
|
|
|
|
## Where to read the code
|
|
|
|
| To understand | Read | Then |
|
|
|---|---|---|
|
|
| the data model | `domain/models.py`, `migrations/001_init.sql` | `domain/states.py` |
|
|
| **the queue** | `repo/tasks.py` — read `_CLAIM_SQL` twice | `services/worker/main.py` |
|
|
| crash safety | `services/worker/handlers.py` | `tests/integration/test_worker.py` |
|
|
| the API contract | `services/api/routes/instances.py` | `tests/integration/test_api.py` |
|
|
| subprocess discipline | `adapters/helm.py::_run` | `tests/integration/test_helm_timeout.py` |
|
|
| day 2 | `domain/windows.py`, `InstanceRepo.list_upgradable` | `services/reconciler/main.py` |
|
|
| how it ships | `.gitea/workflows/ci.yaml`, `deploy/chart/` | [RUNBOOK.md](RUNBOOK.md) |
|
|
|
|
Related: [README.md](README.md) for how to read this repo without spoiling the course,
|
|
[RUNBOOK.md](RUNBOOK.md) for operating it.
|