svcforge: reference implementation
ci / lint (push) Successful in 1m19s
ci / unit (push) Failing after 1m2s
ci / integration (push) Has been skipped
ci / types (push) Successful in 1m37s
ci / security (push) Failing after 38s
ci / dockerfile (push) Successful in 14s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
ci / lint (push) Successful in 1m19s
ci / unit (push) Failing after 1m2s
ci / integration (push) Has been skipped
ci / types (push) Successful in 1m37s
ci / security (push) Failing after 38s
ci / dockerfile (push) Successful in 14s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
Complete working build of the system learn-python/ teaches. 164 tests, mypy --strict clean, domain coverage 99%.
This commit is contained in:
+192
@@ -0,0 +1,192 @@
|
||||
# svcforge runbook
|
||||
|
||||
Four entries. Each starts from an alert firing and ends at either a fix or an escalation.
|
||||
Every command is copy-pasteable; none of them require thinking at 3am, which is the point.
|
||||
|
||||
Set these first:
|
||||
|
||||
```bash
|
||||
set -a; . ~/.config/svcforge/secrets.env; set +a # SVCFORGE_PG_DSN_SESSION for psql
|
||||
alias sfsql='psql "$SVCFORGE_PG_DSN_SESSION"'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Measured numbers
|
||||
|
||||
From `scripts/load.py` + in-process workers on a `FakeProvisioner` (`delay=0.05s`), 200 tasks
|
||||
per run. **These came off local Postgres on the same box, with a sub-millisecond round trip.
|
||||
Supabase's pooler is ~5ms away, so treat these as a ceiling the real thing will not reach** —
|
||||
the shape is what transfers, not the absolute numbers.
|
||||
|
||||
| replicas | pool max_size | worker concurrency | connections used | drain (200 tasks) | throughput |
|
||||
|---:|---:|---:|---:|---:|---:|
|
||||
| 1 | 5 | 4 | 5 | 4.3s | 46.9 task/s |
|
||||
| 2 | 5 | 4 | 10 | 4.2s | 47.6 task/s |
|
||||
| 4 | 5 | 4 | 20 | 2.2s | 92.3 task/s |
|
||||
| 2 | 20 | 16 | 40 | 2.1s | 97.1 task/s |
|
||||
|
||||
Three things this says:
|
||||
|
||||
1. **Going from 1 replica to 2 bought nothing** (46.9 → 47.6). Throughput here is bounded by
|
||||
per-worker concurrency (the semaphore), not by replica count. Adding pods to a saturated
|
||||
semaphore is the most common wrong fix for a slow queue.
|
||||
2. **Concurrency is the knob that moved it** — 4 replicas (20 connections) and 2 replicas at
|
||||
concurrency 16 (40 connections) land in the same place, ~92-97 task/s. The second buys the
|
||||
same throughput for twice the connections, which on the free tier is the worse trade.
|
||||
3. **The budget is `(api + worker replicas) x max_size`**, and it is spent whether or not the
|
||||
connections are busy. The bottom row costs 40 connections for a 2% gain over the row above
|
||||
it. On Supabase free tier, that arithmetic — not throughput — is what decides replica count.
|
||||
|
||||
Nothing failed at any setting, so the real connection ceiling was never hit locally. Finding
|
||||
it against the actual pooler is the experiment worth running: raise `replicas x max_size`
|
||||
until claims slow and `PoolTimeout` appears, and write the number here.
|
||||
|
||||
---
|
||||
|
||||
## Queue stuck
|
||||
|
||||
**Alert:** `SvcforgeQueueDepthRising`
|
||||
|
||||
**Diagnose.** Start here, always:
|
||||
|
||||
```bash
|
||||
sfsql -c "select state, count(*) from tasks group by 1;"
|
||||
```
|
||||
|
||||
Then split the three causes apart — they look identical from the alert and need opposite fixes:
|
||||
|
||||
```bash
|
||||
# Stuck leases: rows 'running' with a locked_at that never advances.
|
||||
sfsql -c "select kind, locked_by, locked_at, last_error from tasks
|
||||
where state='running' order by locked_at limit 10;"
|
||||
|
||||
# No workers: is anything actually consuming?
|
||||
kubectl get pods -l app=worker -o wide
|
||||
kubectl logs -l app=worker --tail=20 --prefix
|
||||
|
||||
# run_after in the future: backoff has parked everything.
|
||||
sfsql -c "select count(*) from tasks where state='queued' and run_after > now();"
|
||||
```
|
||||
|
||||
| What you see | Cause | Fix |
|
||||
|---|---|---|
|
||||
| `running` rows, `locked_at` older than 5m, no worker pods hold those IDs | Workers died mid-task | None. The reconciler resets expired leases within 60s. If it does not, the reconciler is down — check it. |
|
||||
| Zero worker pods, or all `CrashLoopBackOff` | No consumer | Fix the workers. `kubectl describe pod -l app=worker`. |
|
||||
| Everything `queued` with `run_after` far in the future | Backoff, i.e. tasks are failing and retrying | This is not a queue problem. Go to **Provision failing**. |
|
||||
| `queued` rows with `run_after <= now()` and healthy workers | Real: claim is not returning rows | Check pooler connection budget (see **Supabase full**). |
|
||||
|
||||
**Never** hand-edit `state='running'` back to `'queued'`. The lease does that, and doing it
|
||||
by hand while the worker is actually alive gives you two workers on one task — the exact
|
||||
thing the whole design prevents.
|
||||
|
||||
**Escalate** if workers are healthy, leases are fresh, and depth still grows: that is a
|
||||
claim-query or pooler bug, not an ops problem.
|
||||
|
||||
---
|
||||
|
||||
## Provision failing
|
||||
|
||||
**Alert:** `SvcforgeTaskFailed` (tasks reaching the dead-letter state), or `SvcforgeProvisionSlow`
|
||||
(they still succeed, but the p95 has drifted out — usually cluster capacity, diagnosed the same way).
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
sfsql -c "select id, service_type, chart_version, error from instances where state='failed';"
|
||||
sfsql -c "select id, kind, attempts, last_error from tasks where state='failed' order by id desc limit 10;"
|
||||
|
||||
NS=tenant-<team>
|
||||
helm list -n "$NS"
|
||||
kubectl get events -n "$NS" --sort-by=.lastTimestamp | tail -20
|
||||
```
|
||||
|
||||
| `error` looks like | Cause | Fix |
|
||||
|---|---|---|
|
||||
| `chart "..." version "..." not found` | Bad pin in `catalog.yaml` | Correct the version, commit. The next `upgrade`/`provision` picks it up. |
|
||||
| `timed out waiting for the condition` | Cluster capacity — the chart installed but pods never became ready | `kubectl describe pod -n $NS`. Usually `Insufficient cpu/memory` or a PVC pending on Longhorn. |
|
||||
| `Error: ... forbidden: User "system:serviceaccount:svcforge:..."` | RBAC | The worker's ClusterRole is missing a verb. Chart change, not a manual `kubectl edit`. |
|
||||
| `ImagePullBackOff` in events | Registry auth or a gated image | Prefer `bitnamilegacy/*` images, which pull anonymously. |
|
||||
|
||||
After fixing the cause, tasks that already dead-lettered do **not** retry themselves. Requeue
|
||||
deliberately:
|
||||
|
||||
```bash
|
||||
sfsql -c "update tasks set state='queued', attempts=0, run_after=now(), last_error=null
|
||||
where id = <task_id>;"
|
||||
```
|
||||
|
||||
**Escalate** if `error` is empty on a failed instance — that means the failure path itself
|
||||
lost the message.
|
||||
|
||||
---
|
||||
|
||||
## Orphaned release
|
||||
|
||||
**Alert:** `SvcforgeReconcilerStale` — the reconciler has not completed a loop recently, so
|
||||
drift is no longer being *detected* at all. Drift itself is reported in the reconciler's logs
|
||||
and metrics rather than paged on, because it is usually benign and always needs a human to
|
||||
judge. A stale reconciler is the real emergency: nothing is watching.
|
||||
|
||||
The control loop **never auto-deletes a release.** That is deliberate: a bug in the drift
|
||||
check that deletes things is unrecoverable, and one that only reports is a Tuesday.
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
helm list -A -o json | jq -r '.[].name' | sort > /tmp/real
|
||||
sfsql -tAc "select release_name from instances where state in ('ready','provisioning');" | sort > /tmp/want
|
||||
|
||||
comm -23 /tmp/real /tmp/want # in the cluster, not in the DB -> orphan
|
||||
comm -13 /tmp/real /tmp/want # in the DB, not in the cluster -> missing
|
||||
```
|
||||
|
||||
| Direction | Meaning | Action |
|
||||
|---|---|---|
|
||||
| Orphan (cluster only) | A deprovision half-finished, or someone ran `helm install` by hand | Confirm the tenant is gone, then `helm uninstall <name> -n <ns>` **by hand**, and write down that you did. |
|
||||
| Missing (DB only) | Someone deleted a release out from under us | Requeue a `provision` task for that instance. It is idempotent; it will rebuild. |
|
||||
|
||||
**Escalate** before uninstalling anything you did not personally trace to a deleted instance.
|
||||
A wrong `helm uninstall` here deletes a tenant's data.
|
||||
|
||||
---
|
||||
|
||||
## Supabase full
|
||||
|
||||
**Alert:** none — and that is a gap, not a decision. The free tier is 0.5 GB and nothing pages
|
||||
you before you hit it; you find out when writes start failing. Until someone adds a size rule,
|
||||
this entry is driven by the calendar, not by an alert. Check it monthly:
|
||||
|
||||
```bash
|
||||
sfsql -c "select pg_size_pretty(pg_database_size(current_database()));"
|
||||
```
|
||||
|
||||
**Diagnose:**
|
||||
|
||||
```bash
|
||||
sfsql -c "select pg_size_pretty(pg_database_size(current_database()));"
|
||||
sfsql -c "select relname, pg_size_pretty(pg_total_relation_size(relid)) from pg_catalog.pg_statio_user_tables
|
||||
order by pg_total_relation_size(relid) desc limit 5;"
|
||||
sfsql -c "select count(*) from pg_stat_activity;"
|
||||
```
|
||||
|
||||
It is almost always `tasks`. Every provision, upgrade and verify leaves a row forever.
|
||||
|
||||
```bash
|
||||
sfsql -c "delete from tasks where state='done' and created_at < now() - interval '7 days';"
|
||||
sfsql -c "vacuum (analyze) tasks;"
|
||||
```
|
||||
|
||||
`vacuum` alone reclaims space **for reuse by Postgres**, but does not return it to the
|
||||
filesystem — so `pg_database_size` may barely move. That is expected and fine; the space is
|
||||
free for new rows. `vacuum full` does return it, takes an `ACCESS EXCLUSIVE` lock, and will
|
||||
stall every worker for its duration. Only do it in a window, and only if you actually need
|
||||
the bytes back.
|
||||
|
||||
If `count(*) from pg_stat_activity` is near the pooler's ceiling, the cause is arithmetic, not
|
||||
load: `worker_replicas × pool_max_size + api_replicas × pool_max_size`. Lower `max_size` or
|
||||
replicas. **Replica count is a database-capacity decision here**, which is unusual and worth
|
||||
remembering.
|
||||
|
||||
**Escalate** if size is growing with `tasks` already pruned — that means `instances` is
|
||||
growing, i.e. tenants are real, i.e. the free tier is the wrong tier.
|
||||
Reference in New Issue
Block a user