Files
Nguyen Minh Phuc c53734d2bc
ci / lint (push) Successful in 33s
ci / types (push) Successful in 43s
ci / unit (push) Successful in 32s
ci / security (push) Successful in 57s
ci / dockerfile (push) Successful in 7s
ci / chart (push) Successful in 8s
ci / integration (push) Successful in 55s
ci / image (api) (push) Successful in 3m39s
ci / image (reconciler) (push) Successful in 2m53s
ci / image (worker) (push) Successful in 2m14s
ci / bump (push) Successful in 16s
docs: add USER_GUIDE.md, tighten comments, fix CLI needing a DSN
The comment pass is prose-only: every distinct "why" is kept, the
narration around it is not. Verified by AST-comparing each changed file
against HEAD with docstrings stripped — only the two files below differ
in executable code.

Two real fixes fell out of the read-through:

* The CLI documented itself as never touching the database, then called
  load_settings(), which requires SVCFORGE_PG_DSN. It refused to start
  without a Postgres URL it never opens. It now has its own two-field
  ClientSettings; the orphaned api_url/api_token are dropped from
  Settings, where nothing else read them.
* repo/db.py had the DictRow alias comment and the ERROR_MAX_CHARS
  comment run together above the wrong symbol.

USER_GUIDE.md is the caller-facing guide the README only gestured at:
auth, catalog, every endpoint with curl, the lifecycle, the error table,
rate limiting, the CLI, client generation, an end-to-end poll loop.

It records two facts about the live deployment rather than documenting a
flow nobody can run. SVCFORGE_JWKS_URL points at a realm with no IdP
behind it, so the API logs "JWKS warm-up failed" at startup and every
/v1 request is a 401. And `helm repo list` in the worker returns no
repositories, so the three bitnamilegacy/ catalog entries cannot resolve
at provision time; only the oci:// entries can.

make lint clean, 76 unit + 111 integration tests pass.
2026-07-21 14:51:26 +00:00

145 lines
7.7 KiB
Markdown

# svcforge — reference implementation
A complete, working, verified build of the system `../learn-python/` teaches you to build.
## Read this part first
**This folder can ruin the course, and it will if you let it.**
`learn-python/AGENTS.md` opens with a rule: *"Do not write his implementation code for him.
`domain/`, `repo/`, the claim loop, the handlers — those are the course."* This repo is
exactly that code. You asked for it deliberately, and the rule allows you to overrule it —
but the reason for the rule did not go away when you did.
Module 4 is five evenings of getting the claim query wrong, and those five evenings are the
learning. The evening you spend watching two workers grab the same task is the evening
`FOR UPDATE SKIP LOCKED` stops being a phrase and becomes something you understand.
Reading `repo/tasks.py` here takes ninety seconds, teaches you close to nothing, and feels
exactly like learning. That feeling is the trap.
So:
| Use it like this | Not like this |
|---|---|
| Attempt the module. Get stuck. Stay stuck 30 minutes. **Then** diff your version against this one. | Open this first "just to see the shape". |
| Steal the scaffolding — `pyproject.toml`, Dockerfiles, `ci.yaml`, the chart. Nothing is learned by fighting hatchling. | Copy `domain/`, `repo/`, `handlers.py`, or the claim loop. |
| Read the **comments**. They explain *why*, which is the part that transfers. | Read the code. It is the part that doesn't. |
| Use it to check an answer you already produced. | Use it to produce an answer. |
The scaffolding is where struggling teaches you nothing. The domain and the queue are the
whole point. Know which one you're reading.
## What is actually verified
Everything below was executed, not asserted:
| Claim | How it was proven |
|---|---|
| The claim query never double-claims | 50 concurrent workers, 50 tasks, real Postgres. Each claimed exactly once. |
| The transaction story is real | Instance + task roll back together on an abort. |
| SIGTERM drains in flight | Worker finishes a 2s provision after stop is set, exits 0. |
| Handlers are idempotent | Re-running `handle_provision` installs once, not twice. |
| helm timeouts kill the process **group** | A negative control with `proc.kill()` leaks two `sleep 300`s; the real `_run` leaves zero. |
| The state machine is enforced in SQL too | Guard test fails when the guard is removed (control-tested). |
| Redis degrades safely | Rate limit fails open, cache falls through, platform stays up with Redis dead. |
| The images ship a wheel, not an editable | Negative control proved `uv sync` alone ships `/app/libs`; `--no-editable` fixed it. |
| The chart is valid | `helm lint`, `helm template`, `hadolint` — all clean. |
Numbers from a real drain (400 tasks, local Postgres, FakeProvisioner) are in
[RUNBOOK.md](RUNBOOK.md#measured-numbers).
## Running it
**No Docker on this host, deliberately.** This box's `containerd` belongs to a Kubernetes
kubelet; installing `docker.io` would evict the runtime and take every pod with it. So the
integration tests take a DSN from the environment and only fall back to testcontainers
(the CI path) when it is absent:
```bash
sudo -u postgres psql -c "create role svcforge login password 'svcforge' superuser;"
sudo -u postgres psql -c "create database svcforge_test owner svcforge;"
export PATH="$HOME/.local/bin:$PATH"
export SVCFORGE_TEST_DSN="postgresql://svcforge:svcforge@127.0.0.1:5432/svcforge_test"
make dev # uv sync
make lint # ruff + ruff format + mypy --strict
uv run pytest -q -m "not e2e and not slow"
```
Each pytest process clones its own database, so concurrent runs don't truncate each other.
Redis tests marked `slow` hit real Upstash. **Mind the budget**: the free tier is 500K
commands/month = 0.19/sec sustained. The whole suite spends about 35.
```bash
set -a; . ~/.config/svcforge/secrets.env; set +a # never commit these
uv run pytest -q -m slow
uv run python -m scripts.redis_budget # projects month-end burn, exits 1 if over
```
## Using the API
**[USER_GUIDE.md](USER_GUIDE.md)** is the guide for callers: auth, the catalog, every
endpoint with curl, the lifecycle, the error table, the CLI.
The API also documents itself — FastAPI generates OpenAPI from the same models and routes
it serves, so the spec cannot drift the way a hand-written one does.
| What | Where |
|---|---|
| Swagger UI (try requests in the browser) | `https://svcforge.oci-oci.duckdns.org/docs` |
| ReDoc (nicer to read) | `https://svcforge.oci-oci.duckdns.org/redoc` |
| Raw spec, for generating clients | `https://svcforge.oci-oci.duckdns.org/openapi.json` |
Locally, `uv run uvicorn services.api.main:app --factory` then <http://127.0.0.1:8000/docs>.
`tests/integration/test_api.py` pins the description, the tags and the bearer security
scheme, so the docs fail CI if they rot.
## Where things live
| Module | Teaches | Read here |
|---|---|---|
| 0 | toolchain | `pyproject.toml`, `Makefile` |
| 1 | domain, pydantic, Protocol | `libs/svcforge_core/svcforge_core/domain/` |
| 2 | psycopg3, repos, migrations | `repo/{db,instances}.py`, `migrations/001_init.sql` |
| 3 | FastAPI, JWT, one transaction | `services/api/` |
| **4** | **the queue — the core** | **`repo/tasks.py`, `services/worker/`** |
| 5 | subprocess, timeouts, Protocols | `adapters/helm.py` (read `_run` twice) |
| 6 | day-2 fleet upgrades | `domain/windows.py`, `InstanceRepo.list_upgradable` |
| 7 | reconciler, obs | `services/reconciler/`, `obs.py` |
| 8 | packaging, CI, ArgoCD | `services/*/Dockerfile`, `.gitea/workflows/ci.yaml`, `deploy/` |
| 9 | chaos, load, CLI, runbook | `scripts/load.py`, `services/cli/`, `RUNBOOK.md` |
| 10 | Redis as a shortcut | `adapters/redis.py`, `scripts/redis_budget.py` |
## Where this deviates from the spec, and why
Honest list. Each was a real conflict, not a shortcut.
1. **`TaskRepo.enqueue` exists twice.** Module 2 specs `enqueue(conn, ...)`; Module 4 specs
`enqueue(instance_id, kind) -> int`. Both callers are real, so both exist:
`enqueue(conn, ...)` and `enqueue_standalone(...)`. Making `conn` optional would have
hidden the transaction question, which is the one thing that module is about.
2. **The claim query has a CTE wrapper.** The spec calls it byte-identical. The `update ...
for update skip locked` shape is untouched; an outer `select` joins `instances.team` so
the worker can bind `team` to its logs at claim time. The alternative was a second round
trip per task for a column the DB already had.
3. **The day-2 work-list query gained `service_type` and a halted check.** As printed it
compares every service type against one version, and never consults `catalog_versions`
despite the same module requiring halted → 0 rows.
4. **`Role` → `ClusterRole`.** The specced rules grant `create` on `namespaces`, which is
cluster-scoped. A namespaced Role cannot express it. Rules verbatim otherwise.
5. **DELETE is not one transaction.** `InstanceRepo.update_state` owns its connection, so
the CAS and the enqueue can't share one without reaching around the repo. Ordered for
the failure mode instead: CAS first, enqueue second — a crash between leaves `deleting`
with no task, which the reconciler sweeps up. The reverse would tear down a live service.
6. **Traceparent is `-03`, not the spec's `-01`.** Current SDKs set the random-trace-id bit
alongside sampled. The spec's acceptance line is stale.
7. **`make_redis` returns `Redis | None`.** "No Redis configured" has to be runnable.
## What was deliberately NOT built
Because Module 6 says so, and the restraint is the lesson: no `resize`, no `backup`/`restore`,
no `helm rollback` automation, no deprecation timers, no rollouts table, no pause/resume CLI.
A halted rollout is one column, cleared by hand with SQL.