"""Liveness, readiness, metrics. The distinction between the first two is not pedantry, it is the difference between a 30-second blip and a fleet-wide outage: * `/healthz` (liveness) answers "is this process wedged?" A failure here gets the container KILLED. It must therefore touch NOTHING external. Wire it to the DB and a 20-second Postgres failover restarts every pod at once; they come back, find the DB still down, and CrashLoopBackOff with exponential restart delays — so the fleet is now down for minutes after the database recovered. * `/readyz` (readiness) answers "should this pod get traffic?" A failure here only removes it from the Service endpoints. It is allowed to check dependencies, and it recovers by itself the moment the check passes. """ from __future__ import annotations from typing import Any from fastapi import APIRouter, HTTPException, Request, Response, status from prometheus_client import REGISTRY from prometheus_client.exposition import choose_encoder from services.api.deps import PoolDep from services.api.models import ErrorBody router = APIRouter(tags=["ops"]) # No PROMETHEUS_MULTIPROC_DIR here, deliberately: it exists for prefork servers where each # worker process holds a slice of the counters. One uvicorn process per container means # the default in-process registry is already correct, and multiproc mode would add a # shared temp dir, a cleanup obligation, and a class of stale-file bugs for nothing. @router.get("/healthz", status_code=status.HTTP_200_OK) async def healthz() -> dict[str, str]: """Liveness. No I/O. If the event loop can run this, the process is alive.""" return {"status": "ok"} @router.get( "/readyz", responses={503: {"model": ErrorBody, "description": "A dependency is unavailable"}}, ) async def readyz(pool: PoolDep) -> dict[str, str]: """Readiness. Postgres only. Postgres-only is the rule, and Redis is the temptation. Redis holds derived state — rate-limit buckets, caches — and everything degrades gracefully without it. Put it in this check and an Upstash hiccup marks every pod unready, Kubernetes empties the Service, and a cache outage becomes a total API outage. """ try: async with pool.connection() as conn, conn.cursor() as cur: await cur.execute("select 1") row: Any = await cur.fetchone() if row is None: raise RuntimeError("select 1 returned no row") except Exception as exc: # closed pool, timeout, dead DB — all mean the same 'not ready' raise HTTPException( status_code=status.HTTP_503_SERVICE_UNAVAILABLE, detail={"code": "not_ready", "message": "database unavailable"}, ) from exc return {"status": "ready"} @router.get("/metrics", response_class=Response) async def metrics(request: Request) -> Response: """The Prometheus scrape endpoint. A route rather than `app.mount("/metrics", make_asgi_app())`, for two reasons. A Starlette `Mount` compiles to `^/metrics(?P/.*)$`, which does not match a bare `/metrics` — the exact URL every scrape config uses — and a `Mount` is invisible to OpenAPI, while the deliverable asks for `/metrics` in `openapi.json`. The encoding is still prometheus_client's: `choose_encoder` reads the Accept header and picks the exposition format (Prometheus text vs OpenMetrics) with its matching content type. Hand-rolling either is how you end up serving text/plain that a scraper rejects. """ encoder, content_type = choose_encoder(request.headers.get("Accept", "")) return Response(content=encoder(REGISTRY), media_type=content_type)