Files
svcforge/deploy/chart/values.yaml
T
Nguyen Minh Phuc c76154aeaa
ci / lint (push) Successful in 34s
ci / unit (push) Successful in 1m41s
ci / types (push) Successful in 1m41s
ci / dockerfile (push) Successful in 18s
ci / security (push) Successful in 1m27s
ci / chart (push) Failing after 1m11s
ci / integration (push) Successful in 1m10s
ci / image (api) (push) Has been skipped
ci / image (reconciler) (push) Has been skipped
ci / image (worker) (push) Has been skipped
ci / bump (push) Has been skipped
review: fix 26 findings from a 4-agent audit
CORRECTNESS
- lost-lease race: complete()/fail() did not check ownership, so a worker whose
  lease expired could mark a task done while another worker was running it, or
  requeue a task someone else owned. Reproduced, fixed with a CAS on
  (state, locked_by), pinned by two regression tests.
- worker died on report failure: _run_one's docstring claimed no exception
  escapes the TaskGroup; fail()/complete() were outside the guarded block, so a
  DB blip cancelled every sibling provision on the pod.
- claim query used an INNER join, which could strand a just-claimed task and
  report 'queue empty'. LEFT join.
- InstanceRepo.set_error bypassed the state machine and had no callers. Deleted.
- handle_deprovision ignored its CAS result, so a wrong-state instance kept a
  dangling endpoint and got re-provisioned by the drift check 60s later.
- handle_verify re-notified on every retry: five pages for one halt.

DEPLOY-BREAKING
- the migration Job could never succeed: no Dockerfile copied migrations/, and
  migrate.py resolved the path relative to the source tree, which only works for
  an editable install. Added COPY + SVCFORGE_MIGRATIONS_DIR.
- ServiceMonitor selector did not match the Service: API metrics never scraped.
- SvcforgeReconcilerStale fired permanently from every pod, because the gauge is
  module-level and every service exports it as 0. Scoped to the reconciler job.
- SvcforgeTaskFailed latched forever on a monotonic counter. Now increase()[15m].
- the digest guard accepted the all-zeros placeholder.
- worker terminationGracePeriodSeconds was 60s against a 600s helm timeout.

DEAD CODE THAT SHOULD NOT HAVE BEEN
- adapters/k8s.py was never called, so tenant namespaces were never created and
  the first provision for a new team would fail. Wired into handle_provision.
- adapters/redis.py was never imported by any service. Rate limiting is now wired
  into the API, failing open.
- Settings.check_production() had no callers. Given an explicit environment and
  called from every entrypoint.

OBSERVABILITY
- the API never called obs.setup(): no JSON logs, no trace correlation, log_json
  silently inert.
- LogNotifier's structured fields were discarded by the stdlib->structlog bridge.
- bind_task_context cleared the 'service' binding for the life of every task.
- split tasks_failed into task_attempts_failed and tasks_dead_lettered.

SECURITY
- trivy correctly blocked the worker/reconciler images: helm 3.16.2 and kubectl
  1.31.2 carry CRITICAL Go stdlib CVEs. Bumped to helm 3.21.3 and kubectl 1.35.3,
  which also closes a four-minor skew against the v1.35.3 cluster.

TESTS THAT COULD NOT FAIL
- the concurrency cap test passed on a fully serial worker.
- the alert/metric cross-check asserted a hardcoded list instead of reading the
  chart, so it could not catch a rename on the chart side.
- fixed OTel tracer-provider pollution between test files.

DOCS
- ARCHITECTURE.md: mermaid diagrams, user stories, and the helm-vs-ArgoCD
  guarantee (verified with --dry-run=server).
- AGENTS.md + CLAUDE.md.
- prose sweep for back-and-forth phrasing across 19 files.
2026-07-18 12:13:49 +00:00

189 lines
7.0 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# svcforge chart values.
#
# The digests below are the deployment. CI builds each image once, pushes it, reads the
# digest back with `docker buildx imagetools inspect`, and `yq -i`s it into this file as
# its last act. ArgoCD notices the commit and syncs. Nothing else deploys svcforge.
#
# Tags are banned. A tag is a mutable pointer, which means "what is running" and "what
# this file says" can silently diverge. A digest cannot.
nameOverride: ""
fullnameOverride: ""
image:
registry: gitea.oci-oci.duckdns.org
pullPolicy: IfNotPresent
pullSecrets: []
# One repo + one digest per service: CI's matrix builds three images, so there are three
# digests. The zeros are placeholders — a fresh clone must be bumped by CI before it can
# deploy, which is the intended failure mode. Never hand-edit these.
api:
repo: gitea.oci-oci.duckdns.org/gitea_admin/svcforge-api
digest: sha256:0000000000000000000000000000000000000000000000000000000000000000
worker:
repo: gitea.oci-oci.duckdns.org/gitea_admin/svcforge-worker
digest: sha256:0000000000000000000000000000000000000000000000000000000000000000
reconciler:
repo: gitea.oci-oci.duckdns.org/gitea_admin/svcforge-reconciler
digest: sha256:0000000000000000000000000000000000000000000000000000000000000000
api:
replicas: 2
# One process per pod. Module 7 took the "scale with replicas" fix over
# PROMETHEUS_MULTIPROC_DIR, so `uvicorn --workers N` here would corrupt the metrics.
resources:
requests: {cpu: 50m, memory: 128Mi}
limits: {memory: 256Mi}
service:
type: ClusterIP
port: 80
targetPort: 8000
ingress:
enabled: true
className: nginx
annotations:
cert-manager.io/cluster-issuer: letsencrypt-prod
host: svcforge.oci-oci.duckdns.org
tls:
enabled: true
secretName: svcforge-tls
worker:
# Plain replicas. No HPA: the day `replicas: 2` stops keeping up, not before.
replicas: 2
concurrency: 4
resources:
requests: {cpu: 100m, memory: 192Mi}
limits: {memory: 512Mi}
reconciler:
# A singleton, and not by convention — the four checks are not safe to run twice
# concurrently. replicas is deliberately not a value: there is nothing to tune.
intervalSeconds: 60
resources:
requests: {cpu: 50m, memory: 128Mi}
limits: {memory: 256Mi}
# Postgres pool sizing. replicas × maxSize is spent against the Supabase pooler budget:
# api(2 × 5) + worker(2 × 5) + reconciler(1 × 2) = 22 connections. Raise with care.
pool:
minSize: 1
maxSize: 5
migrate:
# backoffLimit: 0 — a failed migration must fail the release, not retry into a
# half-applied schema. Migrations run here and only here; never on app startup.
enabled: true
resources:
requests: {cpu: 50m, memory: 128Mi}
limits: {memory: 256Mi}
auth:
jwksUrl: https://auth.oci-oci.duckdns.org/realms/svcforge/protocol/openid-connect/certs
issuer: https://auth.oci-oci.duckdns.org/realms/svcforge
audience: svcforge
otel:
enabled: true
endpoint: http://alloy.observability.svc.cluster.local:4317
log:
level: info
# The DSNs are pulled from Vault by external-secrets into a Secret the pods envFrom.
# No DSN is ever a chart value, a ConfigMap key, or a CI variable.
externalSecret:
enabled: true
secretStoreRef:
name: vault
kind: ClusterSecretStore
refreshInterval: 1h
# target Secret name; keys land as SVCFORGE_PG_DSN / SVCFORGE_PG_DSN_SESSION / SVCFORGE_REDIS_DSN
targetName: svcforge-secrets
remoteRefs:
- secretKey: SVCFORGE_PG_DSN
key: svcforge/postgres
property: dsn_pooler
- secretKey: SVCFORGE_PG_DSN_SESSION
key: svcforge/postgres
property: dsn_session
- secretKey: SVCFORGE_REDIS_DSN
key: svcforge/redis
property: dsn
rbac:
# The worker helm-installs tenant releases into namespaces it creates. `namespaces` is a
# cluster-scoped resource, so `create namespaces` cannot be granted by a namespaced Role
# — this has to be a ClusterRole. It is still least-privilege: named resources, named
# verbs, no `*`, no cluster-admin, and no rbac.authorization.k8s.io group at all, so the
# worker cannot grant itself anything further.
create: true
serviceAccount:
create: true
annotations: {}
# Chart-native only. A hand-authored ServiceMonitor/PrometheusRule CR is banned — the
# chart owns these, gated by these flags.
serviceMonitor:
enabled: true
interval: 30s
prometheusRule:
enabled: true
rules:
- alert: SvcforgeQueueDepthRising
expr: deriv(svcforge_queue_depth[10m]) > 0
for: 10m
labels:
severity: warning
annotations:
summary: svcforge queue depth is rising and not draining
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#queue-stuck
- alert: SvcforgeProvisionSlow
expr: histogram_quantile(0.95, sum by (le) (rate(svcforge_provision_duration_seconds_bucket[30m]))) > 300
for: 15m
labels:
severity: warning
annotations:
summary: svcforge p95 provision time is over 5 minutes
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#provision-failing
# `increase(...[15m])`, not the raw counter. `sum(counter) > 0` on a monotonic counter
# latches: one dead-lettered task at any point keeps this firing until the pod restarts,
# and a restart silently clears it — so it can never distinguish "failing now" from
# "failed last Tuesday". The dead-letter counter is also the right one: the attempts
# counter increments on ordinary transient retries that later succeed.
- alert: SvcforgeTaskDeadLettered
expr: sum(increase(svcforge_tasks_dead_lettered_total[15m])) > 0
for: 5m
labels:
severity: warning
annotations:
summary: a svcforge task exhausted its retries
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#provision-failing
# Scoped to the reconciler job, and aggregated with max().
#
# RECONCILER_LAST_TICK is a module-level Gauge in obs.py, so EVERY service that imports
# svcforge_core.obs registers and exports it — api and worker included, permanently at
# 0. Unscoped, `time() - 0` is ~1.7e9, so this critical alert fires from the moment the
# chart is installed, from pods that have no reconciler in them. Scope by job, then
# max() so a rolling restart of the single reconciler does not flap it.
- alert: SvcforgeReconcilerStale
expr: time() - max(svcforge_reconciler_last_tick_timestamp_seconds{job=~".*reconciler.*"}) > 300
for: 5m
labels:
severity: critical
annotations:
summary: the svcforge reconciler has not ticked in 5 minutes
runbook_url: https://gitea.oci-oci.duckdns.org/gitea_admin/svcforge/src/branch/master/RUNBOOK.md#orphaned-release
# Not in the spec. Off by default: on a three-node k3s a PDB that cannot be satisfied
# blocks drains, which is worse than the disruption it prevents.
podDisruptionBudget:
enabled: false
minAvailable: 1
nodeSelector: {}
tolerations: []
affinity: {}