4193a18cae
ci / lint (push) Waiting to run
ci / types (push) Blocked by required conditions
ci / unit (push) Blocked by required conditions
ci / integration (push) Blocked by required conditions
ci / security (push) Blocked by required conditions
ci / dockerfile (push) Blocked by required conditions
ci / chart (push) Blocked by required conditions
ci / image (api) (push) Blocked by required conditions
ci / image (reconciler) (push) Blocked by required conditions
ci / image (worker) (push) Blocked by required conditions
ci / bump (push) Blocked by required conditions
helm's list gunzips every release payload to build its table. The only fields
either caller reads are name and namespace:
reconciler/main.py live = {(r.name, r.namespace) for r in releases}
worker/handlers.py releases = {r.name for r in await ...list_releases()}
Both live in the release secret's labels and metadata, so nothing needs
decompressing. Measured from a pod on this cluster: 21ms and 31KB, against
helm's 4392ms.
That is the whole fix for a drift check that was timing out at 330s every tick
and OOMKilling the container at 256Mi. The cause of both was decompressing 96
releases to extract two strings each. It also means the container no longer
cares that a cluster policy mutates its CPU request to 0 — at 21ms there is
nothing left to starve.
Three details carry the correctness, each with a test:
- PartialObjectMetadataList in the Accept header asks for metadata only.
Without it every release's gzipped manifest crosses the wire to be thrown
away, which is the cost this exists to avoid.
- The status selector drops superseded revisions server-side: 96 release
secrets here, 25 live. Failed and pending states are kept, matching what
`helm list` shows, because a failed release does exist and calling it
missing would have the reconciler re-provision on top of it.
- helm writes one secret per revision, up to 10 per release here, so the
newest version label wins. Control-tested with a string compare, which
picks "9" over "10".
Out of cluster there is no ServiceAccount, so it falls back to helm and the
e2e suite and laptop runs are unchanged. An API error raises rather than
falling back: the fallback is for a known-absent ServiceAccount, and quietly
retrying through helm would swap a visible error for the timeout this removes.
chart and app_version are empty on this path rather than wrong — they live
only in the compressed payload, and nothing reads them.