mirror of
https://github.com/certctl-io/certctl.git
synced 2026-10-05 05:39:02 +02:00
Phase 4 of the certctl architecture diligence remediation closure.
Seven findings, all in deploy/helm/certctl/.
DEPL-H2 (High) — ship deploy/helm/certctl/templates/backup-cronjob.yaml
Operator opt-in via backup.enabled=true. Default OFF. CronJob runs
pg_dump --format=custom --no-owner --no-acl --dbname=certctl
matching the canonical shape in
docs/operator/runbooks/postgres-backup.md (so manual and
automated dumps are byte-identical). Sink: PVC (default) OR S3
via aws-cli. Documented as in-cluster-Postgres only — managed DB
deployments rely on their provider's PITR.
DEPL-M1 (Med) — Helm pre-install/pre-upgrade migration hook
deploy/helm/certctl/templates/migration-job.yaml — runs
`certctl-server --migrate-only` before the server Deployment
rolls. The --migrate-only flag (new in cmd/server/main.go) is a
hermetic schema-mutation pass: load config, open DB pool, run
RunMigrations + RunSeed, exit 0. No HTTP listener, no scheduler,
no signing setup.
Server's boot-time RunMigrations call is now gated on
CERTCTL_MIGRATIONS_VIA_HOOK — when set true, the server skips
the boot path (the hook owns the work). Default still runs at
boot, so Compose / VM / bare-metal deploys are unchanged.
migrations.viaHook: false in values.yaml (off by default).
DEPL-M4 (Med) — explicit Postgres StatefulSet strategy fields
deploy/helm/certctl/templates/postgres-statefulset.yaml adds:
spec.updateStrategy.type: OnDelete
spec.podManagementPolicy: OrderedReady
Operator-controlled Postgres upgrades (the OnDelete strategy
means a chart template tweak no longer triggers an immediate
Postgres restart). OrderedReady aligns with the standard
Postgres-on-Kubernetes pattern for any future HA work.
DEPL-M5 (Med) — per-fleet-size resource ladder documentation
deploy/helm/certctl/values.yaml — extended comments next to
server.resources + agent.resources documenting:
"≤ 500 certs / 100 agents" → defaults are validated
"5K certs / 1K agents" → starter suggestions, TBD Phase 8
"50K certs / 10K agents" → starter suggestions, TBD Phase 8
Numbers for the small-fleet case derive from the measured
baselines in docs/operator/performance-baselines.md
(50ms p50, < 3s for 1000-cert inventory walk, etc.). Larger
fleet numbers explicitly marked TBD pending Phase 8 load-test
runs — operators tune empirically until then.
DEPL-L1 (Low) — Helm rollback runbook
docs/operator/runbooks/rollback.md — covers helm rollback
mechanics, the schema-migration manual-cleanup path (when
*.down.sql files apply vs. when full restore is the only safe
path), and the per-migration-class safe-to-rollback table.
DEPL-L2 (Low) — Prometheus AlertManager rules
deploy/helm/certctl/templates/prometheusrules.yaml — opt-in via
monitoring.prometheusRules.enabled=true. Default OFF. Four
starter rules using verified metric names from
internal/api/handler/metrics.go:
CertctlCertificateExpiringSoon (certctl_certificate_expiring_soon)
CertctlAgentOffline ((agent_total - agent_online) > 0 for 1h)
CertctlJobFailureRateHigh (failure rate over 5% for 15m)
CertctlIssuanceFailures (any failures over 15m window)
All thresholds operator-tunable via
monitoring.prometheusRules.thresholds.* in values.
DEPL-L3 (Low) — Prometheus bearer-token setup runbook
docs/operator/runbooks/prometheus-bearer-token.md — documents
the API-key + Secret + values wiring for the RBAC-gated
/api/v1/metrics/prometheus scrape endpoint. End-to-end
procedure with troubleshooting steps + rotation guide.
CI guard: scripts/ci-guards/helm-templates-lint.sh
Six-combo matrix: defaults / backup PVC / backup S3 /
prometheusRules / migrations.viaHook / all-on. Each runs helm
template + checks render success. helm lint also gated.
Wired into the auto-pickup loop in .github/workflows/ci.yml;
azure/setup-helm@b9e51907 (v4.3.0, SHA-pinned per Phase 1
RED-2) installs helm v3.16.0 on the runner.
Verification (all pass):
ls deploy/helm/certctl/templates/{backup-cronjob,migration-job,prometheusrules}.yaml
grep -E 'updateStrategy|podManagementPolicy' deploy/helm/certctl/templates/postgres-statefulset.yaml # 2 matches
helm template deploy/helm/certctl/ --set backup.enabled=true \
--set monitoring.prometheusRules.enabled=true --set migrations.viaHook=true \
| grep -E "kind: (CronJob|PrometheusRule|Job)" # 3 matches
helm lint deploy/helm/certctl/ # 0 failed
ls docs/operator/runbooks/{rollback,prometheus-bearer-token}.md
bash scripts/ci-guards/helm-templates-lint.sh # 6/6 matrix combinations pass
Go build clean (cmd/server compiles, migrate-only path verified by
the build target). YAML validated.
Closes: cowork/certctl-architecture-diligence-audit.html#fix-DEPL-H2
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-M1
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-M4
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-M5
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-L1
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-L2
cowork/certctl-architecture-diligence-audit.html#fix-DEPL-L3
179 lines
8.1 KiB
YAML
179 lines
8.1 KiB
YAML
{{- /*
|
|
Phase 4 DEPL-H2 closure (2026-05-14): opt-in Helm CronJob for
|
|
PostgreSQL backups.
|
|
|
|
OPERATOR OPT-IN. Default `backup.enabled: false`. Turning it on
|
|
requires:
|
|
- In-cluster Postgres (this CronJob does NOT cover managed DB
|
|
services — for AWS RDS / GCP CloudSQL / Azure DB rely on the
|
|
provider's PITR).
|
|
- A sink choice (PVC or S3) configured in values.yaml.
|
|
- For S3: a Secret holding AWS_ACCESS_KEY_ID + AWS_SECRET_ACCESS_KEY
|
|
(or use a service account with IRSA on EKS).
|
|
|
|
The pg_dump invocation matches the canonical shape documented in
|
|
docs/operator/runbooks/postgres-backup.md so a manual run and a
|
|
CronJob run produce byte-identical dumps:
|
|
|
|
pg_dump --format=custom --no-owner --no-acl --dbname=certctl
|
|
|
|
For sink choices beyond PVC + S3 (GCS, Azure Blob, NFS, restic, etc.),
|
|
extend the `aws s3 cp` line below. The Job is intentionally minimal —
|
|
it does ONE thing (capture + ship), not orchestrate retention or
|
|
rotation. Off-host retention is the sink's responsibility (S3 lifecycle
|
|
rules, PVC snapshot retention on the storage class, etc.).
|
|
*/ -}}
|
|
{{- if .Values.backup.enabled }}
|
|
apiVersion: batch/v1
|
|
kind: CronJob
|
|
metadata:
|
|
name: {{ include "certctl.fullname" . }}-postgres-backup
|
|
labels:
|
|
{{- include "certctl.labels" . | nindent 4 }}
|
|
app.kubernetes.io/component: postgres-backup
|
|
spec:
|
|
schedule: {{ .Values.backup.schedule | quote }}
|
|
concurrencyPolicy: Forbid
|
|
successfulJobsHistoryLimit: {{ .Values.backup.successfulJobsHistoryLimit | default 3 }}
|
|
failedJobsHistoryLimit: {{ .Values.backup.failedJobsHistoryLimit | default 1 }}
|
|
startingDeadlineSeconds: {{ .Values.backup.startingDeadlineSeconds | default 300 }}
|
|
jobTemplate:
|
|
spec:
|
|
backoffLimit: {{ .Values.backup.backoffLimit | default 1 }}
|
|
activeDeadlineSeconds: {{ .Values.backup.activeDeadlineSeconds | default 3600 }}
|
|
template:
|
|
metadata:
|
|
labels:
|
|
{{- include "certctl.labels" . | nindent 12 }}
|
|
app.kubernetes.io/component: postgres-backup
|
|
spec:
|
|
restartPolicy: Never
|
|
{{- with .Values.imagePullSecrets }}
|
|
imagePullSecrets:
|
|
{{- toYaml . | nindent 12 }}
|
|
{{- end }}
|
|
serviceAccountName: {{ include "certctl.serviceAccountName" . }}
|
|
securityContext:
|
|
runAsUser: 1000
|
|
runAsGroup: 1000
|
|
runAsNonRoot: true
|
|
fsGroup: 1000
|
|
containers:
|
|
- name: backup
|
|
image: {{ .Values.backup.image | default "postgres:16-alpine" | quote }}
|
|
imagePullPolicy: {{ .Values.backup.imagePullPolicy | default "IfNotPresent" | quote }}
|
|
env:
|
|
- name: PGHOST
|
|
value: {{ include "certctl.fullname" . }}-postgres
|
|
- name: PGPORT
|
|
value: {{ .Values.postgresql.service.port | default 5432 | quote }}
|
|
- name: PGUSER
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: {{ include "certctl.fullname" . }}-postgres
|
|
key: username
|
|
- name: PGPASSWORD
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: {{ include "certctl.fullname" . }}-postgres
|
|
key: password
|
|
- name: PGDATABASE
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: {{ include "certctl.fullname" . }}-postgres
|
|
key: database
|
|
{{- if eq (.Values.backup.sink | default "pvc") "s3" }}
|
|
# S3 sink — operator provides AWS credentials via the
|
|
# Secret referenced in backup.s3.credentialsSecret. The
|
|
# credentials need s3:PutObject + s3:ListBucket on the
|
|
# target bucket only; least-privilege per industry
|
|
# standard.
|
|
- name: AWS_ACCESS_KEY_ID
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: {{ .Values.backup.s3.credentialsSecret.name | quote }}
|
|
key: {{ .Values.backup.s3.credentialsSecret.accessKeyIdKey | default "AWS_ACCESS_KEY_ID" }}
|
|
- name: AWS_SECRET_ACCESS_KEY
|
|
valueFrom:
|
|
secretKeyRef:
|
|
name: {{ .Values.backup.s3.credentialsSecret.name | quote }}
|
|
key: {{ .Values.backup.s3.credentialsSecret.secretAccessKeyKey | default "AWS_SECRET_ACCESS_KEY" }}
|
|
{{- with .Values.backup.s3.region }}
|
|
- name: AWS_DEFAULT_REGION
|
|
value: {{ . | quote }}
|
|
{{- end }}
|
|
{{- end }}
|
|
command:
|
|
- /bin/sh
|
|
- -ceu
|
|
- |
|
|
# Phase 4 DEPL-H2: canonical pg_dump shape per
|
|
# docs/operator/runbooks/postgres-backup.md.
|
|
# Custom-format compressed dump, no ownership /
|
|
# ACL embedded — produces a portable artifact
|
|
# restorable into any Postgres ≥ source major
|
|
# via `pg_restore -d certctl <dump>`.
|
|
set -euo pipefail
|
|
TIMESTAMP="$(date -u +%Y%m%dT%H%M%SZ)"
|
|
DUMP_FILE="/tmp/certctl-${TIMESTAMP}.dump"
|
|
|
|
echo "[backup-cronjob] capturing dump at ${TIMESTAMP}"
|
|
pg_dump --format=custom --no-owner --no-acl --dbname="${PGDATABASE}" \
|
|
> "${DUMP_FILE}"
|
|
|
|
# Integrity check — pg_restore --list parses the
|
|
# dump's table-of-contents; a corrupt dump fails
|
|
# here without shipping garbage off-host. Same
|
|
# check the manual runbook performs.
|
|
echo "[backup-cronjob] verifying dump integrity"
|
|
pg_restore --list "${DUMP_FILE}" > /dev/null
|
|
|
|
{{- if eq (.Values.backup.sink | default "pvc") "s3" }}
|
|
# S3 sink — requires aws-cli. The default
|
|
# postgres:16-alpine image does NOT include
|
|
# aws-cli; operators MUST set
|
|
# backup.image to an image that bundles both
|
|
# (e.g. ghcr.io/your-org/postgres-aws:16) OR
|
|
# override backup.command to install aws-cli at
|
|
# runtime. The line below assumes the image has
|
|
# `aws` on PATH.
|
|
S3_PATH="{{ .Values.backup.s3.bucket }}/{{ .Values.backup.s3.prefix | default "certctl" }}/certctl-${TIMESTAMP}.dump"
|
|
echo "[backup-cronjob] uploading to s3://${S3_PATH}"
|
|
aws s3 cp "${DUMP_FILE}" "s3://${S3_PATH}"
|
|
rm -f "${DUMP_FILE}"
|
|
{{- else }}
|
|
# PVC sink — dump lands at /backups/certctl-${TIMESTAMP}.dump
|
|
# mounted from backup.pvc.claimName. Retention is the
|
|
# PVC's responsibility (storage-class snapshot lifecycle
|
|
# or a separate cleanup CronJob). The Job moves the
|
|
# file from /tmp to /backups atomically; never
|
|
# writes partial dumps into the durable mount.
|
|
FINAL_PATH="/backups/certctl-${TIMESTAMP}.dump"
|
|
echo "[backup-cronjob] persisting to ${FINAL_PATH}"
|
|
mv "${DUMP_FILE}" "${FINAL_PATH}"
|
|
{{- end }}
|
|
echo "[backup-cronjob] done"
|
|
{{- if ne (.Values.backup.sink | default "pvc") "s3" }}
|
|
volumeMounts:
|
|
- name: backups
|
|
mountPath: /backups
|
|
{{- end }}
|
|
resources:
|
|
{{- toYaml (.Values.backup.resources | default dict) | nindent 16 }}
|
|
{{- if ne (.Values.backup.sink | default "pvc") "s3" }}
|
|
volumes:
|
|
- name: backups
|
|
persistentVolumeClaim:
|
|
claimName: {{ .Values.backup.pvc.claimName | quote }}
|
|
{{- end }}
|
|
{{- with .Values.nodeAffinity }}
|
|
affinity:
|
|
nodeAffinity:
|
|
{{- toYaml . | nindent 14 }}
|
|
{{- end }}
|
|
{{- with .Values.backup.tolerations }}
|
|
tolerations:
|
|
{{- toYaml . | nindent 12 }}
|
|
{{- end }}
|
|
{{- end }}
|