Skip to main content

BYOC upgrade and rollback

Procedure for moving a live customer deployment to a new platform version — and back. An upgrade is a re-render + re-plan of the same workspace with a new app_version; it rides the identical gate as the initial provision (hard-mandatory OPA, plan binding, four-eyes, step-up).

Prerequisites​

  • The target image tag exists in the registry the customer account pulls from (immutable tags only — never latest).
  • Release notes checked for migration-bearing changes (new alembic revisions) — migrations are forward-only in production; that shapes the rollback plan (below).
  • A second approver is available.

Upgrade steps​

  1. Bump the version — customer detail → deployment row: PATCH app_version to the new tag (audit-first, step-up).
  2. Re-render the spec — Render spec. An active deployment transitions to upgrading; the spec hash changes only by the version (golden-render property — anything else changing is a red flag: stop and diff the spec).
  3. Plan on the Terraform page. Expect a small diff (ECS task-definition image tags — customer compute is ECS Fargate per ADR 031). Any stateful-resource replacement in the diff is a stop-and-review.
  4. Four-eyes apply the bound plan.
  5. DB migrations — the workloads run alembic on rollout (self-hosted alphaswarm-local installs auto-migrate with backup_on_upgrade taking a pre-upgrade pg_dump). Verify the schema head advanced (below) before calling it done.
  6. Sync the deployment row (upgrading → active).

Post-action verification​

  • Health endpoint reports the new version.
  • alembic_version in the customer DB equals the release's head.
  • License gating unchanged: entitled routes 200; the license header is absent (lease active).
  • Deployment row: active, app_version = new tag, apply run linked.

Rollback​

Rollback = the same governed path with the previous version:

  1. PATCH app_version back to the last-good tag → Render spec → plan → four-eyes apply.
  2. Schema caveat: alembic downgrades are NOT run against customer data. If the bad release carried migrations, roll the app back only if the previous app version tolerates the newer schema (usual case for additive migrations). If it does not, restore from the pre-upgrade backup (backup_on_upgrade dump / RDS snapshot) instead — that is a data-loss decision and needs the customer's sign-off.
  3. Sync the record; annotate meta with the incident link.

Escalation​

  • Rollout stuck (tasks crash-looping on new tag) → per ADR 031, customer compute is ECS Fargate (the EKS path is frozen, not used for BYOC): check the ECS service events / task-definition history in the customer account; rolling the service back to the previous task-definition revision buys time but the terraform record must be reconciled to match before the next plan.
  • Migration failed mid-apply → do not retry blindly; capture the alembic error, restore path per step 2 above.