Customer offboarding and teardown
Procedure for winding down an enterprise customer: contract termination, license revocation, governed infrastructure destroy, data retention, and org suspension — in that order. Destroy is the highest-blast-radius operation in the platform; it only ever runs as a plan-bound destroy through the controller gate (a destroy without a reviewed destroy-plan binding is rejected).
Prerequisites
- Commercial confirmation of the end date + the data-retention terms from the contract (export window, deletion deadline).
- You hold
manage:tenants+manage:infrastructure+terraform:admin; second approver available. - Customer notified of the cutoff schedule.
Steps
- Lifecycle →
offboardingon the customer record (freezes the commercial picture; audit anchor for everything below). - Terminate the contract — contracts table → Terminate. Resolved entitlements empty out; new license issuance for this customer stops (issuance copies entitlements from the active contract).
- Revoke licenses — deployment-scoped revoke for every active lease (see license-issuance-and-revocation). The install decays through expiry + grace; time the destroy after the export window, not the grace window.
- Data export (if contracted) — snapshot the tenant KB silo from the customer account (RDS snapshot + S3 sync) to the agreed hand-off location before any destroy.
- Destroy plan — Terraform page → the
customer-<slug>-<env>workspace → runplan -destroy. Review the full resource list. The OPA customer gate requires stateful deletes to carryalphaswarm:teardown=approved— tag the stateful resources via the reviewed teardown change first (that tag change is itself a plan). - Four-eyes destroy apply — the destroy binding pins the reviewed
destroy plan (
PlanBinding.is_destroy); the second operator approves; apply executes exactly that plan. - Sync + close the record — deployment row
destroying → destroyed. PATCH remaining metadata (final run ids stay linked for audit). - Suspend the org — org status →
suspended(sign-ins stop; data rows remain for the retention window). After the contractual deletion deadline: purge per the tenant-purge flow (/tenants/purge). - Lifecycle →
churned.
Post-action verification
- Customer AWS account: no
alphaswarm-*resources remain (module outputs empty; spot-check RDS/S3/KMS in the console). - Ops account: the state object still exists (retained — it is the audit record of what was destroyed); the workspace row is archived.
- Registry: all leases
revoked; audit ledger shows the full chain (terminate → revoke → destroy plan → approve → destroy). - Org
suspended; customer lifecyclechurned.
Escalation
- Destroy blocked by OPA on stateful resources → that is the gate working; complete the teardown-tag change, never bypass.
- Partial destroy (dependency cycle) → re-plan destroy; the remaining
graph shrinks each pass. Manual console deletes are the last resort
and must be reconciled with
terraform state rm+ documented. - Customer disputes deletion after the fact → the retained state file + audit chain is the evidence trail.
Game-day: BYOC tenant offboard with crypto-shred verification
Phase 4.6(d) game-day from the Unified Infrastructure Control Plane Plan (private
alphaswarm_internalplanning repo) (§4, item 4.6). This extends the teardown procedure above with the terminal crypto-shred step and its unrecoverability proof, rehearsing the offboarding saga's compensation model from ADR 026 (§7) and thesilo-regBYOK/Transit key custody from ADR 027 — Cell isolation tiers.
Objective
Prove that offboarding a BYOC / silo-reg tenant ends in a crypto-shred that
renders the tenant's data unrecoverable — the tenant BYOK CMK is destroyed — and
that this is a human-approved terminal compensation, not an automatic saga
side-effect. Success = the CMK is destroyed and a subsequent read/decrypt of
the tenant's at-rest data fails, with the full offboarding saga audit trail as
evidence.
Grounded mechanisms
- Offboarding saga: the Tenant Provisioning Saga "offboarding reverses with
crypto-shredding as terminal compensation" (plan Phase 2.1). Steps run on the
in-process compensating saga engine; the runner consults the halt store
before every step and refuses destructive steps without an
approval_request_id(pythonic-unified P-7). - Crypto-shred guard: crypto-shredding is a "terminal, human-approved
compensation" and, per the guard, Transit keys are never deleted by
sagas — "Transit keys/CMKs are never deleted by an automatic path". The
shred destroys the tenant BYOK CMK (
tenant_byok_cmk, the AWS-KMS key that encrypts the tenant's RDS/S3 at rest); the per-cell Vault Transit key is retained (the guard protects it from saga deletion). Anchor: plan §3 component map ("crypto-shred guard: Transit keys never deleted by sagas"), ADR 027 Tier-3silo-reg(Vault Transit + BYOK CMK). - Unrecoverability: with the CMK destroyed, envelope keys can no longer be unwrapped, so the encrypted RDS snapshot / S3 objects cannot be decrypted — the data is cryptographically shredded even though ciphertext may still exist.
Preconditions / scope
- A rehearsal BYOC /
silo-regtenant in a disposable account (never a live customer). Its data plane is encrypted with a dedicated BYOK CMK and, where applicable, a per-cell Vault Transit key (ADR 027 Tier 3). - The commercial reversal steps (terminate → revoke → plan-bound destroy → suspend) from the procedure above are complete or rehearsed to the point where the CMK destroy is the remaining terminal step.
- A pre-recorded read/decrypt probe against the tenant's at-rest data that succeeds before the shred (so the post-shred failure is a controlled before/after).
- A signed
approval_request_idfor the destructive crypto-shred step; four-eyes and step-up available.
Roles
| Role | Responsibility |
|---|---|
| Offboarding lead (SRE) | Runs the reverse saga to the terminal step |
| Second approver | Provides the distinct four-eyes approval + approval_request_id |
| Security officer | Witnesses the CMK destruction and the failed decrypt probe |
| Scribe | Captures the saga step audit rows and the before/after probe results |
Step-by-step procedure
- Complete the reversal. Run the offboarding steps above (terminate → revoke → data export if contracted → plan-bound destroy → suspend) against the rehearsal tenant, up to the point the CMK remains.
- Baseline the probe. Run the read/decrypt probe against the tenant's encrypted RDS snapshot / S3 objects and confirm it succeeds (data is recoverable while the CMK lives). Record the result.
- Assert the guard (negative sub-case). Attempt the crypto-shred without
an
approval_request_id; confirm the saga runner refuses it (destructive step needs approval) — proving the terminal compensation cannot fire automatically. - Approve + execute the crypto-shred. With four-eyes + step-up, supply the
approval_request_idand execute the terminal compensation: schedule destruction of / disable the tenant BYOK CMK (tenant_byok_cmk). - Confirm Transit-key retention. Verify the per-cell Vault Transit key was NOT deleted — the crypto-shred guard retains it. Only the tenant CMK is destroyed.
- Prove unrecoverability. Re-run the probe from step 2; the decrypt/read MUST now fail (KMS key unavailable → envelope keys cannot be unwrapped).
- Verify the saga audit trail. Confirm every step (terminate → revoke →
destroy plan → approve → destroy → crypto-shred) landed a ledger row, and that
the crypto-shred row carries the
approval_request_id. - Document in the post-action verification section above.
Success criteria (measurable)
- The tenant BYOK CMK is destroyed (pending-deletion / disabled), and the per-cell Vault Transit key is retained (guard honoured).
- The decrypt/read probe succeeded before and failed after the shred — a controlled proof of unrecoverability.
- The unapproved crypto-shred attempt was refused (no destructive step
without
approval_request_id). - The offboarding saga audit trail is complete and the crypto-shred row is bound to its approval.
Abort / rollback
- Crypto-shred is irreversible by design — once the CMK is destroyed the data is unrecoverable. The rollback window is before step 4: if the export or retention terms are not fully satisfied, halt the saga (the runner consults the halt store before every step) and do not approve the shred.
- If the CMK is only scheduled for deletion (KMS pending-deletion window), cancellation during that window is the sole recovery path; after the window closes there is none.
- Never destroy the Vault Transit key as part of the shred — the guard forbids it, and doing so would exceed the terminal compensation's scope.
Evidence capture (WORM / audit ledger)
- The full offboarding saga audit trail:
terminate → revoke → destroy plan → approve → destroy → crypto-shred, each a hash-chained ledger row. - The crypto-shred row with its bound
approval_request_idand the CMK id/ARN (never key material). - The before/after decrypt-probe results and the retained-Transit-key confirmation.
- Rows fan out to the monolith Postgres ledger and archive to S3 Object Lock
COMPLIANCE WORM; verify the chain with
python -m alphaswarm_controller.terraform.audit_verify(exit≠0 on a break — see the DR replay game-day).
Frequency / owner
- Frequency: semi-annually and on every real BYOC offboarding (the real offboarding is the exercise, with the same evidence bar).
- Owner:
sre-team, security officer co-signs the unrecoverability proof.