Saltar al contenido principal

DR replay runbook

Disaster-recovery rehearsal procedure for AlphaSwarm. Targets:

  • RPO 1 hour for alphaswarm_admin + control-plane services.
  • RTO 4 hours for the same.
  • RPO 15 minutes for trading-relevant data.
  • RTO 1 hour for the same.

The exercise is run quarterly (calendar reminder owned by the platform team). The first exercise is scheduled for the end of Phase 5 of the multi-account overhaul.

Pre-requisites​

  • AWS Organizations + Control Tower applied (Phase 4 complete).
  • ArgoCD app-of-apps applied to dev + staging + prod clusters.
  • Velero installed on every workload cluster (chart at alphaswarm_platform/deployments/kubernetes/helm/velero).
  • ECR cross-region replication active to us-west-2.
  • RDS cross-region read replica green.
  • S3 CRR active on every Parquet + audit-archive bucket.
  • Route 53 health-check failover record set on the manage.alpha-swarm.ai ingress.

Steps​

1. Trigger the failure​

Pick the rehearsal target — typically alphaswarm-dev (never prod). Document the start time in the incident ticket.

# Disable the dev cluster's API server (simulates a control-plane outage).
aws eks update-cluster-config \
--name alphaswarm-dev \
--region us-east-1 \
--resources-vpc-config endpointPrivateAccess=false,endpointPublicAccess=false

2. Confirm impact​

alphaswarm_admin should now show unreachable for the dev cluster under /admin/kubernetes/status. The KillSwitch should still work because it fans out to other clusters too.

3. Bring up the replay cluster​

cd infrastructure/envs/dev
terraform apply -var-file=terraform.tfvars

This re-creates the EKS cluster with the same name + node groups. ArgoCD picks up the new cluster via its Cluster generator (label alphaswarm.io/managed=true).

4. Replay state from Velero​

velero backup-location get
velero restore create dr-replay-$(date +%s) \
--from-backup daily-full-$(velero backup get | tail -1 | awk '{print $1}')

5. Restore RDS​

The cross-region read replica in us-west-2 is promoted to primary; the DR replay points the dev cluster's RDS DSN at the new primary. The Postgres instance comes up with the audit ledger intact so no admin actions are lost.

6. Verify​

  • alphaswarm_admin health should return 200 within 4h.
  • The audit ledger should show the gap as a single contiguous block (no missing rows beyond the RPO window).
  • Paper-trading runs that were active are stamped status=halted by the watchdog.
  • The ArgoCD app-of-apps sync should converge within 15min after the cluster comes back.

7. Document​

Append to the rehearsal log at alphaswarm_docs/docs/operations/dr-rehearsal-log.md with:

  • Start / end timestamps.
  • Actual RPO + RTO measured.
  • Issues encountered + remediations.
  • Sign-off from the security officer.

Game-day: DR replay from encrypted state and WORM ledger​

Phase 4.6(c) game-day from the Unified Infrastructure Control Plane Plan (private alphaswarm_internal planning repo) (§4, item 4.6). This extends the Velero/RDS rehearsal above with the control-plane's own recovery story: restoring a cell from OpenTofu-encrypted state with its per-cell KMS key, and proving the hash-chained WORM audit ledger survived intact. It rehearses the mechanisms decided in ADR 028 — OpenTofu cutover and the audit story in Central Deployment Control — Next Steps.

Objective​

Prove that a cell can be fully rebuilt from its encrypted OpenTofu state using only the per-cell KMS key, and that the tamper-evident audit ledger for that cell verifies clean end-to-end after the recovery. Success = a restored cell whose next plan shows no diff and an audit chain that audit_verify accepts.

Grounded mechanisms​

  • Encrypted state (ADR 028): the OpenTofu encryption{} state block (AWS-KMS key provider, per-cell key) renders only when the resolved binary is tofu. ADR 028 §4 specifies the exact validation this game-day performs: "validating state-encryption round-trips (encrypt → destroy state → restore with the KMS key)" and -json/exit-code parity against the terraform baseline. State is owned by exactly one binary at a time (no dual-write).
  • WORM ledger: the hash-chained JSONL ledger is fanned out to the monolith Postgres ledger and archived to S3 Object Lock COMPLIANCE (SSE-KMS, ≥6y) by WormUploader (python -m alphaswarm_controller.terraform.audit_worm).
  • Integrity verifier: python -m alphaswarm_controller.terraform.audit_verify walks the chain (entry_hash = sha256(prev_hash ‖ canonical(row))) and exits non-zero on a chain break (sev-1) — the pass/fail oracle for this game-day.

Preconditions / scope​

  • A rehearsal cell whose IaC runs under tofu with encryption enabled (never a production silo-reg cell for the first run).
  • Access to the cell's per-cell KMS key (and a way to test the negative case: attempt a restore without the key and confirm it fails).
  • The cell's WORM bucket and a recent archived ledger snapshot are present.
  • Operator holds manage:infrastructure for the restore plan and admin:cluster for any apply; four-eyes for the apply.

Roles​

RoleResponsibility
DR lead (SRE)Runs the restore, holds the incident ticket
Security officerWitnesses the KMS-key-only restore and signs the chain-verify result
ApproverDistinct four-eyes approver for the restore apply
ScribeRecords measured RTO/RPO, the no-diff plan, and the audit_verify exit code

Step-by-step procedure​

  1. Record the target state. Note the cell's encrypted state object location and the tip hash of its WORM-archived ledger.
  2. Simulate loss. Following the state-encryption round-trip discipline, destroy/withdraw the local state so recovery must come from the encrypted backend copy.
  3. Restore with the KMS key. Reconfigure the backend with the cell's per-cell KMS key and run a governed plan on the cell's workspace through the One Gate. OpenTofu decrypts the state with the KMS key provider.
  4. Negative sub-case (key-custody proof). Repeat step 3 with the KMS key denied; confirm the restore fails (state cannot be decrypted) — this proves custody of the key is load-bearing.
  5. Confirm no drift. The restored plan must show no diff versus the pre-loss cell (ADR 028's no-diff migration gate), and -json/exit-code behaviour must match the terraform baseline.
  6. Pull the WORM ledger snapshot for the cell from the Object Lock bucket.
  7. Verify the chain. Run python -m alphaswarm_controller.terraform.audit_verify against the recovered ledger. Exit 0 = intact; exit ≠ 0 = chain break (fail the game-day, raise sev-1).
  8. Tamper sub-case (verifier proof). Mutate one archived row in a copy and re-run audit_verify; confirm it exits non-zero — proving the verifier actually catches breaks.
  9. Document in the rehearsal log alongside the RPO/RTO section above.

Success criteria (measurable)​

  • The cell restored from encrypted state using only the per-cell KMS key; the key-denied sub-case failed to restore.
  • The post-restore plan showed no diff and matched -json/exit-code parity.
  • audit_verify exited 0 on the recovered ledger; the tamper sub-case exited non-zero.
  • Measured RTO/RPO within the targets at the top of this runbook.

Abort / rollback​

  • If the restored plan shows an unexpected diff, do not apply — the state or the encryption context is wrong; stop and reconcile (never -auto-approve a drifted restore).
  • If audit_verify fails, treat the ledger as compromised: preserve the WORM object (Object Lock prevents deletion before retention), open a sev-1, and do not overwrite the archive.
  • Recovery of the cluster/data layer (Velero/RDS) rolls back via the main DR procedure above; this game-day adds no destructive action beyond the disposable rehearsal cell.

Evidence capture (WORM / audit ledger)​

  • The restore plan output (no-diff proof) and the -json parity capture.
  • The audit_verify exit codes for both the clean and tamper sub-cases.
  • The negative key-denied restore failure log.
  • The recovered WORM ledger snapshot itself is the immutable evidence (COMPLIANCE-locked); its tip hash is recorded in the rehearsal log.

Frequency / owner​

  • Frequency: quarterly, folded into the DR rehearsal cadence above; also after any ADR 028 cutover step that flips a cell to tofu/encrypted state.
  • Owner: sre-team, security officer co-signs the chain-verify result.