Saltar al contenido principal

Game-Day: Zero-Orphaned-Order Drain During Trading Hours

Phase 4.6(a) game-day from the Unified Infrastructure Control Plane Plan (private alphaswarm_internal planning repo) (relocated out of this repo on 2026-07-18; see ADR 026) (§4, item 4.6). This is the repeatable exercise that proves the drain choreography leaves no open order behind when a live-trading pod is terminated during market hours. The mechanisms it rehearses are the ones shipped for money-plane workloads — it does not introduce new commands.

Objective​

Prove, with evidence, that terminating a live-trading workload pod during trading hours results in zero open orders at the venue at the moment of termination, because the drain choreography either (a) cancels/flattens every open order before the container exits, or (b) blocks termination until it can. The exit criterion mirrors the plan's Phase 4 gate: the game-day passes only with evidence archived to the WORM ledger.

Grounded mechanisms​

The drain choreography is a layered fail-closed sequence. Each layer is a real mechanism this program built or mandated:

LayerMechanismSource
SchedulerKarpenter karpenter.sh/do-not-disrupt annotation on live-trading pods so the autoscaler never voluntarily evicts them mid-sessionPhase 4.5 progressive-delivery guardrails
AvailabilityPodDisruptionBudget (PDB) blocks voluntary evictions that would drop below the trading-critical replica floorK8s-native drain layer
ContainerpreStop hook issues a cancel/flatten on the pod's open orders before SIGTERM (same cancel/flatten semantics as the bots KillSwitch modes)Bots operator drain path
OperatorThe kopf finalizer is fail-closed: with drain_fail_closed: true the finalizer is KEPT on drain timeout — termination is blocked until drain confirmsAlphaSwarmServiceSpec.drain_fail_closed; pythonic-unified-control-service §3.2, G-4
RolloutStateful trading workloads use a partitioned rolling update, never a canary/blue-green (prohibited on StatefulSets)Phase 4.5; ADR 026 standing prohibitions
CalendarTrading-hours sync windows (ArgoCD) + change-freeze config gate when GitOps may act at allPhase 4.5; pythonic-unified §3.2 (GitOps)

drain_fail_closed is documented as "the operator finalizer is KEPT on drain timeout (termination blocked until drain confirms) — mandatory posture for live-trading workloads" and pairs with drain_timeout_seconds (default 300, max 7200).

Preconditions / scope​

  • Environment: a non-production trading cell running a live-trading workload configured with drain_fail_closed: true and a finite drain_timeout_seconds. Never run first against a production silo-reg cell.
  • Karpenter do-not-disrupt, the PDB, the preStop flatten hook, and the ArgoCD trading-hours sync window are all in effect for the target workload.
  • A venue/paper sandbox with observable open-order state (order book queryable by the workload's account) so "zero open orders at terminate" is measurable.
  • The controller halt surface is reachable for the abort path (/manage/workloads/halt; see Game-Day: Kill-Switch Fleet Halt + Resume).
  • The WORM audit uploader + integrity verifier CronJobs are running (evidence destination).

Roles​

RoleResponsibility
Exercise lead (SRE)Runs the procedure, holds the incident ticket, calls pass/abort
Trading approverConfirms the target account is inside a synthetic session, owns the order-book observation
Platform on-callWatches the operator/finalizer state; drives the abort halt if invoked
ScribeCaptures timestamps, order counts, and the audit_run_ids into the ticket

Step-by-step procedure​

  1. Open the exercise ticket. Record start time, target cell, workload id, drain_timeout_seconds, and the current ArgoCD sync-window state.
  2. Seed open orders. In the synthetic session, place resting orders on the target account so there is a non-zero open-order count to drain. Record the count N_open_before from the order book.
  3. Confirm the guards are armed. Verify the pod carries karpenter.sh/do-not-disrupt, the PDB shows an available budget, and the workload spec reports drain_fail_closed: true.
  4. Trigger a governed termination. Advance the partitioned rolling update by one partition (or, for the manual variant, delete exactly one trading pod). This is the eviction the choreography must survive. Do not use a canary — canary/blue-green on StatefulSets is prohibited.
  5. Observe the drain. The preStop hook fires cancel/flatten; watch the order book converge to zero for that pod's account. The kopf finalizer holds the pod object until drain confirms.
  6. Record the terminate instant. At the moment the container actually exits (finalizer removed), snapshot the venue open-order count N_open_at_terminate for that account.
  7. Negative sub-case (fail-closed proof). Re-run once with a deliberately unreachable venue so flatten cannot complete within drain_timeout_seconds. Confirm the finalizer is KEPT, the pod stays Terminating, and the drain-timeout alert fires — i.e. the system refuses to strand orders.
  8. Resolve the negative sub-case via the abort path (below), then restore the venue and let the drain complete cleanly.
  9. Close the ticket with the measured counts and the audit references.

Success criteria (measurable)​

  • Primary: N_open_at_terminate == 0 for the drained account at the instant the container exits — zero orphaned orders, proven from the venue order book, not inferred.
  • The preStop cancel/flatten completed within drain_timeout_seconds; the finalizer was removed only after confirmation.
  • Negative sub-case: with flatten blocked, the finalizer was KEPT, the pod never terminated, and the drain-timeout alert fired — no silent strand.
  • No Karpenter voluntary eviction and no PDB-violating eviction occurred during a sync-window/freeze period.
  • Every terminate/finalizer transition and the halt (if used) landed a hash-chained audit row (see Evidence capture).

Abort / rollback​

  • Abort trigger: order-book count is not converging to zero, PnL is bleeding, or the negative sub-case must be stopped.
  • Action: engage the fleet halt — POST /manage/workloads/halt (scope workloads:halt) to signal every in-flight WorkloadRun to abort, and for the bot fleet apply a KillSwitch at scope: fleet, mode: flatten per the Kill-Switch Incident Response runbook. For live-trading workloads the halt also sets the order-gate key (kill-switch honesty), so no new orders can be placed while halted.
  • Rollback: pause the partitioned rollout, let the finalizer complete the drain, then resume via the Kill-Switch Fleet Halt + Resume procedure. Never force-delete a Terminating trading pod (--grace-period=0 --force) — that is the exact orphaned-order failure the game-day exists to prevent.

Evidence capture (WORM / audit ledger)​

Every mutating step lands a hash-chained workload_runs row through the gate (JSONL locally → HTTP fan-out to the monolith Postgres ledger → S3 Object Lock COMPLIANCE WORM). For this game-day, archive:

  • The pod terminate + finalizer add/remove events with timestamps.
  • N_open_before and N_open_at_terminate from the order book.
  • The drain-timeout alert payload and finalizer-KEPT proof from the negative sub-case.
  • Any status=HALTED WorkloadRun rows if the abort halt was used, plus the HaltRequest.reason string.
  • The audit_run_ids, verified intact with the integrity verifier (python -m alphaswarm_controller.terraform.audit_verify, exit≠0 on a chain break — see the DR replay game-day).

Frequency / owner​

  • Frequency: quarterly, and before any change to the drain/finalizer path or trading-hours sync-window config.
  • Owner: sre-team (calendar reminder), with the trading approver co-signing the evidence.