Skip to main content

Halt-propagation staging enablement (Sprint S8)

Status: PLANNING — do not flip this flag ON in shared / production envs from this ticket. Defaults remain OFF until an operator explicitly enables staging.

Companion tickets: principal plan S8 (alphaswarm-first-class-module-transformation-plan.md), ADR 036 / 019, and the orchestration rollout orchestration-refactor-rollout.md.

1. Flag inventory (verified)​

Source of truth: alphaswarm/config/settings.py (env prefix ALPHASWARM_).

Settings fieldEnv varDefaultRole
orchestration_kill_propagation_enabledALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLEDFalsePrimary. Enables the workflow-run stall watchdog scan (scan_for_stalled_workflow_runs). Documented in settings as extending watchdog + KillSwitch fan-out into workflow_runs.
agent_watchdog_enabledALPHASWARM_AGENT_WATCHDOG_ENABLEDTrueAND-gate with kill-propagation inside _is_workflow_watchdog_enabled().
agent_watchdog_period_secondsALPHASWARM_AGENT_WATCHDOG_PERIOD_SECONDS60Celery beat period for workflow-stall-watchdog.
agent_stall_threshold_secondsALPHASWARM_AGENT_STALL_THRESHOLD_SECONDS300Age threshold before a running/pending workflow_runs row is halted.
orchestration_halt_check_timeout_secondsALPHASWARM_ORCHESTRATION_HALT_CHECK_TIMEOUT_SECONDS1.0Per-transition halt-check budget in WorkflowRuntime.
orchestration_studio_enabledALPHASWARM_ORCHESTRATION_STUDIO_ENABLEDFalseGates the entire /workflows/* surface including POST /workflows/halt (503 when off). KillSwitch still calls that path; other halt targets continue.
risk_kill_switch_keyALPHASWARM_RISK_KILL_SWITCH_KEYalphaswarm:kill_switchRedis key for the portfolio / global kill-switch signal (orthogonal to workflow watchdog).

.env.example already lists ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false.

Wiring today (do not change in shared envs)​

Hermetic coverage already present (do not weaken defaults):

  • tests/agents/test_orchestration_flags.py — flag defaults OFF
  • tests/tasks/test_workflow_watchdog.py — no-op when flag OFF; halt helper with fake row
  • tests/tasks/test_workflow_watchdog.py::test_workflow_watchdog_dual_gate — documents AND-gate wiring (S8 scaffold)

2. Preconditions (staging only)​

Before enabling in a dedicated staging namespace / compose overlay:

  1. Alembic 0046_workflow_versioning (and later workflow migrations) applied; workflow_runs table present.
  2. Celery worker + beat running with workflow-stall-watchdog registered (alphaswarm/tasks/celery_app.py).
  3. Decide whether KillSwitch workflow halt must succeed: if yes, also enable orchestration_studio_enabled in that same staging overlay (still staging-only).
  4. At least one non-prod WorkflowSpec that can be started and left idle / stalled for measurement.
  5. Operator can complete step-up MFA (rule 52) for halt endpoints.
  6. Observability stack reachable (API logs + Celery logs + optional Grafana/Loki). No money-plane / live trading workloads in the staging cell.

Explicit invariant: production and shared demo envs keep ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false until the evidence checklist in §6 is signed off.

3. Staging-only enablement steps​

Perform only against staging secrets / ConfigMaps / compose overrides.

  1. Confirm current value is false:

    # metadata only — do not dump full env
    docker exec <api> python -c "from alphaswarm.config import settings; print(settings.orchestration_kill_propagation_enabled)"
  2. Set staging override only:

    ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=true

    Optionally (if exercising KillSwitch → /workflows/halt):

    ALPHASWARM_ORCHESTRATION_STUDIO_ENABLED=true
  3. Restart API + Celery worker + Celery beat in staging.

  4. Start a disposable workflow run; leave it without breadcrumbs past agent_stall_threshold_seconds or trigger KillSwitch with step-up.

  5. Verify:

    • Watchdog path: workflow_runs.status='halted', halted=True, Celery task revoked when applicable.
    • KillSwitch path (studio on): POST /workflows/halt returns {ok: true, halted_count: N}; aggregate toast shows workflows in the fan-out summary.
  6. Leave flag ON in staging for a soak window (≥ 24 h recommended) while collecting §4 signals.

Do not commit staging overrides into shared platform values files used by prod.

4. Observability signals​

SignalWherePass criterion
Watchdog task resultCelery result / progress bus{ok: true, halted_count: …} shape from scan_for_stalled_workflow_runs
Halted rowsPostgres workflow_runsstatus='halted', error prefixed watchdog: or fan-out reason
KillSwitch aggregateOperator UI toastWorkflows endpoint success or expected 503 if studio still off
False positivesStaging audit / ops channelNo unexpected halt of healthy runs during soak
LatencyAPI / beat logsHalt observed within ~orchestration_halt_check_timeout_seconds after Redis/DB update (runtime next transition)
Agent watchdog still healthyGET /agents/health / data.agents.healthAgent stall scan unaffected

5. Rollback (staging or accidental shared flip)​

  1. Set ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false.
  2. Restart API + Celery beat (worker reload if settings cached per-process).
  3. Confirm _scan_and_halt_workflow_runs() returns [] with flag off (already unit-tested).
  4. Existing workflow_runs halted rows remain historical — no migration reverse required.
  5. If studio was enabled only for this experiment, restore ALPHASWARM_ORCHESTRATION_STUDIO_ENABLED=false independently.

KillSwitch continues to fan out to agents / paper / bots / rl / quant-agents / assistants / terraform / workloads / lab / ml serving regardless of this flag.

6. Test evidence required before prod​

GateEvidence
Unit / hermeticpytest tests/agents/test_orchestration_flags.py tests/tasks/test_workflow_watchdog.py -q green on CI with defaults OFF
Staging soak≥ 24 h with flag ON; zero unexplained workflow halts
KillSwitch drillDocumented step-up + fan-out result JSON (redact tokens)
Watchdog drillArtificial stalled run halted within 2 * agent_stall_threshold_seconds
Rollback drillFlag OFF restores no-op within one beat period
No money-plane couplingConfirm no live trading / enable_money_plane change in the same change set

Prod enablement is a separate change request after staging sign-off — not part of S8.

7. Operator checklist (copy/paste)​

  • Staging-only target confirmed (not shared demo / prod)
  • workflow_runs migrated
  • Defaults still OFF in checked-in settings.py / .env.example
  • Staging secret/ConfigMap sets kill-propagation true
  • Beat + workers restarted
  • Stall drill + KillSwitch drill recorded
  • Soak window complete
  • Rollback drill recorded
  • Prod left at false