Halt-propagation staging enablement (Sprint S8)
Status: PLANNING — do not flip this flag ON in shared / production envs from this ticket. Defaults remain OFF until an operator explicitly enables staging.
Companion tickets: principal plan S8 (alphaswarm-first-class-module-transformation-plan.md), ADR 036 / 019, and the orchestration rollout orchestration-refactor-rollout.md.
1. Flag inventory (verified)
Source of truth: alphaswarm/config/settings.py (env prefix
ALPHASWARM_).
| Settings field | Env var | Default | Role |
|---|---|---|---|
orchestration_kill_propagation_enabled | ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED | False | Primary. Enables the workflow-run stall watchdog scan (scan_for_stalled_workflow_runs). Documented in settings as extending watchdog + KillSwitch fan-out into workflow_runs. |
agent_watchdog_enabled | ALPHASWARM_AGENT_WATCHDOG_ENABLED | True | AND-gate with kill-propagation inside _is_workflow_watchdog_enabled(). |
agent_watchdog_period_seconds | ALPHASWARM_AGENT_WATCHDOG_PERIOD_SECONDS | 60 | Celery beat period for workflow-stall-watchdog. |
agent_stall_threshold_seconds | ALPHASWARM_AGENT_STALL_THRESHOLD_SECONDS | 300 | Age threshold before a running/pending workflow_runs row is halted. |
orchestration_halt_check_timeout_seconds | ALPHASWARM_ORCHESTRATION_HALT_CHECK_TIMEOUT_SECONDS | 1.0 | Per-transition halt-check budget in WorkflowRuntime. |
orchestration_studio_enabled | ALPHASWARM_ORCHESTRATION_STUDIO_ENABLED | False | Gates the entire /workflows/* surface including POST /workflows/halt (503 when off). KillSwitch still calls that path; other halt targets continue. |
risk_kill_switch_key | ALPHASWARM_RISK_KILL_SWITCH_KEY | alphaswarm:kill_switch | Redis key for the portfolio / global kill-switch signal (orthogonal to workflow watchdog). |
.env.example already lists
ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false.
Wiring today (do not change in shared envs)
Hermetic coverage already present (do not weaken defaults):
tests/agents/test_orchestration_flags.py— flag defaults OFFtests/tasks/test_workflow_watchdog.py— no-op when flag OFF; halt helper with fake rowtests/tasks/test_workflow_watchdog.py::test_workflow_watchdog_dual_gate— documents AND-gate wiring (S8 scaffold)
2. Preconditions (staging only)
Before enabling in a dedicated staging namespace / compose overlay:
- Alembic
0046_workflow_versioning(and later workflow migrations) applied;workflow_runstable present. - Celery worker + beat running with
workflow-stall-watchdogregistered (alphaswarm/tasks/celery_app.py). - Decide whether KillSwitch workflow halt must succeed: if yes, also enable
orchestration_studio_enabledin that same staging overlay (still staging-only). - At least one non-prod
WorkflowSpecthat can be started and left idle / stalled for measurement. - Operator can complete step-up MFA (rule 52) for halt endpoints.
- Observability stack reachable (API logs + Celery logs + optional Grafana/Loki). No money-plane / live trading workloads in the staging cell.
Explicit invariant: production and shared demo envs keep
ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false until the
evidence checklist in §6 is signed off.
3. Staging-only enablement steps
Perform only against staging secrets / ConfigMaps / compose overrides.
-
Confirm current value is false:
# metadata only — do not dump full env
docker exec <api> python -c "from alphaswarm.config import settings; print(settings.orchestration_kill_propagation_enabled)" -
Set staging override only:
ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=trueOptionally (if exercising KillSwitch →
/workflows/halt):ALPHASWARM_ORCHESTRATION_STUDIO_ENABLED=true -
Restart API + Celery worker + Celery beat in staging.
-
Start a disposable workflow run; leave it without breadcrumbs past
agent_stall_threshold_secondsor trigger KillSwitch with step-up. -
Verify:
- Watchdog path:
workflow_runs.status='halted',halted=True, Celery task revoked when applicable. - KillSwitch path (studio on):
POST /workflows/haltreturns{ok: true, halted_count: N}; aggregate toast shows workflows in the fan-out summary.
- Watchdog path:
-
Leave flag ON in staging for a soak window (≥ 24 h recommended) while collecting §4 signals.
Do not commit staging overrides into shared platform values files used by prod.
4. Observability signals
| Signal | Where | Pass criterion |
|---|---|---|
| Watchdog task result | Celery result / progress bus | {ok: true, halted_count: …} shape from scan_for_stalled_workflow_runs |
| Halted rows | Postgres workflow_runs | status='halted', error prefixed watchdog: or fan-out reason |
| KillSwitch aggregate | Operator UI toast | Workflows endpoint success or expected 503 if studio still off |
| False positives | Staging audit / ops channel | No unexpected halt of healthy runs during soak |
| Latency | API / beat logs | Halt observed within ~orchestration_halt_check_timeout_seconds after Redis/DB update (runtime next transition) |
| Agent watchdog still healthy | GET /agents/health / data.agents.health | Agent stall scan unaffected |
5. Rollback (staging or accidental shared flip)
- Set
ALPHASWARM_ORCHESTRATION_KILL_PROPAGATION_ENABLED=false. - Restart API + Celery beat (worker reload if settings cached per-process).
- Confirm
_scan_and_halt_workflow_runs()returns[]with flag off (already unit-tested). - Existing
workflow_runshalted rows remain historical — no migration reverse required. - If studio was enabled only for this experiment, restore
ALPHASWARM_ORCHESTRATION_STUDIO_ENABLED=falseindependently.
KillSwitch continues to fan out to agents / paper / bots / rl / quant-agents / assistants / terraform / workloads / lab / ml serving regardless of this flag.
6. Test evidence required before prod
| Gate | Evidence |
|---|---|
| Unit / hermetic | pytest tests/agents/test_orchestration_flags.py tests/tasks/test_workflow_watchdog.py -q green on CI with defaults OFF |
| Staging soak | ≥ 24 h with flag ON; zero unexplained workflow halts |
| KillSwitch drill | Documented step-up + fan-out result JSON (redact tokens) |
| Watchdog drill | Artificial stalled run halted within 2 * agent_stall_threshold_seconds |
| Rollback drill | Flag OFF restores no-op within one beat period |
| No money-plane coupling | Confirm no live trading / enable_money_plane change in the same change set |
Prod enablement is a separate change request after staging sign-off — not part of S8.
7. Operator checklist (copy/paste)
- Staging-only target confirmed (not shared demo / prod)
-
workflow_runsmigrated - Defaults still OFF in checked-in
settings.py/.env.example - Staging secret/ConfigMap sets kill-propagation
true - Beat + workers restarted
- Stall drill + KillSwitch drill recorded
- Soak window complete
- Rollback drill recorded
- Prod left at
false