Skip to main content

Runbook — Local agent-first data-infra ops

Use this runbook when you or an agent need to inspect or lightly operate data infrastructure from a local shell: native Dagster GraphQL health, failed runs, schedules/sensors, local Definitions validation, and portable orchestration health. This is not the hosted Dagster UI, not alphaswarm_ops_console, and not a Terraform or Helm apply path.

Related:

Operator subagent: alphaswarm_orchestration/.cursor/agents/alphaswarm-orchestration-operator.md

Which surface am I on?​

AlphaSwarm exposes three orchestration names that must not be mixed in flags, MCP tool choice, or incident tickets.

NounMeaningSurface
Native DagsterClassic Definitions in a named code locationGET /dagster/*, alphaswarm-cli data infra *, data.dagster.*, python -m dagster ...
Portable orchestrationEngine-owned TaskDefinition / run projectionGET /orchestration/*, data.orchestration.portable_*
Agent workflowsWorkflowRuntime / workflow_runs watchdogdata.orchestration.health, data.orchestration.list_runs (not Dagster)

data.automation.* is Celery beat, not Dagster. Do not route Dagster schedules through it.

Global constraints for this interface:

  • Pin dagster==1.13.13. Local validation uses the official CLI only: python -m dagster definitions validate -m <module>.
  • No dg / dagster-dg-cli / remote dg launch workflows.
  • No Terraform apply from alphaswarm-cli data infra. Cluster mutation stays on TerraformRuntime / WorkloadRuntime (One Gate).
  • No merged code locations — three live identities stay separate (below).

First-slice commands​

Authenticate once, then use native Dagster verbs (GraphQL via monolith HTTP) or portable health. All examples are PowerShell.

alphaswarm-cli auth login

# Native Dagster — which code location is loaded?
alphaswarm-cli data infra status

# Failed-run triage
alphaswarm-cli data infra runs list --status FAILURE
alphaswarm-cli data infra runs show <run_id>

# Schedules and sensors (start/stop need manage:infrastructure)
alphaswarm-cli data infra schedules list
alphaswarm-cli data infra schedules start daily_full_refresh
alphaswarm-cli data infra sensors list

# Local Definitions load check (subprocess; does not hit remote GraphQL)
alphaswarm-cli data infra validate --module alphaswarm.dagster.definitions
alphaswarm-cli data infra validate --location monolith
alphaswarm-cli data infra validate --location platform

# Static catalog (no GraphQL)
alphaswarm-cli data infra locations

# Portable orchestration engine health (not native Dagster GraphQL)
alphaswarm-cli data infra portable-health

# Composite: Dagster status + portable health + control-plane topology hint
alphaswarm-cli data infra health

# Portable pools/workers (not native Dagster pool YAML)
alphaswarm-cli data infra pools
alphaswarm-cli data infra workers

# Human cancel on portable runs (step-up MFA on the API)
alphaswarm-cli data infra cancel <run_id> --reason "operator stop"
CommandHTTP or local
data infra statusGET /dagster/status
data infra runs list [--status] [--limit]GET /dagster/runs
data infra runs show RUN_IDGET /dagster/runs/{id}
data infra schedules listGET /dagster/schedules
data infra schedules start NAMEPOST /dagster/schedules/{name}/start
data infra schedules stop NAMEPOST /dagster/schedules/{name}/stop
data infra sensors listGET /dagster/sensors
data infra sensors start NAMEPOST /dagster/sensors/{name}/start
data infra sensors stop NAMEPOST /dagster/sensors/{name}/stop
data infra validate [--module] [--location]local python -m dagster definitions validate -m
data infra locationsstatic JSON catalog (monolith / platform / bootstrap)
data infra portable-healthGET /orchestration/health
data infra healthcomposite (above + topology hint)
data infra poolsGET /orchestration/pools
data infra workersGET /orchestration/workers
data infra cancel RUN_IDPOST /orchestration/runs/{id}/cancel

For raw kubectl/docker health on Dagster pods, use alphaswarm-cli services (HTTP to control plane) — not data infra.

Validate each live module separately​

Allowed --module values (or --location monolith / --location platform) for data infra validate:

  1. alphaswarm.dagster.definitions — monolith compose / local dev default.
  2. pipelines.dagster_user_code.definitions — platform Helm user-code deployment. May require PYTHONPATH including alphaswarm_platform when run from a checkout that does not install that package.
alphaswarm-cli data infra validate --module alphaswarm.dagster.definitions

# Platform user-code (set PYTHONPATH when the module is not on the default path)
$env:PYTHONPATH = "C:\path\to\alphaswarm_platform"
alphaswarm-cli data infra validate --module pipelines.dagster_user_code.definitions

Not valid locally: bootstrap, bootstrap-user-code, or --module file. The Helm ConfigMap dagster-bootstrap-user-code is cluster-only; the CLI exits 2 with a clear message.

If validate exits non-zero while GraphQL /dagster/status is reachable, check whether Definitions load is blocked by a separate hardening item (for example WorkflowConfig in the 2026-08-13 Dagster expand plan). This runbook does not replace that fix.

Start/stop scopes (manage:infrastructure)​

ActionScopeStep-up
GET /dagster/* reads (status, runs, schedules, sensors, assets)read:infrastructureNo
POST /dagster/schedules|sensors/{name}/start|stopmanage:infrastructureNo
GET /orchestration/* readsorchestration:read + tenancy headersNo
POST /orchestration/runs/{id}/cancelorchestration:executeYes (RFC 9470)

CLI schedule/sensor toggles call the monolith with your bearer token. A 403 means the principal lacks manage:infrastructure.

Data MCP (data.dagster.*)​

Agents use the Data MCP catalog (stdio or /mcp/data), not direct GraphQL.

Reads — scope read:infrastructure, tenancy_posture=shared:

  • data.dagster.status
  • data.dagster.list_runs (limit, optional status)
  • data.dagster.get_run (run_id)
  • data.dagster.list_schedules
  • data.dagster.list_sensors

Writes — scope manage:infrastructure, mutates=True:

  • data.dagster.start_schedule / stop_schedule
  • data.dagster.start_sensor / stop_sensor

When ctx.actor_kind == "agent", writes queue pending_approval via the existing MCP approval helper — they do not fire immediately.

Portable tools remain on data.orchestration.portable_* (compile, launch, runs, pools, workers). Do not confuse them with native Dagster tools.

Cancel = CLI/API + step-up; MCP cancel is non-autonomous​

Humans cancel portable runs with:

alphaswarm-cli data infra cancel <run_id> --reason "incident" --idempotency-key <uuid>

The API requires fresh step-up MFA. On 401 with WWW-Authenticate: Bearer error="insufficient_user_authentication", the CLI prints step-up required; run alphaswarm-cli auth login and exits 2. It does not auto-retry.

Agents must not cancel via MCP: data.orchestration.portable_cancel remains non-autonomous — it refuses with MFA/step-up guidance and never calls POST /orchestration/runs/{id}/cancel. Escalate to a human operator.

Terraform apply/destroy and workload halt paths are unchanged and out of scope for this CLI group.

Pools: default_limit vs invalid named config​

On Dagster 1.13.13:

  • concurrency.pools.default_limit on the instance is the runtime cap for tagged work.
  • concurrency.pools.config.<name> (named pool config blocks) is invalid on this pin — do not add them.
  • Vendor pool= tags on ops/assets are metadata only; they are not a separate runtime limit.

alphaswarm-cli data infra pools reads GET /orchestration/pools (portable projection). Help text on that command states the above. Doc-only 1/2 limits in YAML remain comments until a future pin proves named pool config.

Three code locations — do not merge​

Never invent one workspace.yaml or combined module that loads all three:

IdentityModule / artifactTypical deployment
Monolithalphaswarm.dagster.definitionsCompose dagster dev -m ...
Platform user-codepipelines.dagster_user_code.definitionsHelm pipelines-user-code
BootstrapConfigMap dagster-bootstrap-user-codeHelm bootstrap-user-code

data infra status reports code_location and module_path for the GraphQL instance you are querying — do not assume Helm user-code is the same module as local compose without checking.

The operator UI Data → Infra hub shows a read-only Code locations banner: three chips (monolith / platform / bootstrap) with the chip matching module_path marked active, plus a topology mismatch hint when control-plane services disagree with GraphQL.

data infra health JSON adds active_location and location_warning using the same identity rules as the UI banner.

Sandbox ≠ prod​

Interactive Dagster sandbox sessions (Dagster sandbox concept; AGENTS hard rule 32; /dagster/sandbox/*) are isolated per session (tempdir, Redis namespace, safe env overrides). They are not the deployed GraphQL instance behind /dagster/status.

  • First-slice data infra verbs talk to deployed native Dagster GraphQL.
  • Sandbox experimentation stays on /dagster/sandbox/* and the sandbox UI/API; there are no data infra sandbox-* CLI verbs in this slice.
  • Do not treat sandbox validate results as production health.

When to invoke which agent​

SituationAgent
Dagster defs load failure, schedule/sensor diagnosis, code-location identity, GraphQL vs validate mismatch, sandbox isolation, portable vs native confusionalphaswarm-orchestration-operator — expect Diagnosis / Evidence / Safe next action / Validation command
Live cluster apply (Terraform, Helm, workload scale/restart), step-up-gated infra mutationalphaswarm-management-engine — One Gate only
LangGraph / WorkflowRuntime / agent crew stallsalphaswarm-agentic-stack-expert
Celery beat / Redis progress / queue depthalphaswarm-queue-cache-operator

The orchestration-operator diagnoses and proposes; it does not replace alphaswarm-cli data infra or MCP tools shipped in this plan.

Operator UI (alphaswarm_client)​

Phase 2a ships a read-only Data → Infra hub in the Vite operator UI (alphaswarm_client, route /data/infra). It mirrors the first-slice CLI reads — no schedule/sensor toggles and no portable cancel in this slice (those land in Task 7).

PanelBacking API
Composite healthGET /dagster/status + GET /orchestration/health
Failed runs (default filter)GET /dagster/runs?status=FAILURE
Run detailGET /dagster/runs/{id}
Schedules / sensorsGET /dagster/schedules, GET /dagster/sensors
Portable pools / workersGET /orchestration/pools, GET /orchestration/workers
Code locations bannerstatic catalog + active chip from GraphQL module_path

Local dev (from an alphaswarm_client checkout):

pnpm --dir alphaswarm_client dev --host 0.0.0.0 --port 3001
# Browser: http://localhost:3001/data/infra

Requires a logged-in session with read:infrastructure (and tenancy headers the BFF forwards). Mutations stay on CLI/API until the Phase 2b UI slice.

CI verification​

GitHub Actions runs path-filtered jobs when data-infra surfaces change:

RepoWorkflowFocus
alphaswarm.github/workflows/data-infra-ops-ci.ymlDagster routes + Data MCP ops tools
alphaswarm_cli.github/workflows/data-infra-ops-ci.ymldata infra CLI + smoke

Reproduce locally before opening a PR (PowerShell, from each repo root with dev deps installed):

# Monolith (sibling packages editable per AGENTS.md agent-bootstrap)
python -m pytest -q `
tests/api/test_dagster_routes.py `
tests/data/mcp/test_dagster_ops_tools.py `
tests/data/mcp/test_orchestration_tools.py

# CLI (after pip install -e ../alphaswarm_config ../alphaswarm_core and pip install -e ".[dev]")
python -m pytest -q tests/test_data_infra.py tests/test_cli_smoke.py

These are focused gates — they do not replace full make agent-verify or the monolith test-monolith job.

Infrastructure provisioning (One Gate — not data infra)​

Cluster provisioning, Helm rollouts, and Terraform apply/destroy never go through alphaswarm-cli data infra. Use the control-plane One Gate surface (TerraformRuntime / AGENTS rule 42):

alphaswarm-cli auth login

# Read workspaces and recent provisioning runs (audit ledger)
alphaswarm-cli cp terraform workspaces
alphaswarm-cli cp terraform runs

# Plan (read-only)
alphaswarm-cli cp terraform plan --workspace-id <id> --spec-version-id <uuid>

# Apply / destroy require fresh step-up MFA on the control-plane API
alphaswarm-cli cp terraform apply --workspace-id <id> --spec-version-id <uuid>
alphaswarm-cli cp terraform destroy --workspace-id <id> --spec-version-id <uuid>

See also Incident response (terraform run audit) and Portable orchestration (engine cutover). For live workload start/stop/scale, invoke alphaswarm-management-engine / WorkloadRuntime — not this runbook's data infra verbs.

Platform deployment (three code locations)​

Helm and compose artifacts live in alphaswarm_platform — docs-first only here; no cluster mutation from this runbook.

IdentityModule / artifactPlatform path
Monolithalphaswarm.dagster.definitionscompose / monolith image (local dev)
Platform user-codepipelines.dagster_user_code.definitionsdeployments/kubernetes/mlops/dagster/values-pipelines-user-code.yaml
BootstrapConfigMap dagster-bootstrap-user-codedeployments/kubernetes/mlops/dagster/user-code-bootstrap-configmap.yaml

Canonical Helm operator notes: alphaswarm_platform/deployments/kubernetes/mlops/dagster/README.md (section Three Dagster code locations). Cross-check with alphaswarm-cli data infra locations (static JSON) before editing workspace YAML — identities stay separate; do not merge repos into one workspace.yaml.

What this is not​

  • Not dg or Dagster+ remote workflows.
  • Not a hosted UI or Dagster UI replacement.
  • Not Terraform apply or kubectl exec from data infra.
  • Not WorkflowRuntime health (data.orchestration.health is a different product surface).