Saltar al contenido principal

Portable orchestration rollout and rollback runbook

Use this runbook to deploy the dual-engine control planes and migrate only the representative first wave. Read the canonical guide and ADR 025 before changing ownership.

Stop condition: never continue if a migrated schedule has zero or more than one active owner, context verification fails, engine state cannot be reconciled, or compatibility tests fail with flags off.

Scope​

This runbook covers:

  • system.context_echo on both engines;
  • data.pipeline_manifest_materialization, default Dagster;
  • agents.workflow_runtime, default Prefect; and
  • schedule ownership handoff for those definitions only.

All other Celery and Argo paths remain unchanged.

Required access​

  • cluster or Compose operator access;
  • AlphaSwarm orchestration:read and write scopes;
  • step-up authorization for cancel, retry, and destructive trigger changes;
  • database migration access through the existing migration job;
  • read access to Dagster and Prefect internal UIs for incident diagnosis; and
  • access to MinIO, OTel, and audit evidence through governed operator paths.

Do not connect to Dagster or Prefect databases for application-level reads.

Pre-deployment evidence​

Record the following in the change ticket:

EvidenceRequired value
Reviewed alphaswarm_orchestration SHAExact commit, not a mutable branch
Dagster versions1.13.13 and integrations 0.29.13
Prefect versions3.7.8 and Kubernetes integration 0.7.10
Alembic stateIncludes 0124_portable_orchestration_ledger, 0130_session_distributed_execution, and 0131_session_resource_operations
Legacy task namesPassing with portable flags off
Queue contractagents/workflows mismatch resolved and verified
Existing ownersCelery Beat entries and all native triggers inventoried
Rollback ownerNamed operator with step-up access
Acceptance poolDedicated Prefect pool and worker; no production template mutation
Acceptance imageReviewed image imports the pinned orchestration package without runtime installation
Session providersEnabled CRDs/operators, namespace allowlist, images, resource profiles, and service accounts

Snapshot current active tasks, Celery Beat entries, Dagster schedules/sensors, Prefect deployments/automations, and reconciliation cursors. The snapshot is evidence only; do not copy native database tables.

Validate documentation and deployment artifacts​

From alphaswarm_docs:

pnpm diagrams:validate
pnpm diagrams:check
python3 -m pytest -q scripts/tests/test_orchestration_diagrams.py

From alphaswarm_platform, render the orchestration profile with the base network definition and real secret-file paths:

docker compose \
-f deployments/compose/docker-compose.base.yml \
-f deployments/compose/docker-compose.orchestration.yml \
--profile orchestration-smoke config >/dev/null
kubectl kustomize deployments/kubernetes/mlops/prefect >/dev/null
kubectl kustomize deployments/kubernetes/mlops/dagster >/dev/null

Run the platform contract tests and pinned Dagster chart lint/template gates from the platform PR before deployment.

Apply additive persistence first​

  1. Back up the AlphaSwarm application database through the standard database runbook.
  2. Apply Alembic migration 0124_portable_orchestration_ledger.
  3. Confirm the ten orchestration tables exist.
  4. Confirm forced RLS is enabled on every table.
  5. Confirm app_runtime cannot update or delete orchestration_events.
  6. Run the repository's RLS, out-of-order event, cursor, and concurrent reconciliation tests.
  7. If session resources are enabled, confirm all eight session-execution tables exist, are forced through RLS, and retain append-only events, immutable lease authority, monotonic fences, and database-derived consumer counts.

Do not downgrade the additive migration during a flag rollback. The tables are safe to retain as readable audit history.

Deploy native control planes disabled for customer traffic​

Compose smoke environment​

Provide the five required secret files, then start the smoke profile:

docker compose \
-f deployments/compose/docker-compose.base.yml \
-f deployments/compose/docker-compose.orchestration.yml \
--profile orchestration-smoke up -d
scripts/smoke/orchestration-compose.sh

The smoke script must prove PostgreSQL, Redis, both APIs, the Dagster GraphQL contract, Prefect health, representative registration, a run, cancellation, and restart reconciliation.

Kubernetes​

  1. Confirm External Secrets are ready before applying workloads.
  2. Apply the Prefect Kustomize bundle.
  3. Reconcile the pinned Dagster HelmRelease and values.
  4. Wait for PostgreSQL, Redis, database migration/bootstrap jobs, Prefect API, background service, Prefect workers, Dagster webservers, daemon, and code location.
  5. Run scripts/smoke/orchestration-kubernetes.sh.
  6. Confirm native UI services are cluster-internal only.
  7. Confirm NetworkPolicies, service accounts, probes, resource limits, topology spread, and PodDisruptionBudgets match the reviewed manifests.
  8. Confirm the acceptance flow uses its dedicated work pool and that deleting the acceptance deployment cannot change the production pool template.
  9. Exec the reviewed worker image before running flows and import alphaswarm_orchestration; do not install a wheel into the running pod.
  10. Apply deployments/kubernetes/mlops/session-runtime before enabling provisioned session clusters. Confirm the alphaswarm-session-runtime Role is namespace scoped, cannot read Secrets, grants Spark only its required pod/Service/ConfigMap/PVC lifecycle, and that Ray/Dask pod templates disable service-account token automount.

The default Dagster values keep the existing Dagster database. Use values-postgres16-cutover.yaml only under its separate database migration and rollback procedure.

Verify the gateway before enabling compatibility routing​

Using an authenticated operator session, verify:

GET /orchestration/health
GET /orchestration/workers
GET /orchestration/pools
GET /orchestration/definitions?limit=100
GET /orchestration/registrations?limit=100
GET /orchestration/triggers?limit=100
GET /orchestration/runs?limit=100

Confirm:

  • unauthenticated access is rejected;
  • cross-workspace reads fail even with guessed IDs;
  • write endpoints require write scopes and idempotency keys;
  • cancel/retry requires step-up and emits audit records;
  • an expired or tampered context envelope is rejected;
  • framework tags contain no secret values; and
  • SSE honors Last-Event-ID and produces the canonical progress projection.

For the SSE reconnect check, open a stream with ?after=snapshot-tail, receive at least one newer event, and reconnect with that newer cursor in Last-Event-ID while leaving the old query parameter unchanged. The first replayed event must follow the header cursor: the header always wins.

Register representative definitions​

For each definition:

  1. create or locate the immutable version;
  2. call validation for both engines;
  3. call compilation for both engines;
  4. review every diagnostic;
  5. register the selected engine with its native trigger disabled;
  6. register the alternate engine only when required for parity testing; and
  7. persist engine/native bindings and confirm they rehydrate after an API restart without a duplicate native registration.

Required defaults:

DefinitionSelected defaultAlternate proof
system.context_echoNeither; explicit test selectionExecute once on each engine
data.pipeline_manifest_materializationDagsterPrefect execution against same manifest
agents.workflow_runtimePrefectSuccessful Dagster compilation

Run shadow verification​

Context parity​

Launch system.context_echo on Dagster and Prefect from the same authenticated session and environment. Compare normalized outputs for:

  • user, organization, team, workspace, project, lab, cell, and role;
  • environment and deployment version;
  • session, request, correlation, experiment, and test IDs;
  • context reference and digest; and
  • W3C trace and parent/child span linkage.

The normalized values must match. Native fields may differ and must remain accessible.

Manifest parity​

Run data.pipeline_manifest_materialization on Dagster, then run the same immutable version and manifest contract on Prefect. Compare:

  • parameter validation;
  • portable step ordering;
  • canonical terminal state;
  • required artifact keys;
  • artifact checksums and sizes; and
  • point-in-time DataBinding provenance.

Agent workflow parity​

Run agents.workflow_runtime on Prefect with a bounded test workflow. Confirm the same definition compiles for Dagster. Verify halt/cancel propagation, budget guardrails, session context, and trace continuity.

Cancellation, retry, and restart drills​

For one non-production run on each engine:

  1. launch and wait for RUNNING;
  2. cancel through POST /orchestration/runs/{run_id}/cancel with a new idempotency key and reason;
  3. confirm the native engine enters its cancellation path;
  4. confirm the canonical state becomes CANCELLED once native cancellation is complete;
  5. launch a controlled failure and retry through POST /orchestration/runs/{run_id}/retry;
  6. confirm attempt increments once and the original engine remains retry owner;
  7. restart the AlphaSwarm API/control service;
  8. confirm registrations and runs are rehydrated from persisted bindings;
  9. page events across the restart and verify unique sequences and native event IDs; and
  10. stop the reconciler during a state transition, restore it, and verify idempotent repair from the persisted cursor.

No step may launch the same logical run in both engines.

Verify session-scoped distributed resources​

Run this drill only for provider kinds enabled in the target environment. A missing CRD/operator is a disabled capability, not a reason to apply an unreviewed cluster-wide manifest during acceptance.

Before the lifecycle drill, health-check GET /internal/session-resources/health with the configured machine bearer credential and manage:infrastructure scope. Confirm an absent credential, invalid credential, or missing scope is rejected, the response declares contract version 1, and controller logs never record the credential, raw endpoints, or secret-reference values. The monolith must use its embedded durable manager and dependency-light HTTP provider; it must not import a Kubernetes client or call the native provider in-process.

  1. Create one signed session resource spec with an idle TTL below its maximum TTL and a session-scoped cache namespace.
  2. Acquire it concurrently from two consumers with distinct idempotency keys.
  3. Confirm one native resource, one immutable lease authority, two active consumer rows, and one provider fence.
  4. Wait for READY; resolve the endpoint by its symbolic reference through the authorized resolver. Confirm no URL, token, or credential appears in the spec, labels, events, or API response.
  5. Submit a bound WorkRequest. On Ray or Dask, prove the remote runtime inventory and context digest match before executing domain work. On Spark, prove the same signed request and point-in-time data binding reaches the application boundary.
  6. Renew with the current lease and consumer token, then replay the same idempotency key. Confirm one provider mutation and the same durable outcome.
  7. Attempt renew and release with the previous fence or consumer token. Both must fail closed without mutating the native resource.
  8. Detach the first consumer and confirm the resource remains ready. Detach the final consumer and confirm the provider is called exactly once.
  9. Cancel an acquire while native creation is in flight. Confirm no late READY commit; the provider must observe or create a tombstone and remove the abandoned native resource before terminal cleanup.
  10. Restart the API/coordinator during renew and release operations. Reconcile from durable reservations and events; confirm no duplicate cluster and no stale release.
  11. Exercise the controller's versioned provision, renew, release, and abandoned-cleanup operations through the HTTP provider. Substitute a wrong spec, lease authority, fence, or service credential and confirm each fails closed before the manager commits state.

For Kubernetes providers, additionally confirm RayCluster, DaskCluster, and SparkConnect resources have deterministic names, tenant-safe labels, allowlisted namespaces, pinned images/resource profiles, service accounts, and the expected NetworkPolicy. Verify Ray head and worker placement is compatible with the pinned image architecture; by default both require amd64. Confirm Ray workers have no controller-authored dashboard probes, while KubeRay reports the desired and available worker counts. Verify resourceVersion conflict retries are bounded and provider status responses contain only allowlisted summaries.

Close known cluster caveats​

The acceptance record must explicitly cover these historical failure modes:

  • Prefect cancellation: cancellation is incomplete while the Kubernetes Job continues running. The adapter must enforce Job deletion after the supported Prefect cancellation request and reconcile canonical state to CANCELLED; a flow that merely remains CANCELLING until natural completion fails the drill.
  • Cross-namespace health: validate the engine-local endpoint first, then an explicit cross-namespace probe that matches the deployed NetworkPolicy and cluster CNI behavior. Do not convert a failed probe into a permanent waiver.
  • Platform dependencies: Flux, External Secrets, the expected OTel namespace/collector, and the configured secret backend must be installed and healthy for production-shaped evidence. Ephemeral secrets and non-fatal OTel warnings are local-development evidence only.
  • Acceptance isolation: use a dedicated work pool and source-managed image; never mutate shared-general or rely on ad-hoc package installation.

Verified Julia k3s acceptance — 2026-07-15​

Platform revision 3b8c242 closed the historical caveats above on the [email protected] k3s cluster:

  • Dagster run 0d407566-4708-4177-8666-12b54e546950 reached SUCCESS in a newly created K8sRunLauncher Job. The Job carried alphaswarm.io/acceptance-owner=dagster-acceptance, ran on aqp-tower, and the isolated release contained no daemon. The acceptance instance used SyncInMemoryRunCoordinator, so it never consumed the production queued-run coordinator.
  • Prefect run 53e9b058-3491-489c-abbf-57b293cb400d reached COMPLETED and emitted one native artifact. Its Job used the immutable reviewed image, mounted prefect-kubernetes-smoke-flow at /opt/alphaswarm-smoke, and ran on aqp-tower.
  • Active Prefect run 1e42e77e-08f6-4da8-8a3f-0aeb5749617c was cancelled through the adapter. The canonical/native state reached CANCELLED and Job heavy-bird-4vv7r was absent immediately after enforcement; the test did not wait for natural flow completion.
  • The Prefect observer was explicitly scoped to the workload namespace, which matched its namespace-scoped Role and produced no cluster-scope authorization errors.
  • Flux, External Secrets, the OpenTelemetry gateway, Dagster, and Prefect all passed the post-run health gate. The cross-namespace OTLP probe succeeded, so the kube-router waiver path was not exercised.
  • Cleanup removed the isolated Helm release, worker, work pool, queue, deployment, ConfigMaps, and owned Jobs while preserving dagster-kubernetes-smoke-definitions. Production replicas were restored to Prefect API/background/worker 2/1/2 and Dagster webserver/daemon/code location 2/1/1, with every Deployment available.

Reproduce the profile with alphaswarm_platform/scripts/smoke/orchestration-kubernetes-execution.sh --live; always follow with --cleanup and the post-restore health gate.

Transfer one schedule owner​

Use a transaction or otherwise serialized control path that enforces the partial unique active-owner constraint.

  1. Confirm the selected native trigger exists and is disabled.
  2. Confirm Celery Beat is the sole active owner.
  3. Disable the one matching Celery Beat entry.
  4. Confirm there is temporarily no active producer and no queued duplicate.
  5. Enable the selected Dagster schedule/sensor or Prefect schedule/automation.
  6. Set the normalized trigger binding active_owner=true.
  7. Query all ownership sources and confirm exactly one owner.
  8. Observe through at least one expected fire window.
  9. Confirm one logical run, one native run binding, and one retry owner.

Never enable the native trigger before disabling its Celery entry.

Enable compatibility routing​

Only after representative acceptance passes:

  1. enable the smallest definition-scoped rollout flag;
  2. keep general Celery compatibility enabled;
  3. call the existing /workflows route and stable task name;
  4. verify the response preserves TaskAccepted, task ID, stream URL, replay, and halt semantics; and
  5. verify the normalized run records the selected native engine.

Do not enable a broad default-engine cutover as part of the first wave.

Observation window​

Monitor:

  • /orchestration/health, worker and pool health;
  • normalized active runs by workspace, state, environment, and age;
  • native Dagster and Prefect run state;
  • event ingestion lag and reconciliation cursor age;
  • duplicate logical/native bindings;
  • schedule fire count and ownership;
  • MinIO checksum or write failures;
  • SSE disconnect/replay errors;
  • worker deadline, cancellation, and idempotency failures; and
  • OTel error rate, latency, queue wait, and trace continuity.

Treat the native control plane as state authority and the normalized ledger as a repairable projection. A projection freshness failure does not justify editing native tables.

Flag-first rollback​

Trigger rollback on duplicate scheduling, duplicate retry, context mismatch, unrecoverable reconciliation lag, native API incompatibility, cancellation failure, cross-tenant exposure, or failed compatibility behavior.

  1. Disable new portable registrations and launches.
  2. Disable the affected native trigger.
  3. Confirm active_owner=false for that binding.
  4. Restore the legacy Celery Beat entry.
  5. Confirm Celery is the only active owner.
  6. For active native runs, choose one authorized action:
    • allow safe completion; or
    • cancel through the AlphaSwarm gateway with step-up and audit.
  7. Keep additive ledger rows and native payload references readable.
  8. Keep both native control planes running long enough to reconcile final state.
  9. Re-run compatibility tests with flags off.
  10. Attach event cursors, audit IDs, run IDs, native IDs, and OTel trace IDs to the incident ticket.

Do not delete normalized history, force engine database state, or downgrade the ledger migration as an incident shortcut.

Incident triage matrix​

SymptomFirst checksSafe action
Duplicate scheduled runsCelery Beat entry, native trigger, active_ownerDisable the newest owner, stop launches, rollback ownership
Normalized state is staleNative run, cursor age, adapter healthRestart reconciler and rehydrate bindings
Native run is missingRegistration ID, native binding, audit eventStop retries; reconcile before deciding rollback
Events repeat after restartPersisted adapter cursor and normalized high-water markStop projection, preserve cursor evidence, repair adapter contract
Context rejectedKey ID, expiry, digest, trusted referenceDo not bypass verification; relaunch with a fresh context
Secret appears in metadataTags, summaries, logs, payload externalizationStop launches, rotate exposed secret, preserve audit evidence
Cancel remains pendingNative engine state and worker JobEnforce adapter Job termination, then reconcile through supported APIs
Prefect worker creates no Jobpool/queue, service account, context verificationRepair configuration; do not submit an unverified Job manually
Session manager is unavailableembedded manager health, Postgres/RLS, controller M2M endpointStop provisioning; repair durability or authenticated provider health before retry
Controller rejects lifecycle callservice credential, contract version, spec/lease authority, fenceDo not bypass auth or mutate Kubernetes directly; reconcile the durable reservation
Resource becomes ready after cancellationreservation token, tombstone, native objectFence the late completion and run cleanup before terminal commit
Stale lease mutates a clusterfence version, endpoint authority, resourceVersionReject the operation; reconcile the current lease without deleting the resource

Completion record​

The change ticket must include:

  • exact package and deployment SHAs;
  • definition content hashes and engine registration IDs;
  • logical and native run IDs for every parity test;
  • artifact checksums;
  • context digest and trace IDs, never secret values;
  • cancellation, retry, restart, reconciliation, and rollback results;
  • schedule owner before, during, and after cutover;
  • links to audit events and dashboards; and
  • the operator who approved each destructive action.

Mark the migration complete only when all acceptance criteria in ADR 025 pass and the observation window contains no duplicate schedule, retry, or cross-tenant event.