Saltar al contenido principal

ADR 016 — Pluggable hybrid local execution layer

Context​

Regulatory, data-gravity, and privacy requirements demand that some AlphaSwarm workloads run on the customer's own machine while the control plane stays in the cloud — running ingestion next to a private dataset, executing agent-emitted strategy code without uploading it, or using a customer's GPU. An external research report ("Architectural Paradigms for Pluggable Hybrid Cloud Execution") plus independent 2025–2026 research recommend a five-part blueprint: outbound reverse-tunnel connectivity, a node abstraction, tiered execution isolation (gVisor/Firecracker), per-tenant identity/isolation, and lease/heartbeat/reaper + idempotency fault tolerance.

A repo-wide audit (the companion design doc) shows AlphaSwarm has already built or scaffolded four of those five:

  • An outbound, NAT-traversing device reverse-tunnel (/tunnel/agent, tunnel_registry) that a local connector (alphaswarm-cli connect up) dials out to, authenticated by a distinct device_credential token — described in code as being "for reverse-tunnel agents" — but today it only relays HTTP to a user's backend; it does not carry workloads.
  • The InfrastructureProvider ABC + WorkloadRuntime seam (rule 45 / ADR 004) with a self-registering metaclass + registry.
  • The alphaswarm_worker Executor/ExecutorRouter runtime with a typed WorkRequest/WorkResult contract, idempotency markers, retry/DLQ, and a money-plane risk gate — but it runs in-process on the host with no sandbox.
  • Per-tenant tenancy + cells: TenantNamespaceSpec/TenantQuotas, TenancyStrategy.Hybrid, the Cell/CellState Deployment-Stamp lifecycle, a SPIFFE backend scaffold (spire_trust_domain = alphaswarm.fund), and a per-tenant token-bucket limiter.

What is missing is the integration glue plus net-new remote-node fault tolerance and isolation: there is no local provider, no typed work channel on the tunnel, no node registry / lease / heartbeat / reaper (run-ledger tables have no owner/lease/attempts columns; stall detection is a wall-clock halt with no requeue and there is zero SELECT … FOR UPDATE SKIP LOCKED), no enforced idempotency key, and no execution isolation on the worker.

Two design forks are load-bearing. (1) The Solution Report recommends mapping each machine to a Virtual Kubelet "virtual node" — but that abstraction only makes sense when kube-scheduler is your control plane; AlphaSwarm's scheduler is Celery routing + WorkloadRuntime + InfrastructureProvider, and Virtual Kubelet remains CNCF Sandbox after 7+ years. (2) The Report recommends gRPC reverse tunnels — but ADR 005 already rejected gRPC ("HTTP/JSON is already understood … until hundreds of req/s") and a working WebSocket tunnel already exists.

Decision​

Adopt a pluggable, customer-hosted Local Execution Node that the cloud control plane drives through a new local InfrastructureProvider — not a parallel control path. Specifically:

  1. ProviderKind.LOCAL + LocalProvider(InfrastructureProvider) in alphaswarm_controller, self-registered via the existing metaclass. All customer-hosted workload ops go through WorkloadRuntime (rule 45): audit row written before dispatch, filter_resources on lists, kill-switch fan-out and telemetry inherited unchanged.
  2. Evolve the existing WebSocket tunnel (do not add gRPC): add a work_* frame family (work_register/dispatch/accept/heartbeat/progress/ result/cancel/usage) alongside the HTTP-relay frames, reusing the socket, stream_id demux, and device_credential. Every customer machine stays outbound-only and never connects to the Celery/Redis broker. gRPC and QUIC are explicitly deferred behind a documented scale trigger.
  3. Reject Virtual Kubelet. Model the machine as a registered, single- tenant ExecutionNode (a dedicated silo / Azure Deployment-Stamp / cell); add an execution_nodes registry and a first-class execution_target dispatch dimension. The controller (not kube-scheduler) matches a run's ResourceSpec to a node's advertised capacity; capacity rejection lets the control plane reschedule (the GitLab-KAS ResourceExhausted pattern).
  4. Tiered execution isolation on the node — process → container → gvisor → microvm (Firecracker/Kata) — selected by an isolation_tier field on the work contract. Agent-/LLM-generated code runs in a microVM (T3), mandatory in multi-tenant mode; the launcher fails closed if a requested tier is unavailable. This realizes the gVisor/Kata/Firecracker + Kyverno require-runtime-class intent already named in RESTRUCTURING_PLAN.md §8.3.
  5. Identity: keep the device-credential pairing flow as the bootstrap; add short-lived, single-run scoped per-work tokens (GitHub-Actions-runner model) and a per-node SPIFFE SVID via the existing scaffold; prefer short-TTL + connect-time denylist over revocation lists; fix the worker's plaintext-token-at-rest gap (OS keyring, rule 53). Identity stays in the IAM hub (rule 27); credentials via CredentialResolver (rule 26).
  6. Fault tolerance: add owner_node_id / lease_expires_at / last_heartbeat_at / attempts / idempotency_key (UNIQUE) to the run-ledger layer (a dedicated node_run_claims table), claimed with SELECT … FOR UPDATE SKIP LOCKED, renewed by work_heartbeat, and reclaimed by a reaper beat task — modelled on the existing OrderOutboxRow transactional-outbox pattern. At-least-once + enforced idempotency; no literal exactly-once (Two Generals).
  7. Multi-tenancy & metering: bind each node to one org_id + Cell; propagate tenant context via RequestContext/claims and strip baggage at the trust boundary; enforce concurrency/rate limits (not CPU quotas) via the existing token-bucket limiter. Meter the control plane / orchestration — never the customer's raw compute (the GitHub/HCP/Datadog/Temporal precedent); the node emits idempotent, event-timestamped, tenant_id-tagged work_usage events that roll up into BillingSummary.

The implementation lands in the alphaswarm_local repo (node agent — no longer empty; it now carries a CLI, tunnel daemon, bootstrap/compose/k8s scaffolding, and a staged/inert exec/laptop_runner.py, though the controller-side LocalProvider and tunnel work_* frame family this ADR proposes are not yet built), alphaswarm_controller (provider + tunnel channel), alphaswarm_core (shared contracts), alphaswarm (registry + ledger + reaper), alphaswarm_worker (isolation launcher + keyring), alphaswarm_auth (per-work tokens + SVID), and alphaswarm_admin (cell binding + metering) — respecting the ADR 005 decision tree and all import boundaries.

Consequences​

Positive

  • Customer-hosted execution is a provider behind rule 45, inheriting audit, filter_resources, kill-switch, and telemetry — no new control path to secure.
  • Reuses the existing tunnel, device credential, worker Executor runtime, cells, tenancy quotas, and SPIFFE scaffold — roughly 60–70 % of the work is already done; the build is mostly integration plus a focused set of net-new primitives.
  • The customer machine is outbound-only and never sees the broker — closing the largest multi-tenant exposure of the current device.{id} queue path.
  • A microVM isolation tier finally contains untrusted/agent-emitted code on customer hardware — the platform's single largest execution-security gap.
  • The new lease/heartbeat/reaper + enforced idempotency benefits the whole platform, not just local execution (today there is no requeue-on-stall and no SKIP LOCKED).

Negative

  • Net-new surface to own: a node agent, a tunnel work channel, a node registry, a reaper, and isolation runtimes. Mitigated by phasing (walking skeleton → fault tolerance → isolation → multi-tenancy → GA) with hard exit criteria.
  • Evolving the WebSocket tunnel rather than adopting gRPC is a deliberate short-term/scale trade-off; revisit at sustained internal req/s.
  • Two execution tiers (co-located broker worker vs customer-hosted node) add operator-facing conceptual surface; documented explicitly.
  • Running isolation runtimes (KVM for microVMs) imposes a host dependency; mitigated by advertising available tiers at registration and refusing code-gen work on gVisor-only nodes.

Alternatives considered​

  • Virtual Kubelet / OpenYurt / KubeEdge — rejected. They assume Kubernetes is the scheduling control plane; AlphaSwarm's is not. They would force every local job into a Pod manifest and run an edge-K8s stack on a laptop. Retained only as references; OpenYurt's YurtHub edge-autonomy idea may inform a future disconnected-cell ADR.
  • Expose the Celery/Redis broker to the customer machine (extend the device.{id} queue path off-network) — rejected. Putting a semi-trusted, NAT'd, multi-tenant-adjacent machine on the shared broker is a security and isolation hazard. Tier B replaces broker draining with scoped tunnel dispatch.
  • gRPC-over-HTTP2 reverse tunnel now — deferred (consistent with ADR 005). The WebSocket tunnel already multiplexes and traverses NAT; gRPC/QUIC are documented future options.
  • Run untrusted code in plain containers only — rejected for multi-tenant Tier B. A shared host kernel is not a sufficient boundary for adversarial / LLM-generated code; microVMs (or at least gVisor) are required.
  • Meter customer compute (CPU/GB-hr) — rejected. There is no cost basis on hardware we don't own, and every comparable vendor bills orchestration / managed units instead.

Implementation references​

  • Full design & phased plan: design/hybrid-local-execution-layer.md
  • Provider seam: alphaswarm_core/src/alphaswarm_core/providers/protocol.py, runtime/workload.py; new alphaswarm_controller/src/alphaswarm_controller/providers/local.py
  • Tunnel: alphaswarm_controller/src/alphaswarm_controller/api/routers/tunnel.py, services/tunnel_registry.py
  • Worker runtime + isolation: alphaswarm_worker/src/alphaswarm_worker/execution/
  • Node agent (new): alphaswarm_local/
  • Identity: alphaswarm_auth/src/alphaswarm_auth/auth/device_service.py; controller spire_backend/spire_trust_domain
  • Ledger + reaper: alphaswarm/persistence/ (OrderOutboxRow pattern in models_orders.py + trading/outbox_relay.py); alphaswarm/tasks/ beat schedule
  • Tenancy + metering: alphaswarm_core/src/alphaswarm_core/models/tenancy.py, topology/models.py (cells); alphaswarm_admin/.../accounts/billing.py, api/routers/tenants.py
  • AGENTS rules: 45 (InfrastructureProvider), 26 (CredentialResolver), 27 (IAM hub), 49 (MCP), 51 (TenancyStrategy), 52 (step-up MFA), 53 (device grant) — AGENTS.md