ADR 015 — Runtime decomposition: cell-based modular monolith over domain microservices
- Status: Accepted (2026-06-10) — operator approval recorded during the KG-platform execution run (P0 gate review, T01-T03 complete; Track C repositories now exist and are scaffolded)
- Authors: Platform team
- Related: ADR 004, ADR 005, ADR 006, ADR 011, ADR 014, RESTRUCTURING_PLAN.md, repository-split.md
Context
The alphaswarm monolith is the platform's largest deployment unit: one
FastAPI process (alphaswarm-core) carrying ~112 route modules, the
in-process MCP routers (/mcp/data, /mcp/codebase, /mcp/ml), and a
Celery task surface of ~57 task modules, backed by 56 ORM model files
and 89 Alembic migrations over a single Postgres, plus Redis in at least
six distinct roles (broker, metadata cache, progress pub/sub, RAG
vectors, ownership event stream, sandbox namespaces). The Settings
singleton exposes ~637 knobs.
The question on the table: should the hosted platform break this runtime into a distributed domain-microservices architecture (agents-service, backtest-service, ml-service, data-service, each with its own datastore), or adopt a different decomposition?
What is already decomposed
The platform is not a greenfield monolith. Substantial decomposition has shipped or is accepted:
| Seam | State | Where |
|---|---|---|
Control plane (alphaswarm-cp) | Shipped — standalone repo, image, /manage/* API; never imports alphaswarm.* | ADR 005, alphaswarm_controller/ |
| Shared kernel | Shipped — wire types, provider ABCs, WorkloadRuntime | alphaswarm_core/ |
| Compute split | Shipped — alphaswarm-worker (queues default,paper,terraform,ingestion,workflows) vs alphaswarm-executor (backtest,training,ml,agents,factors,rag), independent HPAs on alphaswarm_celery_queue_depth | worker-executor-images.md, alphaswarm_platform/deployments/kubernetes/base/ |
| Frontends | Shipped — alphaswarm_client, alphaswarm_ui, alphaswarm_admin, alphaswarm_ide, all HTTP-only | ADR 011 |
| RL / ML boundary packages | Shipped — alphaswarm_rl, alphaswarm_models with deprecation shims; routers/tasks mounted from the external packages | repository-split.md |
| Bots | Boundary package + kopf operator + per-bot pods | ADR 006, alphaswarm_bots/ |
| Edge / cells | Envoy alphaswarm-edge + alphaswarm-tenant-router (ext_authz, rendezvous cell routing), cell registry in topology.yaml, per-cell overlays + ArgoCD ApplicationSet, alphaswarm-cell-data-plane Helm chart | cell-router-cutover.md, alphaswarm_platform/tenant_router/ |
| KB boundary | Accepted (ADR-014) — alphaswarm_kb + alphaswarm_kb_federation designed; repositories now scaffolded and under active development (Track C repo creation has since happened) | ADR 014 |
Invariants any split must respect
Six hard-rule families make naive per-domain services with per-service databases actively harmful here:
- Single Postgres ledger. Every runtime writes immutable
*_spec_versionssnapshots and*_runsledger rows throughLedgerWriter, which stampsexperiment_id/test_idfromRequestContext(rules 13, 15, 17, 24, 34, 41, 43, 57). Splitting the ledger per service destroys the cross-domain experiment umbrella and the audit/replay story. - Single LLM gateway. All LLM calls go through
router_complete(rule 2); telemetry, cost caps, and the semantic cache depend on it. - Single lakehouse write path. All Iceberg writes go through
iceberg_catalog.append_arrowwith medallion validation (rules 3, 21, 46). - DataMCP boundary. Agents never read Postgres/Iceberg directly (rule 22) — the agent↔data seam is already a service-shaped API.
- Kill-switch fan-out. The topbar kill switch fans out to 12+ halt endpoints with a p99 propagation SLO; every new long-running runtime must join the fan-out, and every process fragment multiplies the propagation surface (rules 40, 45, 52).
- Idempotent cross-task state in Postgres only (rule 5) — Celery workers are already stateless and horizontally scalable; the "scaling" benefit of microservices largely exists today via queues.
Options considered
Option 1 — Classic domain microservices
Carve alphaswarm-core into independently deployed services
(agents-svc, backtest-svc, analysis-svc, data-svc, trading-svc, …),
each owning its own database and API, communicating via REST/gRPC and
an event bus.
- Violates invariant 1 (ledger) and 6 unless every service still writes to the shared Postgres — at which point they are not microservices, just N processes sharing one schema and one Alembic chain (a distributed monolith).
- The hot coupling points (
LedgerWriter,router_complete,append_arrow, metadata cache, progress bus) would become N× network hops with retry/outbox machinery the platform doesn't need. - Kill-switch propagation and hash-locked replay would have to be re-engineered across service boundaries.
- The throughput-bound work (backtests, training, agent runs) is already isolated in the executor fleet with queue-depth autoscaling; a backtest-service would duplicate that with more moving parts.
Option 2 — Cell-based modular monolith with selective service extraction (recommended)
Keep one logical application (alphaswarm runtime) but:
- Scale out by cell, not by domain. A cell = one namespace running
the core/worker/executor/beat quartet against a per-cell data plane
(CNPG Postgres, Redis, MinIO, MLflow, Iceberg REST), routed by the
Envoy edge + tenant router (rendezvous hashing on
tenant_id → cell_id). Tiers map onto the existingTenancyStrategylattice (shared-std→ RLS,shared-prem→ schema-per-tenant,silo-reg→ database-per-enterprise). This is RESTRUCTURING_PLAN Phases 3 + 6, already partially provisioned (cells:registry intopology.yaml, cell overlays, ApplicationSet,alphaswarm-cell-data-planechart). - Extract services only along the seams that already have service-shaped contracts — the hash-locked spec runtimes, the MCP HTTP surfaces, the control plane, and the operator pattern — and only when an extraction passes the Future Repo Split Gate in repository-split.md.
Option 3 — Status quo
Keep the single alphaswarm-core Deployment and scale vertically.
Rejected: noisy-neighbor risk across tenants, blast radius of one bad
deploy is the whole fleet, and silo-reg compliance tenants cannot be
served.
Decision
Adopt Option 2. Decomposition proceeds in three tracks, ordered by risk and by whether new repositories are required.
Track A — Process/deployment splits of the existing images (no new repos)
These change alphaswarm_platform/ manifests and entrypoints only; the
code already supports them:
| # | Cut | Detail |
|---|---|---|
| A1 | alphaswarm-beat as a first-class Deployment | Declared in topology.yaml and Terraform but missing from deployments/kubernetes/base/; promote it (replicas: 1, no HPA). |
| A2 | Standalone MCP server Deployments | Serve /mcp/data, /mcp/codebase, /mcp/ml from dedicated pods using the existing image with a scoped ASGI entrypoint. Topology already declares alphaswarm-ml-mcp as a separate service; RFC 9728/8707 audience binding (rule 49) already gives each MCP its own aud. Per-tenant MCP isolation then reuses the alphaswarm-mcp-tenant Helm chart (Phase 5). |
| A3 | Per-queue executor fleets | Split the executor Deployment into per-queue ScaledObjects (KEDA) for backtest, training/ml, agents/rag so GPU-class and CPU-class work scale independently. No code change — queue routing exists in celery_app.py. |
| A4 | paper-trader and ingester-* as first-class K8s units | They exist as compose targets (paper, ingester image stages); give them base manifests + HPAs like worker/executor. |
| A5 | Cell rollout | Execute Phase 3 (cell registry + router live, RequestContext.cell_id propagating) then Phase 6 (per-cell Postgres/MinIO/Redis/MLflow via dual-write migration, ALPHASWARM_CELL_DUAL_WRITE). |
Track B — Deepen existing extractions (existing repos, invasive code changes)
| # | Cut | Detail |
|---|---|---|
| B1 | Sidecar control plane as hosted default | Flip hosted deployments from ALPHASWARM_MANAGEMENT_MODE=embedded to sidecar; alphaswarm-cp already ships standalone. |
| B2 | Ledger/telemetry broker for alphaswarm_rl + alphaswarm_models | Today the extracted packages still import the monolith for LedgerWriter, iceberg_catalog, _progress.emit, ORM. Introduce a narrow HTTP/MCP ledger-write surface (mirroring the controller's HttpAuditSink → /_internal/audit/terraform-runs pattern) so RL/ML workers can run from their own images without importing monolith ORM. This is the gating work for ever running them as separate services. |
| B3 | Bots operator fleet | Continue the ADR 006 path: per-bot pods via quantbot-bot chart, latency-class scheduling (ADR 007), canary PnL gates (ADR 010). Bots are the one domain where per-workload processes are genuinely required (HFT node tiers). |
| B4 | CI boundary gates | Extend the rg-based forbidden-import gates from 2 to all 14+ subprojects (RESTRUCTURING_PLAN §2.1, §4.2) so extracted boundaries cannot silently re-couple. Prerequisite for everything above. |
Track C — Extractions that require new repositories (permission gate)
Per the workspace's repo-per-boundary convention, these need new git repositories and therefore explicit approval before any work begins:
| # | Candidate repo | Justification | Status |
|---|---|---|---|
| C1 | alphaswarm_kb | ADR-014 (accepted) defines the KB boundary package — KBRuntime, hash-locked KBCorpusSpec, adapter trinity. Monolith already mounts its router conditionally and migration 0088_alphaswarm_kb_specs shipped. | Repo created and scaffolded (2026-06-10 onward) |
| C2 | alphaswarm_kb_federation | ADR-014's cross-silo federation gateway — standalone FastAPI, never imports alphaswarm.*; deployable today via compose/docker-compose.kb.yml patterns + Terragrunt silo modules. | Repo created and scaffolded (2026-06-10 onward) |
| C3 | alphaswarm_data | A future data-plane service (ingestion, discovery, catalog) is the largest-blast-radius extraction; the RESTRUCTURING_PLAN sequences it last, after cells and per-tenant object storage. | Repo created (2026-06-19 onward) and under active development, ahead of this ADR's original deferral |
Existing placeholder repos alphaswarm_research ("Services for
Research Plane") and alphaswarm_learning are available landing zones
should the research/learning planes later split; no work is proposed
for them in this ADR.
Target topology
Consequences
Positive
- Tenant isolation, blast-radius reduction, and independent scaling are achieved by cells + queues — the actual goals usually cited for microservices — without breaking the ledger, replay, kill-switch, or hash-lock invariants.
- Every extraction reuses a contract that already exists (spec
runtimes, MCP audiences,
/manage/*, operator CRDs), so no new RPC framework or saga/outbox machinery is invented. Linkerd arrives in Phase 4 for cell mTLS, not for inter-domain RPC. - Track A is pure deployment work and reversible per unit.
Negative / risks
- Per-cell data planes multiply infra cost (mitigated by tiering:
shared backplane for
shared-std/shared-prem). - Track B2's ledger broker adds an HTTP hop to RL/ML run bookkeeping; it must remain async/buffered to keep training loops unaffected.
- The dual-write migration window (Phase 6) is the single riskiest
operation; the rollback path is the
ALPHASWARM_CELL_DUAL_WRITEflag. - Track C repos (
alphaswarm_kb,alphaswarm_kb_federation) have since been created and are under active development; KB code paths in the monolith remain conditional pending full integration.
Explicitly rejected
- Per-domain services with per-service databases (Option 1).
- Extracting
router_completeinto an LLM-gateway service. - Splitting the Postgres ledger or the Alembic chain per service.
- Moving Celery beat scheduling out of the single scheduler.
Rollout order
- Track B4 (CI boundary gates) and Track A1–A2 (beat + MCP pods).
- Track A3–A4 (queue fleets, paper/ingester units) and B1 (sidecar CP).
- Track A5 cells: Phase 3 (router + registry), then Phase 6 (data plane), per the stop conditions in RESTRUCTURING_PLAN §18.2.
- Track B2 (ledger broker) — gate for any future out-of-monolith RL/ML workers.
- Track C — only after repository approval.