Skip to main content

Per-tenant MCP rollout runbook

Phase 5 §8 of RESTRUCTURING_PLAN.md. Walks the cluster operator through deploying per-tenant MCP servers, the gVisor agent-sandbox pool, and the Cell-Bound- Authorization gate at alphaswarm-edge.

Scope​

  1. gVisor RuntimeClass — install via the DaemonSet at alphaswarm_platform/deployments/kubernetes/agent-sandbox/gvisor/.
  2. alphaswarm-agent-sandbox-pool — the gVisor-isolated Deployment at alphaswarm_platform/deployments/kubernetes/agent-sandbox/pool/.
  3. Per-tenant MCP servers — Helm-rendered Deployments from alphaswarm_platform/deployments/helm/alphaswarm-mcp-tenant/ for each shared-prem / silo-reg tenant.
  4. Cell-Bound-Authorization — the second ext_authz step in alphaswarm_platform/build/docker/alphaswarm-edge/envoy.template.yaml.
  5. MCP tool catalog versioning — Alembic 0084 creates the mcp_tool_versions table + adds agent_runs_v2.mcp_tool_descriptor_hashes.

Prerequisites​

  1. Phase 3 cells are registered and at least one is in state=active.

  2. Phase 4 SPIRE control plane is healthy in the cell. Verify:

    kubectl -n spire-system rollout status statefulset/spire-server
    kubectl -n spire-system get pods -l app=spire-agent
  3. The 0084_mcp_tool_versioning migration has been applied (Alembic head has since moved well past it — the monolith is at 0150+ as of this writing). Verify:

    alembic history | grep 0084_mcp_tool_versioning
  4. Phase 2 Kyverno policies are loaded. Verify:

    kubectl get clusterpolicy alphaswarm-require-gvisor-for-agent-sandbox

Step 0 — Install gVisor​

kubectl apply -k alphaswarm_platform/deployments/kubernetes/agent-sandbox/gvisor/
kubectl -n gvisor rollout status daemonset/gvisor-installer --timeout=10m

# Wait for the node labels to appear (the installer marks each node
# `alphaswarm.io/gvisor=installed` after patching containerd):
kubectl get nodes -L alphaswarm.io/gvisor
# Expected: every node ends with `installed`.

Step 1 — Deploy the agent-sandbox pool​

kubectl apply -k alphaswarm_platform/deployments/kubernetes/agent-sandbox/pool/
kubectl -n alphaswarm-agent-sandbox rollout status deployment/alphaswarm-agent-sandbox-pool --timeout=5m

# Confirm gVisor is active inside the pod (the kernel reports as `runsc`):
POD=$(kubectl -n alphaswarm-agent-sandbox get pods -l app=alphaswarm-agent-sandbox-pool -o name | head -1)
kubectl -n alphaswarm-agent-sandbox exec "$POD" -- /bin/sh -c "uname -r; cat /proc/version"
# Expected: kernel version reports as runsc/gVisor.

# Confirm the Kyverno gate is enforced — try to deploy a Pod with the
# `alphaswarm.io/sandbox-required` label but WITHOUT runtimeClassName:gvisor:
cat <<EOF | kubectl apply -f - --dry-run=server
apiVersion: v1
kind: Pod
metadata:
name: sandbox-test
namespace: alphaswarm-agent-sandbox
labels:
alphaswarm.io/sandbox-required: "true"
spec:
containers:
- name: x
image: cgr.dev/chainguard/python:3.11
EOF
# Expected: admission rejected with `cedar_denied: alphaswarm-require-gvisor-...`

Step 2 — Deploy a per-tenant MCP server (silo-reg example)​

# Render the Helm chart for the Acme tenant:
helm upgrade --install \
acme-mcp \
alphaswarm_platform/deployments/helm/alphaswarm-mcp-tenant/ \
--namespace cell-silo-reg-acme \
--set cell_id=cell-silo-reg-acme \
--set tenant_id=tenant_acme \
--set tier=silo-reg

# Verify:
kubectl -n cell-silo-reg-acme get deployments -l alphaswarm.io/tenant-id=tenant_acme
# Expected: alphaswarm-data-mcp-tenant_acme + alphaswarm-codebase-mcp-tenant_acme.

Step 3 — Snapshot the tool catalog​

The MCP server snapshots its tool catalog into mcp_tool_versions on boot. Verify:

kubectl -n cell-silo-reg-acme exec \
$(kubectl -n cell-silo-reg-acme get pods -l alphaswarm.io/mcp-kind=data -o name | head -1) \
-- /bin/sh -c "psql \$ALPHASWARM_POSTGRES_DSN -c 'SELECT tool_name, substring(descriptor_hash, 1, 12) FROM mcp_tool_versions LIMIT 10;'"

Expected output:

       tool_name        | substring
------------------------+--------------
data.catalog.browse | abc123def456
data.entities.search | f0e1d2c3b4a5
...

Step 4 — Verify agent_runs_v2 records descriptor hashes​

After running a backtest that invokes MCP tools, the agent_runs_v2.mcp_tool_descriptor_hashes column carries the set of hashes the run saw:

SELECT id, status, mcp_tool_descriptor_hashes
FROM agent_runs_v2
WHERE workspace_id = '<your-workspace>'
ORDER BY started_at DESC
LIMIT 5;

The hash array MUST be a subset of mcp_tool_versions.descriptor_hash at the matching cell_id. The Phase 7 §10.2 replay harness will verify this invariant.

Step 5 — Validate Cell-Bound-Authorization​

Cross-cell MCP calls now require the Cell-Bound-Authorization header. Without it, alphaswarm-edge returns 403 at the second ext_authz step.

# From outside the cluster, simulate a cross-cell call missing CBA:
curl -sS -XPOST https://manage.alpha-swarm.ai/mcp/data/cell-silo-reg-acme/some.tool \
-H 'authorization: Bearer <jwt>' \
-d '{"args": {}}'
# Expected: 403 with `cell_bound_invalid` in the body.

# With a valid CBA (minted by the source-cell tenant-router):
curl -sS -XPOST https://manage.alpha-swarm.ai/mcp/data/cell-silo-reg-acme/some.tool \
-H 'authorization: Bearer <jwt>' \
-H 'Cell-Bound-Authorization: <cba-jwt>' \
-d '{"args": {}}'
# Expected: tool result.

The CBA validator is co-located in the alphaswarm-tenant-router deployment (POST /cell_bound/v1/check); the alphaswarm-cell-bound-validator Service selects those pods, so no second deployment is needed. The failure_mode_allow: false flag on the ext_authz filter means cross-cell calls still fail closed if the validator is ever unreachable — which is the intended behaviour for the security posture.

Rollback​

Each component is independently revertable:

# Per-tenant MCP — uninstall the Helm release:
helm uninstall acme-mcp -n cell-silo-reg-acme

# Agent sandbox pool — scale to zero:
kubectl -n alphaswarm-agent-sandbox scale deployment alphaswarm-agent-sandbox-pool --replicas=0

# gVisor — DO NOT DROP the installer DaemonSet without first
# removing every Pod with `runtimeClassName: gvisor`, otherwise
# the pods will sit in RunPodSandboxFailed forever.

# Cell-Bound-Authorization — flip ext_authz failure_mode_allow to true
# in the envoy ConfigMap then `kubectl rollout restart -n alphaswarm-edge
# deployment/alphaswarm-edge`. Cross-cell calls then bypass the CBA gate.

Phase 5.5 follow-ups​

  1. alphaswarm-cell-bound-validator service — shipped: the POST /cell_bound/v1/check route is co-located in the alphaswarm-tenant-router Starlette app (alphaswarm_tenant_router.cell_bound.CellBoundVerifier), and the alphaswarm-cell-bound-validator Service selects the same pods, so no separate deployment was needed.
  2. shared-std MCP pool chart — the shared-std tier uses one pool per cell with per-tenant Linux cgroups (cgroups v2 + Pod Security Standards restricted). The Helm chart for the pool is a Phase 5.5 deliverable; the per-tenant chart in this PR targets shared-prem and silo-reg.
  3. Biscuit + TokenExchangeBroker wire-up in AgentRuntime — the helpers in alphaswarm/auth/biscuit.py are standalone today; the AgentRuntime integration that mints + attenuates the biscuit per call is Phase 5.5.
  4. MCP tool versioning replay — mcp_tool_descriptor_hashes recording works in Phase 5; the replay harness that verifies the recorded set matches the live catalog is Phase 7 §10.2.