Saltar al contenido principal

Promotion evidence: trial ledger + evidence bundle

In plain English: if you test enough random strategies, one of them will look brilliant by pure luck. AlphaSwarm defends against this by keeping a tamper-resistant ledger of every experiment a researcher (or research agent) runs — not just the flattering ones — and by requiring statistical evidence that accounts for all those attempts before a strategy is allowed anywhere near real money. A strategy that looked great in 1 test out of 1 is very different from one that looked great in 1 test out of 10,000, and the promotion gates now know the difference.

This page covers the three "ARP" (agentic research promotion) slices landed in August 2026. All enforcement is default-off behind feature flags; see Rollout.

Why this exists​

The Promotion Gates API has always required human approval plus deterministic risk/validation gates before a strategy crosses from the research plane to the money plane. What it could not previously verify was the statistical honesty of the research behind the request:

  • Selection bias: deflated Sharpe ratio (DSR) and probability of backtest overfitting (PBO) statistics are only meaningful if the trial count N reflects every attempt actually made. Self-reported N under-counts.
  • Evidence provenance: gate inputs were supplied by the requester, not read from a system of record.
  • Execution bypass: nothing at the paper/live execution boundary re-checked that a promotion had actually been approved.

The three slices​

Slice 1 — TrialLedger​

A unified ledger of research trials in Postgres (trial_families / trial_attempts, migration 0160), seeded from lab sweep stamps. Every parameter sweep, walk-forward split, and re-run lands as an attempt row under a trial family, so the DSR/PBO denominator N can be derived from the ledger instead of self-reported.

Slice 2 — EvidenceBundle in the gate chain​

The lab's EvidenceBundle (DSR, PBO, trial counts, provenance) is wired into the PromotionGateChain. With strict enforcement on, a promotion request fails closed if the bundle is missing or incomplete, and paper execution refuses to start without an ApprovedPromotion token.

Plane hygiene note: promotion code reaches lab evidence through a registered probe port (alphaswarm/promotion/lab_evidence.py → alphaswarm/lab/evidence/promotion_probe.py), never by importing alphaswarm.lab directly — the plane_money_promotion boundary stays intact.

Slice 3 — Bounded graph tools​

Agents assessing a promotion need lineage context ("what feeds this strategy?", "what else breaks if this dataset is stale?") without an open-ended graph query surface. Four bounded, point-in-time-filtered MCP tools expose exactly that:

ToolQuestion it answers
data.graph.temporal_contextWhat did the graph neighbourhood of this entity look like as of time T?
data.graph.dependency_impactWhat downstream artifacts depend on this node?
data.graph.risk_neighborsWhich risk-relevant entities sit near this node?
data.graph.sync_statusHow fresh is the graph projection itself?

There is deliberately no raw Cypher surface — every tool is bounded in depth, filtered point-in-time, and returns typed payloads.

How a promotion flows now​

Rollout and flags​

FlagDefaultEffect when enabled
ALPHASWARM_TRIAL_LEDGER_N_ENFORCEoffDSR/PBO statistics use ledger-derived N. Does not auto-demote existing promotions.
ALPHASWARM_PROMOTION_LAB_EVIDENCE_ENFORCEoffStrict mode fails promotion closed without a complete EvidenceBundle.
ALPHASWARM_GRAPH_TEMPORAL_TOOLS_ENABLEDoffRegisters the four bounded graph tools.

Recommended order: run the re-verdict report, review deltas with the research owners, enable ledger N in staging, then strict evidence enforcement, then the graph tools for the approval-assistant agents.

See also​