RL Lab — interactive RL builder
Lives at /rl/lab in the AlphaSwarm operator UI
(alphaswarm_client). Combines several surfaces:
| Tab | Purpose | Component |
|---|---|---|
Experiment (/rl/lab) | Compose env + reward + observation + action + agent + ensembler into one RLExperimentSpec on a drag-and-drop canvas, save, train. | RlLabRoute (uses WorkflowEditor + xyflow) |
| Reward / Observation / Agent / Experiment / Backbone / Advantage builders | Pick a registered class for the given rl_kind, fill its kwargs schema, POST a {class, module_path, kwargs} build-spec. | Generic RlBuilder.tsx, parametrised per route (there is no dedicated Environment-builder tab or component). |
| Component library | Browse every registered RL component, filter by tag / source / category. | RlComponentLibrary.tsx |
Routes
| Path | Component |
|---|---|
/rl/lab | RlLabRoute |
/rl/library | RlLibraryRoute (renders RlComponentLibrary) |
/rl/builder/reward | RlRewardBuilderRoute (generic RlBuilder, kind="rl_reward") |
/rl/builder/observation | RlObservationBuilderRoute (generic RlBuilder, kind="rl_observation") |
/rl/builder/agent | RlAgentBuilderRoute (generic RlBuilder, kind="rl_agent") |
/rl/builder/experiment | RlExperimentBuilderRoute (generic RlBuilder, kind="rl_experiment") |
/rl/builder/backbone | RlBackboneBuilderRoute (generic RlBuilder, kind="rl_policy_backbone") |
/rl/builder/advantage | RlAdvantageBuilderRoute (generic RlBuilder, kind="rl_advantage_estimator") |
/rl/runs | RlRunsRoute (RlRunsPage) |
/rl/runs/[id] | RlRunDetailRoute |
/rl/runs/[id]/replay | RlReplayRoute (RlReplayViewer) |
/rl | RlHomeRoute (quick-train, application registry browser). |
/rl/zoo | RlZooRoute — RL agent zoo (RlZooPage). |
There is no /rl/builder/env route or EnvironmentBuilder component; env
composition happens inside the Experiment canvas above. The Experiment
canvas uses the existing
WorkflowEditor +
@xyflow/react stack. The serializer in
alphaswarm_client/src/components/rl/rlSerializer.ts
turns a FlowGraph into an RLExperimentSpec payload by bucketising
nodes via their palette group (env / observation / action / reward /
termination / agent / data pipeline / ensembler).
API surface used
The lab calls the API endpoints in
alphaswarm_rl/src/alphaswarm_rl/api/routes/rl.py:
GET /rl/components— kind counts.GET /rl/components/{kind}— list registered components per kind.POST /rl/lab/preview-reward— reward decomposition.POST /rl/lab/preview-observation— observation shape + features.POST /rl/lab/preview-action— action transform sample.POST /rl/specs— persist a spec.POST /rl/specs/{slug}/run— kick off train / evaluate / paper / replay / walk-forward via the matching Celery task.GET /rl/runs/GET /rl/runs/{id}/.../equity/.../trajectories/.../reward-decomposition/.../episodes/.../actions— runs ledger + step-level data served from DuckDB views over the Iceberg tables.POST /rl/runs/{id}/replay— re-roll a saved policy on a new window.POST /rl/data-pipelines/preview— show first rows + array shapes.
Run replay
The replay viewer (/rl/runs/[id]/replay) loads:
rl.equity_curvesrows for the chosen episode (slider populates from the row count).rl.trajectoriesrows for the chosen episode (each step shows reward + info JSON).
Both come from the DuckDB views generated by
alphaswarm_rl/src/alphaswarm_rl/trajectories/duckdb_views.py.