Your first RL experiment
Goal: from blank RLExperimentSpec to a trained PPO agent with
trajectories persisted to Iceberg, in under 10 minutes on CPU.
Why
The RL stack is AlphaSwarm's most opinionated subsystem: hash-locked
RLExperimentSpec, metaclass-registered components, deterministic
Iceberg trajectory persistence, and a single sanctioned executor
(RLRuntime).
Every RL run produces an immutable rl_runs ledger row and a
replayable trajectory.
Prerequisites
- Quickstart completed.
- A small dev dataset under your local Iceberg catalog. The bundled
alphaswarm_bronze_yfinance_dailynamespace works.
Step 1 — author the spec
Create alphaswarm_rl/configs/experiments/my_first_rl.yaml:
name: MyFirstRLExperiment
description: First-RL tutorial — PPO on a static universe
environment:
rl_alias: StockTradingEnv
symbol: { ticker: SPY, exchange: ARCA, kind: equity }
lookback_bars: 60
initial_cash: 100000
data_pipeline:
rl_alias: IcebergRLDataPipeline
namespace: alphaswarm_bronze_yfinance_daily
start: 2022-01-01
end: 2023-12-31
agent:
rl_alias: SB3Adapter
algorithm: PPO
policy: MlpPolicy
total_timesteps: 50000
rewards:
- { rl_alias: PnLTerm, weight: 1.0 }
- { rl_alias: TurnoverPenaltyTerm, weight: 0.1 }
- { rl_alias: VolatilityPenaltyTerm, weight: 0.05 }
observations:
- { rl_alias: TechnicalIndicatorBuilder }
- { rl_alias: LookbackStackBuilder, length: 20 }
training:
advantage: { rl_alias: GAE, lambda: 0.95, gamma: 0.99 }
backbone: { rl_alias: TransformerBackbone, d_model: 64, n_heads: 4 }
The rl_alias values above are the registered component names in
alphaswarm_rl (verified against
alphaswarm_rl/src/alphaswarm_rl/ —
e.g. rewards live under rewards/, observations under observations/,
environments under envs/). Double-check each component's exact
constructor kwargs against its source file before relying on the block
above verbatim — the top-level environment / data_pipeline nesting
shown here has not been fully verified against RLExperimentSpec's
schema in
alphaswarm_rl/spec.py.
Step 2 — snapshot + train
Persist the spec (reading the YAML into the request body), then dispatch
a training run against it — POST /rl/runs with a bare spec_path does
not exist; the real flow is POST /rl/specs (hash-locks the spec, returns
a slug) followed by POST /rl/specs/{slug}/run:
curl -X POST http://localhost:3000/api/rl/specs \
-H "Content-Type: application/json" \
-d "{\"spec\": $(python -c "import json,yaml,sys; print(json.dumps(yaml.safe_load(open('alphaswarm_rl/configs/experiments/my_first_rl.yaml'))))")}"
curl -X POST http://localhost:3000/api/rl/specs/myfirstrlexperiment/run \
-H "Content-Type: application/json" \
-d '{"target":"train"}'
(The slug defaults to a lowercased, hyphenated form of name — see
RLExperimentSpec._ensure_slug — so MyFirstRLExperiment becomes
myfirstrlexperiment unless you set slug: explicitly in the YAML.)
The response includes a task_id. Tail the progress:
docker compose exec alphaswarm-core python -c "from alphaswarm.ws.broker import subscribe; \
[print(m) for m in subscribe('<task_id>')]"
50k timesteps on a CPU finishes in 5-8 minutes.
Step 3 — inspect the ledger + trajectory store
-- rl_runs ledger
SELECT id, experiment_name, status, total_timesteps, mean_reward
FROM rl_runs ORDER BY created_at DESC LIMIT 5;
The trajectory data lives in Iceberg under
alphaswarm_silver_rl_trajectories.<experiment_hash>:
from pyiceberg.catalog import load_catalog
cat = load_catalog("alphaswarm")
tbl = cat.load_table("alphaswarm_silver_rl_trajectories.<hash>")
df = tbl.scan().to_pandas()
print(df[["episode", "step", "reward", "action"]].head(20))
Step 4 — replay
curl -X POST http://localhost:3000/api/rl/runs/<rl_run_id>/replay \
-d '{"start":"2024-01-01","end":"2024-03-31"}'
Same hash-locked spec, new data window, separate rl_runs row.
Step 5 — halt
curl -X POST http://localhost:3000/api/rl/halt-all
Verify
-
rl_experiment_versionsrow with aspec_hash. -
rl_runsrow with non-NULLmean_reward. - Iceberg trajectory table populated.
- Replay produces a different
rl_runsrow but reuses the samerl_experiment_versionsrow (hash-locked!).
What next
- Concept: RL components — add your own reward term, observation builder, or policy backbone.
- Concept: RL Iceberg trajectories — the persistence contract.
- Tutorial: first agent workflow — hand off RL outputs to an autonomous agent loop.