Saltar al contenido principal

Your first RL experiment

Goal: from blank RLExperimentSpec to a trained PPO agent with trajectories persisted to Iceberg, in under 10 minutes on CPU.

Why​

The RL stack is AlphaSwarm's most opinionated subsystem: hash-locked RLExperimentSpec, metaclass-registered components, deterministic Iceberg trajectory persistence, and a single sanctioned executor (RLRuntime). Every RL run produces an immutable rl_runs ledger row and a replayable trajectory.

See Concept: RL framework.

Prerequisites​

  • Quickstart completed.
  • A small dev dataset under your local Iceberg catalog. The bundled alphaswarm_bronze_yfinance_daily namespace works.

Step 1 — author the spec​

Create alphaswarm_rl/configs/experiments/my_first_rl.yaml:

name: MyFirstRLExperiment
description: First-RL tutorial — PPO on a static universe
environment:
rl_alias: StockTradingEnv
symbol: { ticker: SPY, exchange: ARCA, kind: equity }
lookback_bars: 60
initial_cash: 100000
data_pipeline:
rl_alias: IcebergRLDataPipeline
namespace: alphaswarm_bronze_yfinance_daily
start: 2022-01-01
end: 2023-12-31
agent:
rl_alias: SB3Adapter
algorithm: PPO
policy: MlpPolicy
total_timesteps: 50000
rewards:
- { rl_alias: PnLTerm, weight: 1.0 }
- { rl_alias: TurnoverPenaltyTerm, weight: 0.1 }
- { rl_alias: VolatilityPenaltyTerm, weight: 0.05 }
observations:
- { rl_alias: TechnicalIndicatorBuilder }
- { rl_alias: LookbackStackBuilder, length: 20 }
training:
advantage: { rl_alias: GAE, lambda: 0.95, gamma: 0.99 }
backbone: { rl_alias: TransformerBackbone, d_model: 64, n_heads: 4 }

The rl_alias values above are the registered component names in alphaswarm_rl (verified against alphaswarm_rl/src/alphaswarm_rl/ — e.g. rewards live under rewards/, observations under observations/, environments under envs/). Double-check each component's exact constructor kwargs against its source file before relying on the block above verbatim — the top-level environment / data_pipeline nesting shown here has not been fully verified against RLExperimentSpec's schema in alphaswarm_rl/spec.py.

Step 2 — snapshot + train​

Persist the spec (reading the YAML into the request body), then dispatch a training run against it — POST /rl/runs with a bare spec_path does not exist; the real flow is POST /rl/specs (hash-locks the spec, returns a slug) followed by POST /rl/specs/{slug}/run:

curl -X POST http://localhost:3000/api/rl/specs \
-H "Content-Type: application/json" \
-d "{\"spec\": $(python -c "import json,yaml,sys; print(json.dumps(yaml.safe_load(open('alphaswarm_rl/configs/experiments/my_first_rl.yaml'))))")}"

curl -X POST http://localhost:3000/api/rl/specs/myfirstrlexperiment/run \
-H "Content-Type: application/json" \
-d '{"target":"train"}'

(The slug defaults to a lowercased, hyphenated form of name — see RLExperimentSpec._ensure_slug — so MyFirstRLExperiment becomes myfirstrlexperiment unless you set slug: explicitly in the YAML.)

The response includes a task_id. Tail the progress:

docker compose exec alphaswarm-core python -c "from alphaswarm.ws.broker import subscribe; \
[print(m) for m in subscribe('<task_id>')]"

50k timesteps on a CPU finishes in 5-8 minutes.

Step 3 — inspect the ledger + trajectory store​

-- rl_runs ledger
SELECT id, experiment_name, status, total_timesteps, mean_reward
FROM rl_runs ORDER BY created_at DESC LIMIT 5;

The trajectory data lives in Iceberg under alphaswarm_silver_rl_trajectories.<experiment_hash>:

from pyiceberg.catalog import load_catalog
cat = load_catalog("alphaswarm")
tbl = cat.load_table("alphaswarm_silver_rl_trajectories.<hash>")
df = tbl.scan().to_pandas()
print(df[["episode", "step", "reward", "action"]].head(20))

Step 4 — replay​

curl -X POST http://localhost:3000/api/rl/runs/<rl_run_id>/replay \
-d '{"start":"2024-01-01","end":"2024-03-31"}'

Same hash-locked spec, new data window, separate rl_runs row.

Step 5 — halt​

curl -X POST http://localhost:3000/api/rl/halt-all

Verify​

  • rl_experiment_versions row with a spec_hash.
  • rl_runs row with non-NULL mean_reward.
  • Iceberg trajectory table populated.
  • Replay produces a different rl_runs row but reuses the same rl_experiment_versions row (hash-locked!).

What next​