Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 9 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,10 +76,11 @@ AdaptiveRL is an educational reinforcement-learning project in which a PPO agent
- **Deterministic Evaluation**: Reusable evaluation pipeline with reproducible seed control.
- **Untrained Random Policy Baseline**: Built-in non-learning baseline to scientifically validate policy improvement.
- **Obstacle-Density Experiment**: Controlled testing across 4, 6, and 8 obstacles to demonstrate environmental difficulty scaling.
- **Command-Line Interface (CLI)**: Typer-based CLI for training, evaluation, environment inspection, and trajectory demonstration.
- **PPO Learning-Curve Benchmark**: Train fresh PPO models across configurable timestep budgets and export evaluation metrics as JSON/CSV with optional plots.
- **Command-Line Interface (CLI)**: Typer-based CLI for training, evaluation, learning-curve benchmarking, environment inspection, and trajectory demonstration.
- **Streamlit + Plotly 3D GUI**: Interactive browser-based presentation flight deck with a trajectory playback scrubber and live sensor visualization.
- **Training Checkpoints**: Automatic model weight checkpointing (`.zip`) and JSON metadata export.
- **Automated Test Suite**: 49 unit and integration tests verifying kinematics, environment spaces, training lifecycle, and GUI charts.
- **Automated Test Suite**: Unit and integration tests verify kinematics, environment spaces, training lifecycle, benchmark outputs, and GUI charts.

---

Expand Down Expand Up @@ -353,7 +354,7 @@ For detailed per-test execution traces:
python -m pytest -v
```

The repository includes **49 automated unit and integration tests** verifying:
The repository includes automated unit and integration tests verifying:
- 3D kinematics equations and aerodynamic drag
- 29-dimensional observation space bounds
- Analytical 16-ray LiDAR raycasts and obstacle clearance
Expand Down Expand Up @@ -396,7 +397,7 @@ When a training run completes, artifacts are automatically written to disk:
*(e.g., `artifacts/models/drone_ppo_demo_final.zip`)*
- **Training Metadata & Loss/Reward Log**:
`artifacts/metadata/{experiment_name}_training.json`
*(contains total timesteps, duration in seconds, mean reward, and per-episode return lists)*
*(contains total timesteps, `training_time_seconds` for PPO optimization only, broader `duration_seconds` through model serialization, mean reward, and per-episode return lists)*
- **Periodic Checkpoints** (if configured):
`artifacts/checkpoints/{experiment_name}/`

Expand Down Expand Up @@ -677,6 +678,8 @@ ARL/
│ ├── __init__.py # Package version declaration
│ ├── cli.py # Typer CLI implementation
│ ├── config.py # Pydantic configuration schemas and YAML loader
│ ├── benchmarking/
│ │ └── learning_curve.py # PPO budget sweep, evaluation, and JSON/CSV/plot exports
│ ├── algorithms/
│ │ ├── base.py # BaseAlgorithm abstract interface
│ │ ├── ppo.py # Stable-Baselines3 PPO wrapper
Expand All @@ -694,12 +697,13 @@ ARL/
│ └── training/
│ ├── callbacks.py # Episode metric logging and checkpoint callbacks
│ └── trainer.py # PPOTrainer training orchestrator
└── tests/ # 49 automated unit and integration tests
└── tests/ # Automated unit and integration tests
├── test_cli.py
├── test_configuration.py
├── test_drone.py
├── test_evaluation.py
├── test_gui.py
├── test_learning_curve_benchmark.py
└── test_training.py
```

Expand Down
93 changes: 90 additions & 3 deletions docs/EXPERIMENT.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,94 @@ This document details the experimental methodology, hypotheses, benchmark variab

---

## 3. Expected Results (Hypothesized Prior to Testing)
## 3. PPO Learning-Curve Benchmark

The budget benchmark trains a fresh PPO model from the same base configuration at each requested training budget. Every model is evaluated with the same ordered evaluation seed groups, episode count per seed, environment parameters, algorithm settings, and deterministic-action setting; evaluation uses the saved model and a separate fresh environment. `--eval-seeds` selects the seed groups; `--episodes` is the number of episodes run within each group and does not determine how many seeds are evaluated.

```bash
adaptive-rl benchmark budgets \
--config configs/drone_ppo.yaml \
--budgets 5000,10000,25000,50000 \
--training-seed 42 \
--eval-seeds 42,43,44,45,46 \
--episodes 20 \
--deterministic
```

The command reports the budget list and output locations when complete. By default, machine-readable artifacts are written beneath `artifacts/benchmarks/`:

```text
Learning Curve Benchmark
PPO learning-curve benchmark complete
Budgets: 5,000, 10,000, 25,000, 50,000
Training seed: 42
Evaluation seeds: [42, 43, 44, 45, 46]
JSON: artifacts/benchmarks/learning_curve_budget.json
CSV: artifacts/benchmarks/learning_curve_budget.csv
Plot: not generated

artifacts/benchmarks/
├── learning_curve_budget.json
├── learning_curve_budget.csv
└── learning_curve/
├── budget_5000/models/ppo_budget_5000_final.zip
├── budget_10000/models/ppo_budget_10000_final.zip
└── ...
```

JSON contains benchmark settings, one result object per requested budget, pooled metrics, per-seed summaries, cross-seed Student's t statistics, and plot-ready series. CSV contains the pooled per-budget performance values. Pass `--plot` to additionally render `learning_curve_budget.png`; Matplotlib is imported only when plotting is requested and must be installed for that optional output. Generated plot figures are closed after saving.

The built-in benchmark defaults are budgets `[5000, 10000, 25000, 50000]`, training seed `42`, evaluation seed groups `[42, 43, 44, 45, 46]`, and `20` episodes per seed. A `benchmark` section in the YAML supplies these values instead; explicit CLI options override the corresponding config values. Legacy `evaluation.eval_episodes` does not control the number of seed groups or the benchmark episode count. Thus, without overrides, the default evaluation runs five seed groups with twenty episodes each, not twenty seed groups with twenty episodes each.

`budget_timesteps` records the requested budget, while `trained_timesteps` records the actual environment interactions reported by Stable-Baselines3. For example, budget `65` with PPO `n_steps: 64` trains to `128` steps because PPO collects complete rollouts. Compare results using `trained_timesteps` when budgets are not aligned to rollout sizes.

`training_time_seconds` measures only the call to `PPOAlgorithm.train()` using a monotonic clock. It excludes environment/model setup, final model serialization, metadata writing, evaluation, JSON/CSV export, and plotting. Training metadata also retains the broader legacy `duration_seconds` lifecycle measure, which is not the benchmark training-time metric. Neither duration is hardware-independent.

The named benchmark metrics (`success_rate`, `collision_rate`, `timeout_rate`, `mean_reward`, `std_reward`, and `mean_episode_length`) are pooled descriptive summaries over all evaluated episodes for a budget. Reward standard deviation is the sample standard deviation across pooled episode returns and is unavailable (`null` in JSON, blank in CSV) with fewer than two episodes. Success and collision rates use episodes that reported the corresponding outcome field; timeout rate is based only on Gymnasium's actual `truncated` signal. The JSON additionally retains per-seed summaries and cross-seed Student's t statistics from the reusable evaluator; these are distinct from the pooled metrics and are not estimates based on the pooled episode sample. Within-seed reward and episode-length standard deviations follow the evaluator's existing population-standard-deviation convention; cross-seed uncertainty is then calculated over those seed summaries using sample-standard-deviation and Student's t conventions.

Interpret the curves jointly: rising success rate and mean reward with a falling collision or timeout rate suggest improvement; flat metrics may indicate a plateau. A timeout is counted only when Gymnasium returns `truncated=True`, not merely because an episode has a particular length. The same seed groups and settings make evaluation conditions comparable, but do not remove variation from training or guarantee bit-for-bit results across hardware, PyTorch versions, or CUDA kernels.

For a CI-sized run, copy the experiment YAML and set PPO `n_steps: 64` and `batch_size: 32` in that copy. Then run a short evaluation:

```bash
cp configs/drone_ppo_demo.yaml /tmp/drone_ppo_ci.yaml
# Edit /tmp/drone_ppo_ci.yaml: set n_steps to 64 and batch_size to 32.
adaptive-rl benchmark budgets --config /tmp/drone_ppo_ci.yaml --budgets 64,128 --episodes 1
```

The committed demo config uses `n_steps: 1024`, so those tiny budgets would be rounded up to its rollout boundary; keep the shipped training hyperparameters unchanged and use the copied config only for this CI-sized run.

---

## 4. Multi-Seed Evaluation and Confidence Intervals

Evaluation over several independent environment seeds helps show how policy performance varies with randomized starts and obstacles, instead of depending on one seed sequence. `--episodes` is the number of episodes run for each listed seed. Each requested seed owns a disjoint block of actual environment reset seeds (`seed * episodes_per_seed + episode_index`), avoiding overlap between adjacent requested seed groups; the requested seed and actual per-episode reset seed are both recorded. Duplicate requested seeds are rejected to avoid overweighting a repeated condition.

```bash
adaptive-rl evaluate \
--config configs/drone_ppo.yaml \
--model artifacts/models/drone_ppo_final.zip \
--seeds 0 1 2 3 4 \
--episodes 10 \
--deterministic
```

The existing invocation remains single-seed and uses the configuration seed unless overridden with `--seed`:

```bash
adaptive-rl evaluate --config configs/drone_ppo.yaml --episodes 10
adaptive-rl evaluate --config configs/drone_ppo.yaml --seed 7 --episodes 10
```

`--seed` and `--seeds` are mutually exclusive. In multi-seed mode, `--episodes` is per seed, and `--compare-random` is not supported. The command writes `artifacts/evaluation_multiseed.json` and `artifacts/evaluation_multiseed.csv` by default; `--output-report` and `--output-csv` can select alternate destinations.

The JSON retains raw episode records (requested seed, episode index, actual reset seed, return, episode length, outcomes, truncation, and path length when the environment reports positions), per-seed summaries, aggregate metrics, and evaluation metadata. The CSV is a stable, aggregate-only table with one row per metric and columns `metric`, `mean`, `std`, `ci95_lower`, `ci95_upper`, `sample_count`, `seed_count`, `episodes_per_seed`, and `total_episodes`.

Cross-seed means and confidence intervals are calculated from the per-seed summaries, not pooled episodes. The standard deviation is the sample standard deviation (`ddof=1`); two-sided 95% confidence intervals use Student's t critical values and `mean ± t * s / sqrt(n)`. Missing values are excluded per metric. With fewer than two valid seeds, sample standard deviation and CI bounds are `null`/unavailable; they are not replaced with zero. The interval describes uncertainty in the estimated mean across the evaluated seeds under the independent, representative-seed and approximate t-model assumptions. It is not proof that one policy is superior. Identical seeds and deterministic actions reproduce equivalent episode results when the policy and environment implementation are unchanged.

---

## 5. Expected Results (Hypothesized Prior to Testing)

1. **Random Action Baseline**:
- Success Rate: $0.0\%$ (probability of randomly stumbling into a $1.5\text{ m}$ sphere across a $13,500\text{ m}^3$ arena without striking walls is practically zero).
Expand All @@ -50,7 +137,7 @@ This document details the experimental methodology, hypotheses, benchmark variab

---

## 4. Actual Measured Results (Empirical Verification)
## 6. Actual Measured Results (Empirical Verification)

All results below were generated through genuine Python 3.12 CPU execution using the canonical project commands:
```bash
Expand Down Expand Up @@ -87,7 +174,7 @@ adaptive-rl experiment-density --model artifacts/models/drone_ppo_demo_final.zip

---

## 5. Reproducibility Guarantee
## 7. Reproducibility Guarantee

To independently reproduce the identical metrics on any student laptop:
```bash
Expand Down
26 changes: 26 additions & 0 deletions src/adaptive_rl/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,11 @@
a simulated 3D drone through obstacles toward a target waypoint.
"""

from typing import Any

from adaptive_rl.config import (
AlgorithmConfig,
BenchmarkConfig,
ConfigError,
EnvironmentConfig,
EvaluationConfig,
Expand All @@ -25,20 +28,43 @@

__version__ = "0.1.0"

_BENCHMARK_EXPORTS = {
"LearningCurveBenchmarkResult",
"LearningCurvePoint",
"plot_learning_curve",
"run_learning_curve_benchmark",
"validate_budgets",
}


def __getattr__(name: str) -> Any:
if name in _BENCHMARK_EXPORTS:
from adaptive_rl import benchmarking

return getattr(benchmarking, name)
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")


__all__ = [
"__version__",
"AlgorithmConfig",
"BenchmarkConfig",
"ConfigError",
"DefaultOutcomePolicy",
"EnvironmentConfig",
"EpisodeMetrics",
"EpisodeMetricsAccumulator",
"EvaluationConfig",
"ExperimentConfig",
"LearningCurveBenchmarkResult",
"LearningCurvePoint",
"OutcomePolicy",
"TrainingConfig",
"compute_rate",
"extract_episode_metrics",
"load_config",
"plot_learning_curve",
"run_learning_curve_benchmark",
"save_config",
"validate_budgets",
]
17 changes: 17 additions & 0 deletions src/adaptive_rl/benchmarking/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
"""Benchmarking utilities for AdaptiveRL."""

from adaptive_rl.benchmarking.learning_curve import (
LearningCurveBenchmarkResult,
LearningCurvePoint,
plot_learning_curve,
run_learning_curve_benchmark,
validate_budgets,
)

__all__ = [
"LearningCurveBenchmarkResult",
"LearningCurvePoint",
"plot_learning_curve",
"run_learning_curve_benchmark",
"validate_budgets",
]
Loading