Implementation of REINFORCE, A2C, and DDPG using PyTorch and Gymnasium for classic control tasks.
This repository contains modular, well-documented implementations of three fundamental policy gradient algorithms:
- REINFORCE (Vanilla + Baseline): Monte Carlo policy gradient
- A2C (Advantage Actor-Critic): Synchronous actor-critic with online learning
- DDPG (Deep Deterministic Policy Gradient): Off-policy for continuous control
| Environment | Action Space | Obs Space | Algorithms |
|---|---|---|---|
| CartPole-v1 | Discrete (2) | Continuous (4D) | REINFORCE, A2C |
| Pendulum-v1 | Continuous [-2, 2] | Continuous (3D) | DDPG |
| Algorithm | Environment | Episodes to Solve | Final Avg Reward |
|---|---|---|---|
| REINFORCE (Vanilla) | CartPole-v1 | 262 | 450.80 |
| REINFORCE (Baseline) | CartPole-v1 | 495 | 461.60 |
| A2C | CartPole-v1 | 699 | 453.50 |
| DDPG | Pendulum-v1 | 172 | -195.39 |
✅ REINFORCE vanilla solved fastest (262 episodes) - simple but effective
✅ Baseline version reduced variance significantly
✅ A2C enables online learning without episode waits
✅ DDPG efficiently handles continuous actions with replay buffer
git clone https://github.com/belay-cell/RL-Policy-Gradient-Algorithms.git
cd RL-Policy-Gradient-Algorithms
pip install -r requirements.txtimport gymnasium as gym
from models import SoftmaxPolicy
from reinforce import REINFORCE
env = gym.make("CartPole-v1")
policy = SoftmaxPolicy(env.observation_space.shape[0], env.action_space.n)
agent = REINFORCE(policy, lr=1e-3, solve_criteria=450, episode_limit=1000)
total_rewards = agent.train(env)import gymnasium as gym
from models import DDPGActor, DDPGCritic
from utils import ReplayBuffer
from ddpg import DDPG
env = gym.make("Pendulum-v1")
actor = DDPGActor(3, 1, action_range=2.0)
critic = DDPGCritic(3, 1)
buffer = ReplayBuffer(10000)
agent = DDPG(buffer, actor, critic, actor_lr=1e-4, critic_lr=1e-3)
total_rewards = agent.train(env)RL-Policy-Gradient-Algorithms/
├── models.py # Neural network architectures
├── reinforce.py # REINFORCE (vanilla + baseline)
├── a2c.py # Advantage Actor-Critic
├── ddpg.py # Deep Deterministic Policy Gradient
├── utils.py # Replay buffer utilities
├── requirements.txt # Dependencies
├── README.md # This file
└── LICENSE # MIT License
SoftmaxPolicy (Discrete Actions - CartPole)
state (4D) → FC(128) → ReLU → FC(64) → ReLU → FC(2) → Softmax → action_probs
Critic (Value Function - A2C)
state → FC(128) → ReLU → FC(64) → ReLU → FC(1) → value
DDPGActor (Continuous Actions - Pendulum)
state (3D) → FC(128) → ReLU → FC(64) → ReLU → FC(1) → Tanh → action * 2.0
DDPGCritic (State-Action Value)
concat(state, action) → FC(128) → ReLU → FC(64) → ReLU → FC(1) → Q-value
Policy Gradient Theorem:
∇θ J(θ) = 𝔼[∇θ log πθ(a|s) · G_t]
With Baseline:
θ ← θ + α ∑_t (G_t - baseline) ∇θ log πθ(a_t|s_t)
Advantage Function:
A(s_t, a_t) = R_{t+1} + γV(s_{t+1}) - V(s_t)
Update Rules:
Actor: θπ ← θπ + α ∇θπ[A · log π(a|s)]
Critic: θv ← θv - α ∇θv[V(s) - (R + γV(s'))]^2
Key Features:
- Deterministic policy:
a = μ(s) - Experience replay (capacity: 10,000)
- Target networks with soft updates (τ=0.005)
- Gaussian exploration noise
| Parameter | REINFORCE | A2C | DDPG |
|---|---|---|---|
| Learning Rate | 1e-3 | Actor:1e-3, Critic:1e-3 | Actor:1e-4, Critic:1e-3 |
| Discount (γ) | 0.99 | 0.99 | 0.99 |
| Batch Size | Full episode | 1 (online) | 32 |
| Replay Buffer | - | - | 10,000 |
| Soft Update (τ) | - | - | 0.005 |
| Solve Criteria | 450 | 450 | -200 |
- Sutton & Barto. Reinforcement Learning: An Introduction (2nd ed.), Chapter 13
- Mnih et al. Asynchronous Methods for Deep Reinforcement Learning (2016)
- Lillicrap et al. Continuous Control with Deep Reinforcement Learning (2015)
Originally developed for KAIST CS377 Reinforcement Learning, refactored into professional-grade codebase.
MIT License - See LICENSE
Belay Zeleke
GitHub: @belay-cell
Interested in MLOps, Reinforcement Learning, and Production ML Systems
Built with PyTorch 🔥 | Trained on Gymnasium Environments 🎮