Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

EE675: Introduction to Reinforcement Learning

Coursework for EE675 (Introduction to Reinforcement Learning). Both assignments study the same predator-prey pursuit game, first solved exactly with dynamic programming and then approximately with a learned neural policy.

Contents

Folder Topic Methods
Assignment 1 Predator-prey MDP on an N x N grid MDP formulation, iterative policy evaluation, value iteration, policy iteration, model-based kernel estimation
Assignment 2 Predator-prey game on a 4 x 4 grid REINFORCE (policy gradient) with a neural-network policy, plain SGA vs Adam, hyperparameter sweep

Each assignment folder has its own README.md, and every question in Assignment 1 has a README.md with exact run instructions.

The common environment

  • A predator and a prey occupy cells of a grid. Time is discrete.
  • Actions are {Stay, Up, Down, Left, Right}. Moves that would leave the grid keep the mover in place.
  • The prey moves to a uniformly random valid neighbouring cell (or stays) at every step.
  • The predator gets reward +1 when it lands on the prey's cell and 0 otherwise.
  • In Assignment 1 the prey respawns uniformly at random after being caught; the discount factor is gamma = 0.99. Assignment 2 uses episodic rollouts on a fixed 4 x 4 grid.

Assignment 1 represents the state as (predator, prey) with |S| = N^4 and |A| = 5, and stores the transition kernel and reward matrix as sparse SciPy matrices so that grids up to N = 25 (390,625 states) stay tractable. Assignment 2 feeds the four normalised coordinates into an MLP that outputs a softmax distribution over the five actions.

Dependencies

# Assignment 1
pip install numpy scipy matplotlib

# Assignment 2
pip install torch numpy matplotlib pyyaml

Layout

.
├── Assignment 1/
│   ├── README.md
│   ├── EE675_Assignment_1.pdf
│   ├── Question 1/   MDP formulation and policy evaluation
│   ├── Question 2/   value iteration and policy iteration
│   └── Question 3/   model-based kernel estimation
└── Assignment 2/
    ├── README.md
    ├── EE675_Assignment_2.pdf
    ├── nn_policy.py, gradient_estimate.py, simulator.py, utils.py
    ├── train_sga.py, train_adam.py, sweep.py, config.yaml
    └── learning-curve PNGs and sweep_results.txt

About

Reinforcement learning coursework: a predator-prey MDP solved with dynamic programming and model-based kernel estimation, and a REINFORCE policy-gradient agent trained with SGA and Adam.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages