Coursework for EE675 (Introduction to Reinforcement Learning). Both assignments study the same predator-prey pursuit game, first solved exactly with dynamic programming and then approximately with a learned neural policy.
| Folder | Topic | Methods |
|---|---|---|
Assignment 1 |
Predator-prey MDP on an N x N grid |
MDP formulation, iterative policy evaluation, value iteration, policy iteration, model-based kernel estimation |
Assignment 2 |
Predator-prey game on a 4 x 4 grid |
REINFORCE (policy gradient) with a neural-network policy, plain SGA vs Adam, hyperparameter sweep |
Each assignment folder has its own README.md, and every question in Assignment 1 has a README.md
with exact run instructions.
- A predator and a prey occupy cells of a grid. Time is discrete.
- Actions are
{Stay, Up, Down, Left, Right}. Moves that would leave the grid keep the mover in place. - The prey moves to a uniformly random valid neighbouring cell (or stays) at every step.
- The predator gets reward
+1when it lands on the prey's cell and0otherwise. - In Assignment 1 the prey respawns uniformly at random after being caught; the discount factor is
gamma = 0.99. Assignment 2 uses episodic rollouts on a fixed4 x 4grid.
Assignment 1 represents the state as (predator, prey) with |S| = N^4 and |A| = 5, and stores
the transition kernel and reward matrix as sparse SciPy matrices so that grids up to N = 25
(390,625 states) stay tractable. Assignment 2 feeds the four normalised coordinates into an MLP that
outputs a softmax distribution over the five actions.
# Assignment 1
pip install numpy scipy matplotlib
# Assignment 2
pip install torch numpy matplotlib pyyaml.
├── Assignment 1/
│ ├── README.md
│ ├── EE675_Assignment_1.pdf
│ ├── Question 1/ MDP formulation and policy evaluation
│ ├── Question 2/ value iteration and policy iteration
│ └── Question 3/ model-based kernel estimation
└── Assignment 2/
├── README.md
├── EE675_Assignment_2.pdf
├── nn_policy.py, gradient_estimate.py, simulator.py, utils.py
├── train_sga.py, train_adam.py, sweep.py, config.yaml
└── learning-curve PNGs and sweep_results.txt