A high-performance, deterministic Reinforcement Learning engine written in C++ that masters perfect-information board games through Iterative Self-Play. It utilizes a custom memory pool for rapid tree traversal and integrates a PyTorch neural network via LibTorch for state evaluation.
The project bridges low-level systems programming with a machine learning pipeline:
- Core Engine: Written in C++ for maximum execution speed and strict memory control.
- Neural Bridge: Utilizes the LibTorch (PyTorch C++) API to convert board states into tensors and run inference directly inside the C++ runtime.
- Search Algorithm: Implements Monte Carlo Tree Search (MCTS) utilizing the PUCT (Predictor Upper Confidence Bound applied to Trees) formula to balance exploitation and exploration.
- Training Pipeline: A Python-based PyTorch script that consumes C++ generated
.csvself-play data, trains a ResNet-style neural network, and exports the optimized weights via TorchScript (.pt) back to the C++ engine.
Standard dynamic memory allocation (new/delete) creates massive bottlenecks and heap fragmentation during millions of MCTS node expansions.
- To solve this, the engine uses a custom
NodePoolallocator. - It pre-allocates a massive contiguous block of memory (
std::vector<MCTSNode>) at runtime startup. - Node allocation is reduced to an
$O(1)$ index increment operation, ensuring zero latency during deep tree searches and guaranteeing cache locality.
- CMake (3.10+)
- Python 3 with a configured
venv - LibTorch (CPU version)
- Clone the repository and create a build directory:
mkdir build && cd build - Run CMake and compile:
cmake .. && make - Execute the engine:
./rl_engine