Seven algorithms for detecting driver state from video: Rule-Based, SVM, Random Forest, MLP, CNN, LSTM, and TCN.
Task: detect driver drowsiness and impaired alertness from video in real time.
Solution: extract facial landmarks with MediaPipe FaceMesh, compute eye aspect ratio (EAR), mouth aspect ratio (MAR), head pose, and gaze direction per frame. Classify driver state using seven algorithms — a rule-based threshold baseline, three per-frame ML models (SVM, Random Forest, MLP), and three temporal sequence models (CNN, LSTM, TCN) that look at sliding windows of recent frames. All sequence models run on pure NumPy without PyTorch or TensorFlow.
core/ - FaceMesh wrapper, feature extraction (EAR/MAR/pose/gaze), temporal buffer
models/ - All 7 algorithms + model factory
models/nn/ - NumPy-only layers (Dense, Conv1D, LSTM, Adam) — no PyTorch/TensorFlow
utils/ - Config, dataset builder, logger, temporal smoother
main.py - Run any trained model on video/webcam
train.py - Train and compare all models
compare.py - Fair side-by-side comparison of all models on the same input
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt# Train
python train.py --video test_video.mp4
python train.py --video test_video.mp4 --models svm mlp
python train.py --video test_video.mp4 --models cnn lstm tcn --epochs 50
# Run
python main.py --video test_video.mp4 --algorithm rule_based
python main.py --video 0 --algorithm cnn
# Compare all models on the same input
python compare.py --webcam
python compare.py --video test_video.mp4 --algorithms svm lstm tcn --log output/compare.csvSequence models (CNN/LSTM/TCN) need window_size frames of history before predicting; until then they report Alert. compare.py runs one capture loop and feeds every model the exact same frame for a fair comparison.