- Real-Time Gesture Capture: Uses a webcam to capture images of hand gestures.
- Dynamic Labeling: Allows users to assign labels to gestures during capture.
- Data Storage: Saves images in a structured folder format based on labels.
- Model Training: Trains a Convolutional Neural Network (CNN) on the captured data.
- Real-Time Prediction: Predicts gestures in real-time using a webcam feed.
- LLM Integration: Need to be worked on
- Python 3.8+
- Libraries:
opencv-pythonmediapipenumpytensorflowsklearnopenai(for LLM integration)
Install the required libraries using:
pip install opencv-python mediapipe numpy tensorflow scikit-learn openai.
├── CollectedData/ # Folder containing labeled gesture images
│ ├── Label1/
│ │ ├── 1.jpg
│ │ ├── 2.jpg
│ ├── Label2/
│ ├── 1.jpg
│ ├── 2.jpg
|
├── models/ # Folder containing labeled gesture images
│ ├── model1/
| ├── model2/
| ├── model3/ # Trained gesture recognition model
├── label_class.npy # Label encoder for mapping labels
└── Final(Revised).ipynb # Main script for capturing, training, and predicting
└── README.md # Project overview and instructions
└── frontend
└── backend
Run the script to start capturing gestures:
python main.py- Press
's'to save a labeled gesture. - Press
'q'to quit.
Captured images will be stored in CollectedData/ under subfolders named after labels.
Use the captured data to train a CNN model:
# In your script, call the train_model function
train_model('CollectedData/', 'model.h5', 'label_encoder_classes.npy')This saves the trained model as model.h5(change the name as per the model to be used eg:- model1.h5, model2.h5, model3.h5) and label encoder as label_class.npy.
Run the script for real-time prediction:
real_time_prediction('model.h5', 'label_class.npy')- The predicted label will be displayed on the webcam feed.
- Press
's'to save the predicted label tosaved_labels.txt.
Further a basic frontend and backend has been developed using Node.js and express.js which uses the pretrained model to predict the gestures in real time and convert them into text and pass it to the backend where it can be used for further processing by the LLM model.
-TO RUN THE FRONTEND
cd frontend
npm install
npm start-TO RUN THE BACKEND
cd backend
npm install
uvicorn main:app --reload-
Gesture Capture:
- Captures and labels hand gesture images.
- Displays landmarks using Mediapipe.
- Saves cropped hand images for training.
-
Model Training:
- Builds and trains a CNN model.
- Saves the trained model and label encoder.
-
Real-Time Prediction:
- Uses the trained model to predict gestures in real time.
- Displays predictions on the webcam feed.
-
LLM Integration:
- Basic help of Gemini API for LLM .
- Extend the dataset with more gestures and labels.
- Implement multi-hand gesture recognition.
- Use of video data for training and prediction.
- Implement a more advanced model for gesture recognition.
- Enhance the UI for gesture capture and prediction.
- Integrate additional AI models for more advanced recognition.
- Mediapipe for hand tracking and landmarks.
- TensorFlow/Keras for deep learning.

