This repository contains a Computer Vision (CV) pipeline for detecting good vs. bad prints on medical product labels. The project includes:
- Deep CNN classification using DenseNet121
- OCR extraction with rule-based validation
- Multimodal deep learning that combines image features and OCR-based features
├── data # Data files
│ ├── print_images_task/
│ └── labels(in).csv
│ └── possible_values.json
├── data_load.py # Script to split data into train/test sets
├── data_augmentation.py # Script for augmenting training images
├── data_transform.py # Common transforms used for model training and inference
├── test/ # Auto-generated test data (from data_load.py)
│ └── images/
│ └── labels.csv
├── train/ # Auto-generated train data (from data_load.py & data_augmentation.py)
│ └── aug/
│ └── images/
│ └── train_total/
│ └── labels.csv
│ └── labels_aug.csv
│ └── train_total_labels.csv
├── ocr_extract.py # Runs EasyOCR on images and saves results
├── ocr_results/ # Auto-generated OCR output directory(from ocr_extract.py)
│ └── train/
│ └── test/
├── ocr_to_fields.py # Maps OCR text to structured fields using rules
├── ocr_train_fields.csv # Auto-generated (from ocr_to_fields.py)
├── ocr_test_fields.csv # Auto-generated (from ocr_to_fields.py)
├── train.py
├── train_prediction.csv # Auto-generated (from train.py)
├── test_prediction.csv # Auto-generated (from train.py)
├── train_multimodal.py
├── models # Folder where trained models are stored
│ └── densenet121_reg.pth
│ └── multimodal_densenet.pth
├── inference.py
├── inference_ocr.py
├── inference_multimodal.py
├── metricts.txt # Auto-generated (from train.py & train_multimodal.py)
├── requirements.txt # List of dependencies required to run the project.
├── documentation.pdf # Additional documentation
├── binary_classifier_trials.ipynb # different models are tested, compared, and evaluated
└── README.md
-
Open your terminal: Launch your terminal application.
-
Choose a directory: Navigate to the directory where you want to clone the project using the
cdcommand.cd path/to/your/directory -
Clone the repository: Copy the project into your local machine by running the following command in your terminal. Make sure to replace with the actual URL of the repository you want to clone.
git clone <repository-url>
-
Navigate to the project directory: Once the cloning process is complete, move into the project directory.
cd repository-name -
Check the files: You can list the files in the directory to see all the project files you have cloned.
ls
-
Create a virtual environment:
python3 -m venv {virtual_environment_name}- Activate the virtual environment:
- On Windows:
{virtual_environment_name}\Scripts\activate - On macOS or Linux:
source {virtual_environment_name}/bin/activate
- On Windows:
- Make sure to install all necessary dependencies first by running:
pip install -r requirements.txt- Load raw images
python3 data_load.py- Run augmentation
python3 data_augmentation.py- Prepare transforms and dataloaders
python3 data_transform.py- Run OCR on datasets
python3 ocr_extract.pyThis generates .ocr.json files inside ocr_results/ directory
- Convert OCR results into structured CSV files
For training images:
python3 ocr_to_fields.py --ocr_dir ocr_results/train --possible data/possible_values.json --out_csv ocr_train_fields.csvFor test images:
python3 ocr_to_fields.py --ocr_dir ocr_results/test --possible data/possible_values.json --out_csv ocr_test_fields.csvThese files (ocr_train_fields.csv, ocr_test_fields.csv) are used as inputs to the multimodal model.
Approach 1 – Pure CNN Classification
Train the CNN model:
python3 train.pyAfter run model is saved as models/densenet121_reg.pth
Run inference on a single image:
python3 inference.py --image test/images/good_28.bmpApproach 2 – CNN + OCR (Rule-Based Hybrid)
python3 inference_ocr.py --image train/images/good_1.bmpApproach 3 – Multimodal Deep Learning (CNN + OCR Vector)
Train the multimodal model:
python3 train_multimodal.pyAfter run model is saved as models/multimodal_densenet.pth
Run multimodal inference
python3 inference_multimodal.py --image train/images/good_1.bmpAfter each training run, performance metrics are saved (and updated) in metrics.txt
For more detailed explanation and methodology, please refer to the full documentation in documentation.pdf