Skip to content

About

This repo contains a Computer Vision (CV) pipeline for detecting good vs. bad prints on medical product labels

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Data Science Task

This repository contains a Computer Vision (CV) pipeline for detecting good vs. bad prints on medical product labels. The project includes:

  • Deep CNN classification using DenseNet121
  • OCR extraction with rule-based validation
  • Multimodal deep learning that combines image features and OCR-based features

Project Structure

├── data                        # Data files  
│   ├── print_images_task/
│   └── labels(in).csv
│   └── possible_values.json
├── data_load.py                # Script to split data into train/test sets
├── data_augmentation.py        # Script for augmenting training images
├── data_transform.py           # Common transforms used for model training and inference
├── test/                       # Auto-generated test data (from data_load.py)
│   └── images/
│   └── labels.csv
├── train/                      # Auto-generated train data (from data_load.py & data_augmentation.py)
│   └── aug/
│   └── images/
│   └── train_total/
│   └── labels.csv
│   └── labels_aug.csv
│   └── train_total_labels.csv
├── ocr_extract.py             # Runs EasyOCR on images and saves results
├── ocr_results/               # Auto-generated OCR output directory(from ocr_extract.py)
│    └── train/
│    └── test/
├── ocr_to_fields.py           # Maps OCR text to structured fields using rules
├── ocr_train_fields.csv       # Auto-generated (from ocr_to_fields.py)  
├── ocr_test_fields.csv        # Auto-generated (from ocr_to_fields.py)
├── train.py    
├── train_prediction.csv       # Auto-generated (from train.py)  
├── test_prediction.csv        # Auto-generated (from train.py) 
├── train_multimodal.py   
├── models                     # Folder where trained models are stored
│   └── densenet121_reg.pth
│   └── multimodal_densenet.pth  
├── inference.py   
├── inference_ocr.py 
├── inference_multimodal.py              
├── metricts.txt               # Auto-generated (from train.py & train_multimodal.py)       
├── requirements.txt           # List of dependencies required to run the project.
├── documentation.pdf          # Additional documentation
├── binary_classifier_trials.ipynb # different models are tested, compared, and evaluated
└── README.md

Part 1: Cloning the Project

  1. Open your terminal: Launch your terminal application.

  2. Choose a directory: Navigate to the directory where you want to clone the project using the cd command.

    cd path/to/your/directory
  3. Clone the repository: Copy the project into your local machine by running the following command in your terminal. Make sure to replace with the actual URL of the repository you want to clone.

    git clone <repository-url>
  4. Navigate to the project directory: Once the cloning process is complete, move into the project directory.

    cd repository-name
  5. Check the files: You can list the files in the directory to see all the project files you have cloned.

    ls
  6. Create a virtual environment:

 python3 -m venv {virtual_environment_name}
  1. Activate the virtual environment:
    • On Windows:
      {virtual_environment_name}\Scripts\activate
    • On macOS or Linux:
      source {virtual_environment_name}/bin/activate
  2. Make sure to install all necessary dependencies first by running:
 pip install -r requirements.txt

Part 2: Prepare the Dataset

  1. Load raw images
   python3 data_load.py
  1. Run augmentation
  python3 data_augmentation.py
  1. Prepare transforms and dataloaders
  python3 data_transform.py

Part 3: OCR Pipeline

  1. Run OCR on datasets
   python3 ocr_extract.py

This generates .ocr.json files inside ocr_results/ directory

  1. Convert OCR results into structured CSV files

For training images:

   python3 ocr_to_fields.py --ocr_dir ocr_results/train --possible data/possible_values.json --out_csv ocr_train_fields.csv

For test images:

  python3 ocr_to_fields.py --ocr_dir ocr_results/test --possible data/possible_values.json --out_csv ocr_test_fields.csv

These files (ocr_train_fields.csv, ocr_test_fields.csv) are used as inputs to the multimodal model.

Part 4: Model Training & Inference

Approach 1 – Pure CNN Classification

Train the CNN model:

  python3 train.py

After run model is saved as models/densenet121_reg.pth

Run inference on a single image:

  python3 inference.py --image test/images/good_28.bmp

Approach 2 – CNN + OCR (Rule-Based Hybrid)

  python3 inference_ocr.py --image train/images/good_1.bmp

Approach 3 – Multimodal Deep Learning (CNN + OCR Vector)

Train the multimodal model:

   python3 train_multimodal.py

After run model is saved as models/multimodal_densenet.pth

Run multimodal inference

  python3 inference_multimodal.py --image train/images/good_1.bmp

After each training run, performance metrics are saved (and updated) in metrics.txt

For more detailed explanation and methodology, please refer to the full documentation in documentation.pdf

About

This repo contains a Computer Vision (CV) pipeline for detecting good vs. bad prints on medical product labels

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages