Skip to content

About

🧠 Convolutional Neural Network implemented entirely from scratch in NumPy. Features manual convolution, ReLU, max pooling, backpropagation, MNIST training, FastAPI deployment, and an interactive web interface for real-time handwritten digit classification.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🧠 NeuralVis β€” Handwritten Digits Classifier

A Convolutional Neural Network built entirely from scratch in NumPy, with a live interactive visualisation dashboard.

Live Demo GitHub Python React FastAPI


✨ What is this?

NeuralVis is a full-stack machine learning project that implements a Convolutional Neural Network (CNN) entirely from scratch β€” no PyTorch, no TensorFlow, no Keras. Every single mathematical operation (convolution, backpropagation, gradient descent) is hand-coded using NumPy only.

The project then wraps this raw CNN in a polished interactive web interface where you can:

  • ✏️ Draw any digit (0–9) on a canvas
  • ⚑ Watch the CNN classify it in real time with live probability scores
  • πŸ”¬ Inspect every internal layer β€” raw feature maps, ReLU activations, pooled maps, and hidden neuron activations
  • πŸ—οΈ Explore the architecture in an interactive node graph
  • πŸ”„ Follow the data pipeline step by step through the network

⚠️ Latency Notice: The hosted demo on Render uses a free-tier server that spins down when idle. For real-time, low-latency predictions, run the project locally (see below).


🎬 Live Demo

https://handwritten-digits-classifier-yx3h.onrender.com

The first request after a period of inactivity may take ~30 seconds while the server wakes up. All subsequent predictions will be fast.


🧠 The Concepts

This project implements a classic CNN architecture. Here's what every layer does:

1. πŸ“₯ Input β€” 28 Γ— 28

The MNIST dataset images are 28Γ—28 grayscale images. Pixel values are normalised to the range [0, 1] (divided by 255). When you draw on the canvas, your drawing is preprocessed into this exact format before being sent to the model.

2. πŸ”² Convolution Layer β€” 16 filters, 3Γ—3 kernel, stride 1

The core operation of any CNN. 16 learnable 3Γ—3 filters slide across the 28Γ—28 input image. Each filter looks at a small 3Γ—3 patch of pixels at a time and computes a weighted sum β€” effectively detecting a specific low-level pattern (an edge, a curve, a corner). Each filter produces one feature map of shape 26Γ—26 (because a 3Γ—3 filter on a 28Γ—28 image with no padding fits in 26 positions per axis). The result is 16 separate feature maps.

Output shape: (16, 26, 26)
Parameters:   16 Γ— (3Γ—3 weights + 1 bias) = 160

The convolution at each position (i, j) is:

$$\text{output}[i,j] = \sum_{m=0}^{2} \sum_{n=0}^{2} \text{image}[i+m,\ j+n] \times \text{kernel}[m,n] + \text{bias}$$

3. ⚑ ReLU Activation β€” Rectified Linear Unit

Applied element-wise after convolution. Any negative value becomes zero, positive values pass through unchanged:

$$\text{ReLU}(x) = \max(0, x)$$

This introduces non-linearity into the network. Without activation functions, stacking multiple layers would still only be capable of computing a linear transformation β€” no matter how deep the network. ReLU also helps avoid the vanishing gradient problem.

4. πŸ”½ Max Pooling β€” 2Γ—2 pool, stride 2

Divides each 26Γ—26 feature map into non-overlapping 2Γ—2 patches and keeps only the maximum value from each patch. This:

  • Halves the spatial dimensions from 26Γ—26 β†’ 13Γ—13
  • Makes the representation translation-invariant (the model detects the presence of a feature regardless of its exact position)
  • Reduces the number of parameters and computation for subsequent layers
Output shape: (16, 13, 13)
Parameters:   0 (no learnable weights)

5. πŸ“ Flatten

Unrolls the 3D tensor (16, 13, 13) into a single 1D vector of length 2704 (16 Γ— 13 Γ— 13 = 2704). This bridges the convolutional part of the network to the fully-connected (dense) layers.

6. πŸ”— Hidden Dense Layer β€” 2704 β†’ 64 neurons

A fully-connected layer where every one of the 2704 inputs connects to each of the 64 output neurons. Each neuron computes:

$$z = W \cdot x + b$$

This layer learns high-level combinations of the spatial features extracted by the conv layer.

Parameters: 2704 Γ— 64 weights + 64 biases = 173,120

7. ⚑ ReLU Activation (again)

Same as before β€” introduces non-linearity into the dense layer.

8. 🎯 Output Dense Layer β€” 64 β†’ 10 neurons

Produces 10 raw scores (logits), one per digit class (0–9). The neuron with the highest score corresponds to the model's predicted digit.

Parameters: 64 Γ— 10 weights + 10 biases = 650

9. πŸ“Š Softmax

Converts the 10 raw logits into a probability distribution that sums to 1.0. Each value represents the model's confidence for that digit class:

$$P(y = k) = \frac{e^{z_k}}{\sum_{j=0}^{9} e^{z_j}}$$

The implementation subtracts the maximum logit before exponentiating for numerical stability.


πŸ” Backpropagation & Training

The network is trained using mini-batch stochastic gradient descent with cross-entropy loss:

$$\mathcal{L} = -\sum_{k} y_k \log(\hat{y}_k)$$

Gradients are propagated backwards through every layer β€” softmax, dense, ReLU, flatten, pooling, and convolution β€” computing dW, db, and d_input for each. Weights are updated with a fixed learning rate of 0.005.

Training Configuration

Parameter Value
Dataset MNIST (provided in backend/dataset/)
Training samples 10,000 images
Epochs 20
Batch size 32
Learning rate 0.005
Weight save frequency Every 5 epochs (when test accuracy improves)
Random seed 0

Pre-trained weights are included in backend/weights/ β€” you don't need to retrain from scratch. The model loads them automatically when the server starts.


πŸ—‚οΈ Project Structure

CNN from scratch/
β”œβ”€β”€ backend/
β”‚   β”œβ”€β”€ api.py              # FastAPI server β€” wraps the CNN with HTTP endpoints
β”‚   β”œβ”€β”€ main.py             # The raw CNN: all training + inference logic (NumPy only)
β”‚   β”œβ”€β”€ requirements.txt    # Python dependencies
β”‚   β”œβ”€β”€ weights/            # βœ… Pre-trained model weights (included)
β”‚   β”‚   β”œβ”€β”€ kernels.npy
β”‚   β”‚   β”œβ”€β”€ conv_biases.npy
β”‚   β”‚   β”œβ”€β”€ W1.npy
β”‚   β”‚   β”œβ”€β”€ b1.npy
β”‚   β”‚   β”œβ”€β”€ W2.npy
β”‚   β”‚   └── b2.npy
β”‚   └── dataset/            # βœ… MNIST dataset binary files (included)
β”‚       β”œβ”€β”€ train-images.idx3-ubyte
β”‚       β”œβ”€β”€ train-labels.idx1-ubyte
β”‚       β”œβ”€β”€ t10k-images.idx3-ubyte
β”‚       └── t10k-labels.idx1-ubyte
β”‚
β”œβ”€β”€ frontend/
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ components.tsx  # The entire React frontend in one file
β”‚   β”‚   β”œβ”€β”€ index.css       # Global styles & Tailwind directives
β”‚   β”‚   └── vite-env.d.ts   # Vite type references
β”‚   β”œβ”€β”€ index.html
β”‚   β”œβ”€β”€ vite.config.ts
β”‚   β”œβ”€β”€ tailwind.config.js
β”‚   β”œβ”€β”€ postcss.config.js
β”‚   β”œβ”€β”€ tsconfig.json
β”‚   └── package.json
β”‚
β”œβ”€β”€ main.ipynb              # Original Jupyter notebook used for development
└── .gitignore

πŸ› οΈ Tech Stack

Layer Technology
CNN / Training Python, NumPy
API Server FastAPI, Uvicorn
Image Processing Pillow
Frontend Framework React 18, TypeScript
Build Tool Vite
Styling Tailwind CSS
Animations Framer Motion
State Management Zustand
Deployment Render (Backend: Web Service, Frontend: Static Site)

πŸš€ Running Locally

Prerequisites

  • Python 3.11+
  • Node.js 18+ and npm
  • Git

Step 1 β€” Clone the Repository

git clone https://github.com/AA-KH/Handwritten-Digits-Classifier.git
cd Handwritten-Digits-Classifier

Step 2 β€” Start the Backend

# Navigate to the backend folder
cd backend

# Create and activate a virtual environment (recommended)
python -m venv .venv
source .venv/bin/activate      # On Windows: .venv\Scripts\activate

# Install Python dependencies
pip install -r requirements.txt

# Start the API server
uvicorn api:app --host 0.0.0.0 --port 8000

You should see:

βœ… Model weights loaded successfully
   kernels: (16, 3, 3), W1: (2704, 64), W2: (64, 10)
INFO:     Uvicorn running on http://0.0.0.0:8000

You can verify the backend is healthy at: http://localhost:8000/health


Step 3 β€” Start the Frontend

Open a second terminal and run:

# Navigate to the frontend folder
cd frontend

# Install Node dependencies
npm install

# Start the Vite development server
npm run dev

You should see:

  VITE v5.x.x  ready in 300ms

  ➜  Local:   http://localhost:5173/

Open http://localhost:5173 in your browser. The frontend automatically proxies all /api requests to the backend on port 8000.


Step 4 β€” (Optional) Retrain the Model

If you want to retrain the CNN from scratch, run main.py from inside the backend/ directory:

cd backend
python main.py

⚠️ Note: Training runs 20 epochs with batches of 32 samples on 10,000 training images. The model evaluates on the test set every 5 epochs and saves updated weights to backend/weights/ only when test accuracy improves. Training is CPU-only and may take some time.


πŸ“‘ API Endpoints

Once the backend is running, these endpoints are available:

Method Endpoint Description
GET /health Health check β€” confirms model is loaded
POST /predict Accepts a base64 PNG, returns predictions + all layer activations
GET /architecture Returns static CNN architecture metadata

πŸ“ Notes

  • Pre-trained weights are included β€” the model loads them immediately on startup. No training required to use the app.
  • The MNIST dataset is included in backend/dataset/ β€” needed only if you want to retrain.
  • The model was trained for 20 epochs on 10,000 images with weights saved every 5 epochs whenever test accuracy improved.
  • The CNN is intentionally kept simple (single conv layer) to make every layer's behaviour clearly observable in the visualiser.
  • main.py is the original, untouched training script. api.py wraps it for HTTP access without modifying it.

Built with 🀍 β€” pure NumPy, no ML frameworks.

About

🧠 Convolutional Neural Network implemented entirely from scratch in NumPy. Features manual convolution, ReLU, max pooling, backpropagation, MNIST training, FastAPI deployment, and an interactive web interface for real-time handwritten digit classification.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages