NeuralVis is a full-stack machine learning project that implements a Convolutional Neural Network (CNN) entirely from scratch β no PyTorch, no TensorFlow, no Keras. Every single mathematical operation (convolution, backpropagation, gradient descent) is hand-coded using NumPy only.
The project then wraps this raw CNN in a polished interactive web interface where you can:
- βοΈ Draw any digit (0β9) on a canvas
- β‘ Watch the CNN classify it in real time with live probability scores
- π¬ Inspect every internal layer β raw feature maps, ReLU activations, pooled maps, and hidden neuron activations
- ποΈ Explore the architecture in an interactive node graph
- π Follow the data pipeline step by step through the network
β οΈ Latency Notice: The hosted demo on Render uses a free-tier server that spins down when idle. For real-time, low-latency predictions, run the project locally (see below).
https://handwritten-digits-classifier-yx3h.onrender.com
The first request after a period of inactivity may take ~30 seconds while the server wakes up. All subsequent predictions will be fast.
This project implements a classic CNN architecture. Here's what every layer does:
The MNIST dataset images are 28Γ28 grayscale images. Pixel values are normalised to the range [0, 1] (divided by 255). When you draw on the canvas, your drawing is preprocessed into this exact format before being sent to the model.
The core operation of any CNN. 16 learnable 3Γ3 filters slide across the 28Γ28 input image. Each filter looks at a small 3Γ3 patch of pixels at a time and computes a weighted sum β effectively detecting a specific low-level pattern (an edge, a curve, a corner). Each filter produces one feature map of shape 26Γ26 (because a 3Γ3 filter on a 28Γ28 image with no padding fits in 26 positions per axis). The result is 16 separate feature maps.
Output shape: (16, 26, 26)
Parameters: 16 Γ (3Γ3 weights + 1 bias) = 160
The convolution at each position (i, j) is:
Applied element-wise after convolution. Any negative value becomes zero, positive values pass through unchanged:
This introduces non-linearity into the network. Without activation functions, stacking multiple layers would still only be capable of computing a linear transformation β no matter how deep the network. ReLU also helps avoid the vanishing gradient problem.
Divides each 26Γ26 feature map into non-overlapping 2Γ2 patches and keeps only the maximum value from each patch. This:
- Halves the spatial dimensions from
26Γ26β13Γ13 - Makes the representation translation-invariant (the model detects the presence of a feature regardless of its exact position)
- Reduces the number of parameters and computation for subsequent layers
Output shape: (16, 13, 13)
Parameters: 0 (no learnable weights)
Unrolls the 3D tensor (16, 13, 13) into a single 1D vector of length 2704 (16 Γ 13 Γ 13 = 2704). This bridges the convolutional part of the network to the fully-connected (dense) layers.
6. π Hidden Dense Layer β 2704 β 64 neurons
A fully-connected layer where every one of the 2704 inputs connects to each of the 64 output neurons. Each neuron computes:
This layer learns high-level combinations of the spatial features extracted by the conv layer.
Parameters: 2704 Γ 64 weights + 64 biases = 173,120
Same as before β introduces non-linearity into the dense layer.
Produces 10 raw scores (logits), one per digit class (0β9). The neuron with the highest score corresponds to the model's predicted digit.
Parameters: 64 Γ 10 weights + 10 biases = 650
Converts the 10 raw logits into a probability distribution that sums to 1.0. Each value represents the model's confidence for that digit class:
The implementation subtracts the maximum logit before exponentiating for numerical stability.
The network is trained using mini-batch stochastic gradient descent with cross-entropy loss:
Gradients are propagated backwards through every layer β softmax, dense, ReLU, flatten, pooling, and convolution β computing dW, db, and d_input for each. Weights are updated with a fixed learning rate of 0.005.
| Parameter | Value |
|---|---|
| Dataset | MNIST (provided in backend/dataset/) |
| Training samples | 10,000 images |
| Epochs | 20 |
| Batch size | 32 |
| Learning rate | 0.005 |
| Weight save frequency | Every 5 epochs (when test accuracy improves) |
| Random seed | 0 |
Pre-trained weights are included in
backend/weights/β you don't need to retrain from scratch. The model loads them automatically when the server starts.
CNN from scratch/
βββ backend/
β βββ api.py # FastAPI server β wraps the CNN with HTTP endpoints
β βββ main.py # The raw CNN: all training + inference logic (NumPy only)
β βββ requirements.txt # Python dependencies
β βββ weights/ # β
Pre-trained model weights (included)
β β βββ kernels.npy
β β βββ conv_biases.npy
β β βββ W1.npy
β β βββ b1.npy
β β βββ W2.npy
β β βββ b2.npy
β βββ dataset/ # β
MNIST dataset binary files (included)
β βββ train-images.idx3-ubyte
β βββ train-labels.idx1-ubyte
β βββ t10k-images.idx3-ubyte
β βββ t10k-labels.idx1-ubyte
β
βββ frontend/
β βββ src/
β β βββ components.tsx # The entire React frontend in one file
β β βββ index.css # Global styles & Tailwind directives
β β βββ vite-env.d.ts # Vite type references
β βββ index.html
β βββ vite.config.ts
β βββ tailwind.config.js
β βββ postcss.config.js
β βββ tsconfig.json
β βββ package.json
β
βββ main.ipynb # Original Jupyter notebook used for development
βββ .gitignore
| Layer | Technology |
|---|---|
| CNN / Training | Python, NumPy |
| API Server | FastAPI, Uvicorn |
| Image Processing | Pillow |
| Frontend Framework | React 18, TypeScript |
| Build Tool | Vite |
| Styling | Tailwind CSS |
| Animations | Framer Motion |
| State Management | Zustand |
| Deployment | Render (Backend: Web Service, Frontend: Static Site) |
- Python 3.11+
- Node.js 18+ and npm
- Git
git clone https://github.com/AA-KH/Handwritten-Digits-Classifier.git
cd Handwritten-Digits-Classifier# Navigate to the backend folder
cd backend
# Create and activate a virtual environment (recommended)
python -m venv .venv
source .venv/bin/activate # On Windows: .venv\Scripts\activate
# Install Python dependencies
pip install -r requirements.txt
# Start the API server
uvicorn api:app --host 0.0.0.0 --port 8000You should see:
β
Model weights loaded successfully
kernels: (16, 3, 3), W1: (2704, 64), W2: (64, 10)
INFO: Uvicorn running on http://0.0.0.0:8000
You can verify the backend is healthy at: http://localhost:8000/health
Open a second terminal and run:
# Navigate to the frontend folder
cd frontend
# Install Node dependencies
npm install
# Start the Vite development server
npm run devYou should see:
VITE v5.x.x ready in 300ms
β Local: http://localhost:5173/
Open http://localhost:5173 in your browser. The frontend automatically proxies all /api requests to the backend on port 8000.
If you want to retrain the CNN from scratch, run main.py from inside the backend/ directory:
cd backend
python main.py
β οΈ Note: Training runs 20 epochs with batches of 32 samples on 10,000 training images. The model evaluates on the test set every 5 epochs and saves updated weights tobackend/weights/only when test accuracy improves. Training is CPU-only and may take some time.
Once the backend is running, these endpoints are available:
| Method | Endpoint | Description |
|---|---|---|
GET |
/health |
Health check β confirms model is loaded |
POST |
/predict |
Accepts a base64 PNG, returns predictions + all layer activations |
GET |
/architecture |
Returns static CNN architecture metadata |
- Pre-trained weights are included β the model loads them immediately on startup. No training required to use the app.
- The MNIST dataset is included in
backend/dataset/β needed only if you want to retrain. - The model was trained for 20 epochs on 10,000 images with weights saved every 5 epochs whenever test accuracy improved.
- The CNN is intentionally kept simple (single conv layer) to make every layer's behaviour clearly observable in the visualiser.
main.pyis the original, untouched training script.api.pywraps it for HTTP access without modifying it.
Built with π€ β pure NumPy, no ML frameworks.