Search your images by meaning, not filename. Drop images into a folder, search "cherry blossoms" later, and find them instantly.
It uses Jina CLIP v2 to encode both images and text into the same vector space, stores embeddings in LanceDB, and provides a web interface to search, upload, and manage your image collection.
No API keys needed. The embedding model runs locally.
- Text-to-image search - describe what you're looking for in natural language
- Image-to-image search - upload a reference image to find visually similar ones
- Auto-indexing - drop images into the
input_images/folder and they're automatically embedded via file watcher - Metadata - add descriptions and source URLs to any image
- CLI search - query your collection from the terminal via Unix socket
- Web UI - React frontend with grid view, image modal (copy/download/find similar), drag-and-drop upload
- Jina CLIP v2 for multimodal embeddings (text + image in same space, 512-dim)
- LanceDB for vector storage and search
- FastAPI for the REST API
- watchdog for filesystem monitoring
- React 19 + Tailwind CSS v4 + shadcn/ui for the frontend
- uv for Python dependency management
-
Clone the repository:
git clone https://github.com/SamIsTheFBI/semantic-image-search.git cd semantic-image-search/ -
Set up the backend:
cd backend uv syncFirst run will download the Jina CLIP v2 model weights (~2 GB). Subsequent runs are instant.
-
Set up the frontend:
cd frontend pnpm install
Start both servers in separate terminals:
cd backend
source .venv/bin/activate
python server.pyThe API server starts on http://localhost:8000.
cd frontend
pnpm devThe dev server starts on http://localhost:5173 and proxies API requests to the backend.
Three ways to add images:
- Web UI - go to the Upload page, drag-and-drop or paste images
- Filesystem - drop files directly into
backend/input_images/. The file watcher picks them up automatically and generates embeddings - API -
POST /uploadwith a multipart file
Supported formats: JPG, PNG, WEBP, GIF, BMP.
-
Text search - type a natural language query like "red car on highway" or "pencil sketch of a cat"
-
Image search - click the image icon in the search bar and upload a reference image
-
Find similar - click any result, then hit "Find similar" in the detail modal
-
CLI - with the backend running, use the CLI client:
cd backend python search.py "sunset over mountains" --limit 5 -v
| Method | Endpoint | Description |
|---|---|---|
GET |
/search?q=...&limit=10 |
Text-to-image semantic search |
POST |
/search-by-image |
Image-to-image similarity search (multipart upload) |
POST |
/upload |
Upload and index an image |
PATCH |
/description/{filename} |
Update description and source URL |
GET |
/images/{filename} |
Serve stored images (static files) |
Edit backend/config.yaml:
jina_clip:
model: "jinaai/jina-clip-v2"
truncate_dim: 512
paths:
input_dir: "input_images"
storage_dir: "data"
search:
default_limit: 10backend/
server.py # uvicorn entrypoint
main.py # standalone mode (Unix socket + file watcher)
search.py # CLI search client
config.yaml # configuration
src/
api/app.py # FastAPI routes (upload, search, metadata)
core/
config.py # YAML config loader
interfaces.py # VectorStore abstract base class
models.py # SearchResult dataclass
search_server.py # Unix socket search server
providers/
jina_clip.py # Jina CLIP v2 embedding provider
storage/
lancedb_store.py # LanceDB vector store implementation
watcher/
file_watcher.py # watchdog-based auto-indexing
frontend/
src/
pages/
SearchPage.tsx # text + image search
UploadPage.tsx # drag-drop / paste upload
components/
SearchBar.tsx # search input + image search button
ResultsGrid.tsx # masonry-style image grid
ImageModal.tsx # detail view (metadata, copy, download, find similar)
UploadZone.tsx # drop zone with inline metadata editing
-
Indexing: images land in
input_images/(via upload or filesystem). The file watcher detects new files, passes them through Jina CLIP v2 to get a 512-dim embedding, and stores the record in LanceDB. -
Text search: user query is encoded via Jina CLIP's text encoder into the same 512-dim space. LanceDB runs vector similarity search against stored image embeddings.
-
Image search: uploaded reference image is encoded via Jina CLIP's image encoder. Same vector search follows.
-
Results: frontend displays matches in a grid. Clicking an image shows metadata (dimensions, file size, match score), and lets you edit description/source URL, copy/download, or find similar images.
Beginning of my 11th grade, I received my first smartphone as I went to hostel and on it I used Pinterest a lot. I was fascinated by how it got me just the image I wanted. And similar images wasn't just for namesake it actually worked correctly (now every site's got that correctly but at the time this was magic to me). Used it for so many sketching hours. Then I discovered Booru image board sites and I slowly got accustomed to just finding things by proper tags. So, anyway I wanted to replicate that and have my own Pinterest-Booru kinda thing. Will add tags to this someday when it gets less tricky to handle.