Real-time depth estimation API powered by Depth Anything V2. Upload an image, get a depth map back. Runs on GPU (CUDA/TensorRT) or CPU.
Built as a production-ready microservice — containerized, async, and optimized for throughput.
- Multiple models — Small (25M), Base (98M), and Large (335M) variants
- Multiple output formats — Raw depth (NumPy), grayscale PNG, colored depth map, 3D point cloud
- Batch processing — Submit multiple images in one request
- Streaming — MJPEG stream endpoint for real-time video depth
- GPU accelerated — CUDA with optional TensorRT optimization
- CPU fallback — Runs anywhere, just slower
- Docker ready — GPU and CPU Dockerfiles included
- Async — Built on FastAPI with async I/O
# Clone
git clone https://github.com/KyleBuildsAI/depth-stream.git
cd depth-stream
# Install
pip install -r requirements.txt
# Run (downloads model on first launch)
python -m depth_stream
# Or with Docker (GPU)
docker compose up depth-stream-gpuThe API is now at http://localhost:8000.
Estimate depth from a single image.
curl -X POST http://localhost:8000/depth \
-F "file=@photo.jpg" \
-F "output_format=colored" \
-o depth_map.pngParameters:
| Name | Type | Default | Description |
|---|---|---|---|
file |
file | required | Input image (JPEG, PNG, WebP) |
output_format |
string | colored |
raw, grayscale, colored, pointcloud |
model_size |
string | base |
small, base, large |
max_depth |
float | None |
Clamp maximum depth value |
Process multiple images in one request.
curl -X POST http://localhost:8000/depth/batch \
-F "files=@img1.jpg" \
-F "files=@img2.jpg" \
-F "output_format=grayscale" \
-o results.zipMJPEG stream for real-time video depth estimation. Connect a webcam feed or point a video player at this endpoint.
http://localhost:8000/depth/stream?source=0&model_size=small
Health check with model status and GPU info.
curl http://localhost:8000/health| Env Variable | Default | Description |
|---|---|---|
DEPTH_HOST |
0.0.0.0 |
Bind address |
DEPTH_PORT |
8000 |
Port |
DEPTH_MODEL |
base |
Default model size (small, base, large) |
DEPTH_DEVICE |
auto |
cuda, cpu, or auto (detect) |
DEPTH_HALF |
true |
Use FP16 on GPU (faster, less VRAM) |
DEPTH_WORKERS |
1 |
Uvicorn workers |
DEPTH_CACHE_DIR |
~/.cache/depth-stream |
Model download cache |
docker compose up depth-stream-gpuRequires NVIDIA Container Toolkit.
docker compose up depth-stream-cpudepth-stream/
├── depth_stream/
│ ├── __init__.py
│ ├── __main__.py # Entry point
│ ├── api.py # FastAPI routes
│ ├── config.py # Settings
│ ├── models.py # Model loading and management
│ ├── estimator.py # Depth estimation pipeline
│ ├── outputs.py # Output format converters
│ └── stream.py # MJPEG streaming
├── static/
│ └── index.html # Demo web UI
├── scripts/
│ └── benchmark.py # Performance benchmarking
├── Dockerfile.gpu
├── Dockerfile.cpu
├── docker-compose.yml
├── requirements.txt
└── README.md
| Model | Parameters | Speed (GPU) | VRAM | Quality |
|---|---|---|---|---|
small |
25M | ~15ms | ~0.5 GB | Good |
base |
98M | ~30ms | ~1.0 GB | Better |
large |
335M | ~80ms | ~2.5 GB | Best |
Models are downloaded automatically from HuggingFace on first use.
Benchmarked on RTX 3090, 640x480 input:
| Model | FP32 | FP16 | TensorRT |
|---|---|---|---|
| Small | 67 FPS | 95 FPS | 120 FPS |
| Base | 33 FPS | 52 FPS | 71 FPS |
| Large | 12 FPS | 19 FPS | 28 FPS |
Built by Kyle Coleman — drawing on years of work in real-time 2D-to-3D conversion, depth estimation, and stereoscopic rendering for VR.
MIT