A centralized feature store for managing, versioning, and serving ML features at scale. Supports batch and real-time feature computation, storage, and retrieval for production ML pipelines.
- Feature Computation: Spark, Databricks
- Storage: DynamoDB (real-time), S3 (batch)
- Registry: MLflow, Feature Store API
- Serving: REST API, gRPC
- Versioning: Git-based feature definitions
- Monitoring: Drift detection, quality metrics
✅ Feature versioning and reproducibility ✅ Batch and real-time feature computation ✅ Automatic feature serving with low latency ✅ Data drift detection & monitoring ✅ Feature lineage tracking ✅ Multi-environment support (dev/staging/prod) ✅ Easy integration with ML frameworks (sklearn, TensorFlow, PyTorch)
- Core: Python, PySpark
- Storage: DynamoDB, S3, PostgreSQL
- ML Frameworks: MLflow, scikit-learn, TensorFlow
- APIs: FastAPI
- Language: Python
pip install -r requirements.txt
python -m feature_store.init
python scripts/compute_features.py --env prod
python scripts/serve.py # Start API server├── feature_store/
│ ├── core/ # Core feature store logic
│ ├── compute/ # Feature computation engines
│ ├── storage/ # Storage backends
│ ├── serving/ # Feature serving API
│ └── monitoring/ # Drift & quality monitoring
├── features/ # Feature definitions (YAML)
├── scripts/ # Utility scripts
├── tests/ # Tests
└── README.md
MIT