A comprehensive Data Engineering and Machine Learning Pipeline designed to predict customer churn in telecommunications companies. This solution enables proactive customer retention strategies through data-driven insights and predictive analytics.
The pipeline follows a stage-based architecture with clear separation of concerns:
Raw Data โ Ingestion โ Validation โ Storage โ Feature Engineering โ ML Training โ Model Deployment
โ โ โ โ โ โ โ
Kaggle Download Quality Parquet Features Models API
Files & Extract Checks Files & Store & Reports Endpoints
- ๐ฅ Data Ingestion - Download and extract data from multiple sources
- ๐ Data Validation - Ensure data quality and business rule compliance
- ๐พ Data Storage - Efficient storage with versioning and metadata
- ๐ง Feature Engineering - Transform raw data into ML-ready features
- ๐ค Machine Learning - Train and evaluate churn prediction models
- ๐ Model Deployment - Deploy models for production use
- Python 3.8+
- Required packages:
pandas,numpy,scikit-learn,prefect
python test_demo_pipeline.pypython run_ml_training_demo.pypython main.pyDMML/
โโโ ๐ business_requirements/ # Business documentation
โโโ ๐ docs/ # Technical documentation
โโโ ๐ง src/ # Source code by pipeline stage
โ โโโ data_ingestion/ # Stage 1: Data Ingestion
โ โโโ data_validation/ # Stage 2: Data Validation
โ โโโ data_storage/ # Stage 3: Data Storage
โ โโโ feature_engineering/ # Stage 4: Feature Engineering
โ โโโ ml_training/ # Stage 5: Machine Learning
โ โโโ model_deployment/ # Stage 6: Model Deployment
โโโ ๐ช feature_store/ # Feature storage and metadata
โโโ ๐ my_data_lake/ # Raw data storage
โโโ ๐ version_control/ # Dataset and model versioning
โโโ ๐ค models/ # Trained model artifacts
โโโ ๐งช tests/ # Test suite
โโโ ๐ main.py # Main pipeline orchestration
-
Kaggle - blastchar/telco-customer-churn
- Format: CSV
- Records: ~7,000 customers
- Content: Demographics, services, billing, churn status
-
Kaggle - abdallahwagih/telco-customer-churn
- Format: XLSX
- Records: ~7,000 customers
- Content: Extended customer attributes and behavioral data
- Completeness: >95% for critical fields, >80% for others
- Accuracy: >98% for financial fields, >90% for categorical
- Consistency: Unified formats across data sources
- Validated raw data with quality reports
- Combined dataset from multiple sources
- Data dictionary and metadata
- 50+ engineered features across categories:
- Demographic: Age groups, family size, location clusters
- Service: Service bundles, contract duration, upgrade history
- Behavioral: Usage patterns, payment behavior, support interactions
- Financial: Revenue trends, payment reliability, price sensitivity
- Models: Logistic Regression, Random Forest
- Performance: >75% F1-score target
- Outputs: Serialized models, preprocessors, performance reports
- Reduce Churn Rate: Target 15-25% reduction within 12 months
- Improve ROI: Achieve 3:1 ROI on retention initiatives
- Enhance Customer Experience: Improve satisfaction scores by 20%
- Churn Rate: Monthly tracking with trend analysis
- Revenue Retention: Monthly recurring revenue retained
- Customer Lifetime Value: Average revenue per customer
- Operational Efficiency: Support cost reduction
Edit ml_training_config.py to customize:
ML_TRAINING_CONFIG = {
"training_mode": "basic", # Always basic for demo
"feature_selection_k": 10, # Number of features
"timeout_minutes": 10, # Execution timeout
"enable_mlflow": False, # Disabled for demo
"models_to_train": ["logistic_regression", "random_forest"],
"cross_validation_folds": 3, # Reduced for speed
"enable_hyperparameter_tuning": False, # Disabled for demo
}- Data Sources: Configure in
main.py - Feature Engineering: Settings in feature store configuration
- Model Training: Parameters in ML training config
- Output Paths: Configurable storage locations
tests/
โโโ unit/ # Unit tests for individual components
โโโ integration/ # Integration tests for stage interactions
โโโ e2e/ # End-to-end pipeline tests
โโโ fixtures/ # Test data and fixtures
# Run all tests
python -m pytest tests/
# Run specific test categories
python -m pytest tests/unit/
python -m pytest tests/integration/
python -m pytest tests/e2e/
# Run with coverage
python -m pytest --cov=src tests/business_requirements/customer_churn_prediction_requirements.md: Business problem, objectives, success criteria
docs/technical_architecture.md: System design and architecturedocs/project_structure.md: Code organization and structuredocs/api_reference.md: API documentation and usage
DEMO_README.md: Demo pipeline usage and configurationFEATURE_ENGINEERING_README.md: Feature engineering detailsVERSION_CONTROL_DOCUMENTATION.md: Version control system usage
# Ensure you've run the main pipeline first
python main.pypip install scikit-learn pandas numpy prefect# Increase timeout in ml_training_config.py
ML_TRAINING_CONFIG["timeout_minutes"] = 15# Verify pipeline order is correct
python test_pipeline_order.py# Enable verbose logging
python -c "
import logging
logging.basicConfig(level=logging.DEBUG)
from run_ml_training_demo import SimpleChurnTrainer
trainer = SimpleChurnTrainer()
"- Identify Stage: Determine which pipeline stage the feature belongs to
- Create Component: Add new component to appropriate stage folder
- Update Tests: Add unit and integration tests
- Update Documentation: Document new functionality
- Integration: Ensure compatibility with other stages
- Type Hints: Use Python type hints for all functions
- Documentation: Comprehensive docstrings for all classes and methods
- Testing: >90% test coverage for critical components
- Linting: Follow PEP 8 and project-specific style guidelines
- Local Development: Docker containers for consistency
- Version Control: Git with feature branch workflow
- Code Quality: Linting, formatting, and pre-commit hooks
- Containerization: Docker containers for easy deployment
- Orchestration: Kubernetes for container management
- CI/CD: Automated testing and deployment pipelines
- Monitoring: Comprehensive logging and metrics collection
- Check Documentation: Review relevant documentation files
- Run Tests: Verify system functionality with test suite
- Check Issues: Look for similar problems in issue tracker
- Create Issue: Report bugs or request features
- Fork Repository: Create your own fork
- Create Branch: Work on feature or bug fix
- Add Tests: Include tests for new functionality
- Submit PR: Create pull request with description
- End-to-End Execution: <30 minutes target
- Data Processing: <10 minutes for 10K records
- Model Training: <5 minutes for demo models
- API Response: <100ms for real-time predictions
- Accuracy: >80% target
- Precision: >75% for churn prediction
- Recall: >70% for churn detection
- F1-Score: >75% target
- Advanced feature engineering algorithms
- Hyperparameter tuning and optimization
- Model performance monitoring
- Real-time prediction API
- Additional ML algorithms (XGBoost, Neural Networks)
- Automated model retraining
- A/B testing framework
- Business intelligence dashboards
- Multi-tenant architecture
- Advanced analytics and insights
- Integration with CRM systems
- Predictive maintenance capabilities
This project is licensed under the MIT License - see the LICENSE file for details.
- Kaggle: For providing the customer churn datasets
- Open Source Community: For the excellent tools and libraries
- Contributors: All those who have contributed to this project
This Customer Churn Prediction Pipeline provides a robust, scalable foundation for implementing data-driven customer retention strategies. The modular architecture ensures maintainability and extensibility for future enhancements.