Skip to content

Latest commit

ย 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

Customer Churn Prediction Pipeline

๐ŸŽฏ Project Overview

A comprehensive Data Engineering and Machine Learning Pipeline designed to predict customer churn in telecommunications companies. This solution enables proactive customer retention strategies through data-driven insights and predictive analytics.


๐Ÿ—๏ธ Architecture Overview

The pipeline follows a stage-based architecture with clear separation of concerns:

Raw Data โ†’ Ingestion โ†’ Validation โ†’ Storage โ†’ Feature Engineering โ†’ ML Training โ†’ Model Deployment
   โ†“           โ†“          โ†“          โ†“            โ†“              โ†“              โ†“
Kaggle     Download    Quality   Parquet     Features      Models       API
Files      & Extract   Checks    Files       & Store       & Reports    Endpoints

Pipeline Stages

  1. ๐Ÿ“ฅ Data Ingestion - Download and extract data from multiple sources
  2. ๐Ÿ” Data Validation - Ensure data quality and business rule compliance
  3. ๐Ÿ’พ Data Storage - Efficient storage with versioning and metadata
  4. ๐Ÿ”ง Feature Engineering - Transform raw data into ML-ready features
  5. ๐Ÿค– Machine Learning - Train and evaluate churn prediction models
  6. ๐Ÿš€ Model Deployment - Deploy models for production use

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.8+
  • Required packages: pandas, numpy, scikit-learn, prefect

1. Test the Demo Pipeline

python test_demo_pipeline.py

2. Run Demo ML Training Only

python run_ml_training_demo.py

3. Run Complete Pipeline (with Demo ML Training)

python main.py

๐Ÿ“ Project Structure

DMML/
โ”œโ”€โ”€ ๐Ÿ“‹ business_requirements/          # Business documentation
โ”œโ”€โ”€ ๐Ÿ“š docs/                          # Technical documentation
โ”œโ”€โ”€ ๐Ÿ”ง src/                           # Source code by pipeline stage
โ”‚   โ”œโ”€โ”€ data_ingestion/               # Stage 1: Data Ingestion
โ”‚   โ”œโ”€โ”€ data_validation/              # Stage 2: Data Validation
โ”‚   โ”œโ”€โ”€ data_storage/                 # Stage 3: Data Storage
โ”‚   โ”œโ”€โ”€ feature_engineering/          # Stage 4: Feature Engineering
โ”‚   โ”œโ”€โ”€ ml_training/                  # Stage 5: Machine Learning
โ”‚   โ””โ”€โ”€ model_deployment/             # Stage 6: Model Deployment
โ”œโ”€โ”€ ๐Ÿช feature_store/                 # Feature storage and metadata
โ”œโ”€โ”€ ๐Ÿ“Š my_data_lake/                  # Raw data storage
โ”œโ”€โ”€ ๐Ÿ”– version_control/               # Dataset and model versioning
โ”œโ”€โ”€ ๐Ÿค– models/                        # Trained model artifacts
โ”œโ”€โ”€ ๐Ÿงช tests/                         # Test suite
โ””โ”€โ”€ ๐Ÿš€ main.py                        # Main pipeline orchestration

๐Ÿ“Š Data Sources

Primary Datasets

  1. Kaggle - blastchar/telco-customer-churn

    • Format: CSV
    • Records: ~7,000 customers
    • Content: Demographics, services, billing, churn status
  2. Kaggle - abdallahwagih/telco-customer-churn

    • Format: XLSX
    • Records: ~7,000 customers
    • Content: Extended customer attributes and behavioral data

Data Quality Requirements

  • Completeness: >95% for critical fields, >80% for others
  • Accuracy: >98% for financial fields, >90% for categorical
  • Consistency: Unified formats across data sources

๐ŸŽฏ Expected Outputs

1. Clean Datasets for EDA

  • Validated raw data with quality reports
  • Combined dataset from multiple sources
  • Data dictionary and metadata

2. Transformed Features for ML

  • 50+ engineered features across categories:
    • Demographic: Age groups, family size, location clusters
    • Service: Service bundles, contract duration, upgrade history
    • Behavioral: Usage patterns, payment behavior, support interactions
    • Financial: Revenue trends, payment reliability, price sensitivity

3. Deployable Churn Prediction Models

  • Models: Logistic Regression, Random Forest
  • Performance: >75% F1-score target
  • Outputs: Serialized models, preprocessors, performance reports

๐Ÿ“ˆ Business Impact

Primary Objectives

  • Reduce Churn Rate: Target 15-25% reduction within 12 months
  • Improve ROI: Achieve 3:1 ROI on retention initiatives
  • Enhance Customer Experience: Improve satisfaction scores by 20%

Success Metrics

  • Churn Rate: Monthly tracking with trend analysis
  • Revenue Retention: Monthly recurring revenue retained
  • Customer Lifetime Value: Average revenue per customer
  • Operational Efficiency: Support cost reduction

๐Ÿ”ง Configuration

ML Training Configuration

Edit ml_training_config.py to customize:

ML_TRAINING_CONFIG = {
    "training_mode": "basic",           # Always basic for demo
    "feature_selection_k": 10,          # Number of features
    "timeout_minutes": 10,              # Execution timeout
    "enable_mlflow": False,             # Disabled for demo
    "models_to_train": ["logistic_regression", "random_forest"],
    "cross_validation_folds": 3,        # Reduced for speed
    "enable_hyperparameter_tuning": False,  # Disabled for demo
}

Pipeline Configuration

  • Data Sources: Configure in main.py
  • Feature Engineering: Settings in feature store configuration
  • Model Training: Parameters in ML training config
  • Output Paths: Configurable storage locations

๐Ÿงช Testing

Test Suite Organization

tests/
โ”œโ”€โ”€ unit/                    # Unit tests for individual components
โ”œโ”€โ”€ integration/             # Integration tests for stage interactions
โ”œโ”€โ”€ e2e/                    # End-to-end pipeline tests
โ””โ”€โ”€ fixtures/               # Test data and fixtures

Run Tests

# Run all tests
python -m pytest tests/

# Run specific test categories
python -m pytest tests/unit/
python -m pytest tests/integration/
python -m pytest tests/e2e/

# Run with coverage
python -m pytest --cov=src tests/

๐Ÿ“š Documentation

Business Documentation

  • business_requirements/customer_churn_prediction_requirements.md: Business problem, objectives, success criteria

Technical Documentation

  • docs/technical_architecture.md: System design and architecture
  • docs/project_structure.md: Code organization and structure
  • docs/api_reference.md: API documentation and usage

User Guides

  • DEMO_README.md: Demo pipeline usage and configuration
  • FEATURE_ENGINEERING_README.md: Feature engineering details
  • VERSION_CONTROL_DOCUMENTATION.md: Version control system usage

๐Ÿšจ Troubleshooting

Common Issues

1. Feature Store Not Found

# Ensure you've run the main pipeline first
python main.py

2. Dependencies Missing

pip install scikit-learn pandas numpy prefect

3. Timeout Issues

# Increase timeout in ml_training_config.py
ML_TRAINING_CONFIG["timeout_minutes"] = 15

4. Pipeline Order Issues

# Verify pipeline order is correct
python test_pipeline_order.py

Debug Mode

# Enable verbose logging
python -c "
import logging
logging.basicConfig(level=logging.DEBUG)
from run_ml_training_demo import SimpleChurnTrainer
trainer = SimpleChurnTrainer()
"

๐Ÿ”„ Development Workflow

Adding New Features

  1. Identify Stage: Determine which pipeline stage the feature belongs to
  2. Create Component: Add new component to appropriate stage folder
  3. Update Tests: Add unit and integration tests
  4. Update Documentation: Document new functionality
  5. Integration: Ensure compatibility with other stages

Code Quality Standards

  • Type Hints: Use Python type hints for all functions
  • Documentation: Comprehensive docstrings for all classes and methods
  • Testing: >90% test coverage for critical components
  • Linting: Follow PEP 8 and project-specific style guidelines

๐Ÿš€ Deployment

Development Environment

  • Local Development: Docker containers for consistency
  • Version Control: Git with feature branch workflow
  • Code Quality: Linting, formatting, and pre-commit hooks

Production Environment

  • Containerization: Docker containers for easy deployment
  • Orchestration: Kubernetes for container management
  • CI/CD: Automated testing and deployment pipelines
  • Monitoring: Comprehensive logging and metrics collection

๐Ÿ“ž Support and Contributing

Getting Help

  1. Check Documentation: Review relevant documentation files
  2. Run Tests: Verify system functionality with test suite
  3. Check Issues: Look for similar problems in issue tracker
  4. Create Issue: Report bugs or request features

Contributing

  1. Fork Repository: Create your own fork
  2. Create Branch: Work on feature or bug fix
  3. Add Tests: Include tests for new functionality
  4. Submit PR: Create pull request with description

๐Ÿ“Š Performance Metrics

Pipeline Performance

  • End-to-End Execution: <30 minutes target
  • Data Processing: <10 minutes for 10K records
  • Model Training: <5 minutes for demo models
  • API Response: <100ms for real-time predictions

Model Performance

  • Accuracy: >80% target
  • Precision: >75% for churn prediction
  • Recall: >70% for churn detection
  • F1-Score: >75% target

๐Ÿ”ฎ Future Enhancements

Short-term (3-6 months)

  • Advanced feature engineering algorithms
  • Hyperparameter tuning and optimization
  • Model performance monitoring
  • Real-time prediction API

Medium-term (6-12 months)

  • Additional ML algorithms (XGBoost, Neural Networks)
  • Automated model retraining
  • A/B testing framework
  • Business intelligence dashboards

Long-term (12+ months)

  • Multi-tenant architecture
  • Advanced analytics and insights
  • Integration with CRM systems
  • Predictive maintenance capabilities

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.


๐Ÿ™ Acknowledgments

  • Kaggle: For providing the customer churn datasets
  • Open Source Community: For the excellent tools and libraries
  • Contributors: All those who have contributed to this project

This Customer Churn Prediction Pipeline provides a robust, scalable foundation for implementing data-driven customer retention strategies. The modular architecture ensures maintainability and extensibility for future enhancements.

About

No description or website provided.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages