A Databricks Lakehouse project for integrating child-company sales data into a parent-company FMCG analytics star schema using PySpark, Delta Lake, AWS S3, and Databricks Workflows.
-
Updated
Jul 5, 2026 - Jupyter Notebook
A Databricks Lakehouse project for integrating child-company sales data into a parent-company FMCG analytics star schema using PySpark, Delta Lake, AWS S3, and Databricks Workflows.
End-to-end Azure data pipeline (ADF → ADLS → Databricks/PySpark → Synapse Serverless → Power BI) using Medallion architecture.
End-to-end Data Engineering project using Databricks to build a Bike Data Lakehouse with Bronze, Silver, and Gold layers
Scalable movie recommender system built with Apache Spark and Item-Based Collaborative Filtering using the MovieLens 1M dataset.
End-to-end Formula 1 telemetry analytics pipeline using PySpark, Spark MLlib, Databricks, and Power BI.
Data Engineer focused on building scalable, reliable, and production-ready data pipelines.
Production-style Real-Time IoT Streaming Lakehouse built with Databricks, PySpark, Delta Lake and Structured Streaming, featuring Medallion Architecture, Watermarking, Window Aggregations, Anomaly Detection, Monitoring and Performance Optimization.
End-to-end educational data engineering pipeline using Python, PySpark, data cleaning, transformations, window functions, and Parquet output.
Một hệ thống giám sát và phát hiện gian lận thẻ tín dụng thời gian thực sử dụng Apache Spark, Machine Learning và Streamlit Dashboard.
To associate your repository with the apche-spark topic, visit your repo's landing page and select "manage topics."