A Modern ETL data Pipeline using
- Dagster
- Python
- Pandas
- PyTest
- Automate Lieage and Metadata Tracking
This project demonstrates a modern ETL (Extract, Transform, Load) data pipeline using Dagster, Python, and Pandas. The pipeline is designed to efficiently process and transform data, making it ready for analysis and reporting. Dagster is used as the orchestration tool to manage the workflow, ensuring that each step of the pipeline is executed in the correct order and handling any dependencies between tasks. Python and Pandas are utilized for data manipulation and transformation, providing powerful tools for cleaning, aggregating, and analyzing data. This project serves as a template for building robust and scalable data pipelines.
The ETL framework has been designed to process both fact and dimension data using a madelion architecture.
Aims to Ingest Fact Data from Various Source System
Keeps Fact data cleansed and transformed for future Processing
In order to ingest Dimension Data we have a seperate module called dim_ingestion
Likewise, Fact Transformation we have Dimension processing.
A curation in a merged Layer between Fact and Dimension Data whcih can be further served for Reporting or application to consume
- database Files are saved Under Data folder
- Clone this repo - git clone https://github.com/dipanjannet/dagster_pipeline.git
- Create a Virtual Environment
- python -m venv dagster_tutorial
- Navigate to that Location : cd dagster_pipeline\dagster_tutorial\Scripts
- Activate the Virtual Environment : .\Activate.ps1
- Install necessary deedency : pip install dagster dagster-webserver pandas pytest
- Navigate to : cd .\data_pipeline\
- Run : dagster dev
A Curation Layer will be a VIEW or Materialized Table(s) Based on the use Case / Access Control
All Data Assets are Monitored through automated checks and access is controlled via RBAC, Row level Access control etc.
External Sources --> Ingestion --> Dimension Data --> Transformation --> Processed Dimension Data --> Curation
External Sources --> Ingestion --> Fact Data --> Transformation --> Processed Fact Data --> Curation
Curation Layer --> Access Layer --> Analytics, Reporting, Applications Curation Layer --> Access Layer --> DuckDB Curation Layer --> Access Layer --> Spark Curation Layer --> Access Layer --> Trino

