An end-to-end analytics project that transforms Formula 1 telemetry into clean performance indicators, machine-learning predictions, driving-behaviour clusters, and Power BI-ready reporting tables.
The project uses a Bronze-Silver-Gold architecture in PySpark: raw race telemetry is ingested and profiled, validated and enriched, then aggregated into business-ready tables. It also classifies speed categories and discovers driving patterns with unsupervised learning.
| Path | Purpose |
|---|---|
F1 Complete Pipeline/F1_Pipeline.ipynb |
Main Databricks/PySpark notebook: ingestion, transformation, EDA, ML, clustering, and exports. |
F1 Complete Pipeline/formula1_race.zip |
Compressed Formula 1 telemetry dataset. Extract the CSV before running the notebook. |
F1 Complete Pipeline/Dashboard/Formula1_Dashboard.pbix |
Power BI dashboard built from the analytical outputs. |
F1 Complete Pipeline/f1_report.pdf |
Project report. |
Formula 1 telemetry CSV
|
v
Bronze: raw Spark DataFrame, schema and quality checks
|
v
Silver: deduplication, validation, cleaning and derived features
|
+-----------------------+
| |
v v
Gold: KPI tables ML and clustering
| |
+-----------+-----------+
v
Power BI-ready CSV exports and dashboard
- Bronze-Silver-Gold (medallion) pipeline: separates raw, cleaned, and consumption-ready data.
- Spark CSV ingestion and SQL views: loads the source as distributed DataFrames and exposes temporary views for SQL exploration.
- Data quality profiling: inspects schema, volume, cardinality, and missing values.
- Data cleaning and validation: removes duplicate/incomplete rows and retains realistic telemetry ranges for speed, RPM, gear, throttle, brake, DRS, and lap number.
- Feature engineering: creates
speed_category(Low / Medium / High),braking_status, anddrs_status. - Gold-layer aggregation: summarizes performance by driver, race, and session using telemetry-record counts and speed, RPM, throttle, braking, and DRS KPIs.
- Descriptive statistics: count, mean, minimum, maximum, and standard deviation.
- Grouped performance analysis by driver, race, session, gear, speed category, braking state, and DRS state.
- Pearson correlations between RPM and speed, throttle and speed, and brake and speed.
The classification task predicts the engineered speed_category from RPM, gear, throttle, brake, DRS, and accelerometer features (acc_x, acc_y, acc_z).
- Label encoding:
StringIndexerconverts speed-category labels to numeric labels. - Feature assembly:
VectorAssemblercreates Spark ML feature vectors. - Holdout validation: reproducible 80/20 train-test split (
seed=42). - Models: Logistic Regression, Decision Tree Classifier, and Random Forest Classifier (50 trees,
seed=42). - Evaluation: accuracy, weighted precision, weighted recall, F1 score, prediction distributions, and Random Forest feature importance.
- Standardization:
StandardScalerwith mean centering and standard-deviation scaling. - K-Means clustering: three clusters (
k=3,seed=42) built from RPM, speed, gear, throttle, brake, DRS, and accelerometer measurements. - Cluster evaluation: silhouette score, cluster-size distribution, and cluster profiles.
- Cluster interpretation: rule-based labels identify high-speed/aggressive, frequent-braking, and balanced driving behaviour.
- Python
- Apache Spark / PySpark and Spark SQL
- Databricks
- Spark MLlib
- Power BI
- Extract
F1 Complete Pipeline/formula1_race.zipand upload the CSV to your Databricks workspace. - Open
F1 Complete Pipeline/F1_Pipeline.ipynbin Databricks. - Update
CSV_PATH(and the matching CSV read path) to the uploaded dataset location. - Run the cells in order. The notebook creates and uses the
formula1database, builds temporary views, trains the models, and exports dashboard tables. - Connect
F1 Complete Pipeline/Dashboard/Formula1_Dashboard.pbixto the exported CSV files, if paths differ from the original environment.
The notebook exports these CSV tables for reporting:
driver_performance.csvrace_performance.csvsession_analysis.csvmodel_evaluation.csvfeature_importance.csvcluster_analysis.csv