Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Formula 1 Telemetry Big Data Pipeline

An end-to-end analytics project that transforms Formula 1 telemetry into clean performance indicators, machine-learning predictions, driving-behaviour clusters, and Power BI-ready reporting tables.

The project uses a Bronze-Silver-Gold architecture in PySpark: raw race telemetry is ingested and profiled, validated and enriched, then aggregated into business-ready tables. It also classifies speed categories and discovers driving patterns with unsupervised learning.

Project contents

Path Purpose
F1 Complete Pipeline/F1_Pipeline.ipynb Main Databricks/PySpark notebook: ingestion, transformation, EDA, ML, clustering, and exports.
F1 Complete Pipeline/formula1_race.zip Compressed Formula 1 telemetry dataset. Extract the CSV before running the notebook.
F1 Complete Pipeline/Dashboard/Formula1_Dashboard.pbix Power BI dashboard built from the analytical outputs.
F1 Complete Pipeline/f1_report.pdf Project report.

Architecture

Formula 1 telemetry CSV
        |
        v
Bronze: raw Spark DataFrame, schema and quality checks
        |
        v
Silver: deduplication, validation, cleaning and derived features
        |
        +-----------------------+
        |                       |
        v                       v
Gold: KPI tables          ML and clustering
        |                       |
        +-----------+-----------+
                    v
        Power BI-ready CSV exports and dashboard

Methods used

Data engineering

  • Bronze-Silver-Gold (medallion) pipeline: separates raw, cleaned, and consumption-ready data.
  • Spark CSV ingestion and SQL views: loads the source as distributed DataFrames and exposes temporary views for SQL exploration.
  • Data quality profiling: inspects schema, volume, cardinality, and missing values.
  • Data cleaning and validation: removes duplicate/incomplete rows and retains realistic telemetry ranges for speed, RPM, gear, throttle, brake, DRS, and lap number.
  • Feature engineering: creates speed_category (Low / Medium / High), braking_status, and drs_status.
  • Gold-layer aggregation: summarizes performance by driver, race, and session using telemetry-record counts and speed, RPM, throttle, braking, and DRS KPIs.

Exploratory data analysis

  • Descriptive statistics: count, mean, minimum, maximum, and standard deviation.
  • Grouped performance analysis by driver, race, session, gear, speed category, braking state, and DRS state.
  • Pearson correlations between RPM and speed, throttle and speed, and brake and speed.

Supervised machine learning

The classification task predicts the engineered speed_category from RPM, gear, throttle, brake, DRS, and accelerometer features (acc_x, acc_y, acc_z).

  • Label encoding: StringIndexer converts speed-category labels to numeric labels.
  • Feature assembly: VectorAssembler creates Spark ML feature vectors.
  • Holdout validation: reproducible 80/20 train-test split (seed=42).
  • Models: Logistic Regression, Decision Tree Classifier, and Random Forest Classifier (50 trees, seed=42).
  • Evaluation: accuracy, weighted precision, weighted recall, F1 score, prediction distributions, and Random Forest feature importance.

Unsupervised machine learning

  • Standardization: StandardScaler with mean centering and standard-deviation scaling.
  • K-Means clustering: three clusters (k=3, seed=42) built from RPM, speed, gear, throttle, brake, DRS, and accelerometer measurements.
  • Cluster evaluation: silhouette score, cluster-size distribution, and cluster profiles.
  • Cluster interpretation: rule-based labels identify high-speed/aggressive, frequent-braking, and balanced driving behaviour.

Technologies

  • Python
  • Apache Spark / PySpark and Spark SQL
  • Databricks
  • Spark MLlib
  • Power BI

Running the notebook

  1. Extract F1 Complete Pipeline/formula1_race.zip and upload the CSV to your Databricks workspace.
  2. Open F1 Complete Pipeline/F1_Pipeline.ipynb in Databricks.
  3. Update CSV_PATH (and the matching CSV read path) to the uploaded dataset location.
  4. Run the cells in order. The notebook creates and uses the formula1 database, builds temporary views, trains the models, and exports dashboard tables.
  5. Connect F1 Complete Pipeline/Dashboard/Formula1_Dashboard.pbix to the exported CSV files, if paths differ from the original environment.

Dashboard outputs

The notebook exports these CSV tables for reporting:

  • driver_performance.csv
  • race_performance.csv
  • session_analysis.csv
  • model_evaluation.csv
  • feature_importance.csv
  • cluster_analysis.csv

About

End-to-end Formula 1 telemetry analytics pipeline using PySpark, Spark MLlib, Databricks, and Power BI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Contributors

Languages