Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Formula 1 Race Win Prediction using Machine Learning

Project Overview

This project analyzes historical Formula 1 race data from 1950 to 2024 and explores whether historical driver and constructor performance can be used to predict race wins.

The project follows a complete data science workflow:

  • Import and inspect multiple CSV datasets
  • Perform Exploratory Data Analysis (EDA)
  • Merge and combine data from different sources
  • Create historical performance features through Feature Engineering
  • Prepare the data for machine learning
  • Train classification models
  • Optimize model hyperparameters using GridSearchCV
  • Optimize classification thresholds
  • Compare model performance using precision, recall, F1-score and accuracy

The main objective is not only to build a predictive model, but also to understand which factors are most strongly associated with winning a Formula 1 race.


Dataset

The project uses the Formula 1 World Championship 1950–2020 dataset by Rohan Rao, extended with the available race data through 2024.

The dataset consists of multiple CSV files covering different aspects of Formula 1 history.

Main datasets used include:

  • Races
  • Drivers
  • Constructors
  • Results
  • Qualifying
  • Circuits
  • Driver standings
  • Constructor standings
  • Lap times
  • Pit stops
  • Sprint results
  • Status

The use of multiple related datasets makes it possible to combine race, driver, constructor and circuit information.


Exploratory Data Analysis (EDA)

The exploratory analysis investigates historical patterns within Formula 1.

Key areas analyzed include:

  • Number of races over time
  • Number of participating drivers
  • Race distribution by country
  • Most frequently used circuits
  • Grand Prix frequency
  • Driver victories
  • Constructor victories
  • Nationality of race winners
  • Starting grid positions
  • Position changes during races
  • Relationship between pole position and race wins

Key Findings

Pole Position

Starting from pole position provides a significant advantage.

However, pole position does not guarantee victory.

Approximately 42.34% of pole positions resulted in a race win in the analyzed dataset.

This demonstrates that qualifying performance is an important factor, but race results are influenced by additional factors.

Position Changes

Drivers frequently gained or lost positions between their starting grid position and their final race position.

This highlights that starting position and final race result are related, but not identical.

Successful Constructors

Historically successful constructors such as Ferrari, McLaren, Mercedes and Red Bull account for a large share of Formula 1 race victories and championship points.

Circuits and Countries

A relatively small number of circuits and countries have hosted a large proportion of Formula 1 races.

Monza, Monaco and Silverstone are among the circuits with the highest number of hosted races.


Feature Engineering

The Feature Engineering stage transforms historical race information into variables that can be used by machine learning models.

A central principle of the feature engineering process is avoiding data leakage.

Historical performance features are calculated using information available before the current race.

This allows the machine learning model to simulate a realistic prediction scenario.

Driver Features

The following historical driver features were created:

  • previous_races
  • previous_wins
  • previous_win_rate
  • previous_points
  • previous_avg_position
  • previous_podiums
  • previous_podium_rate

These features describe the driver's historical performance before the current race.

Constructor Features

Historical constructor performance was also incorporated:

  • previous_constructor_wins
  • previous_constructor_points
  • previous_constructor_races
  • previous_constructor_win_rate

Race Features

Current race-related features include:

  • grid
  • pole_position
  • year

The target variable is:

  • win

where:

  • 0 = driver did not win
  • 1 = driver won the race

Machine Learning Approach

Prediction Task

The machine learning task is a binary classification problem.

The objective is:

Predict whether a Formula 1 driver will win a given race.

The target variable is highly imbalanced because only a small proportion of all race results represent victories.

Approximately:

  • 95.78% → no win
  • 4.22% → win

Because of this imbalance, accuracy alone is not sufficient for evaluating model performance.

Particular attention is therefore given to:

  • Precision
  • Recall
  • F1-score

for the winning class.


Train/Test Split

Because Formula 1 is a chronological dataset, a random train/test split was avoided.

Instead, a time-based split was used:

Dataset Years
Training 1950–2019
Testing 2020–2024

This approach simulates a realistic prediction scenario where historical data is used to predict future races.

The final datasets contain:

  • 26,759 observations
  • 14 machine learning features

Data Preprocessing

The feature previous_avg_position contains missing values for drivers who had not participated in a previous race.

These missing values were handled using:

SimpleImputer with median strategy

The imputer was fitted only on the training data and then applied to the test data.

This prevents information from the test period from influencing the preprocessing process.


Baseline Model

A DummyClassifier was used as a baseline model.

The baseline always predicts the most frequent class.

Although this produces a relatively high accuracy due to the class imbalance, it fails to identify race winners.

This demonstrates why accuracy alone is not an appropriate metric for this prediction task.


Models

Two ensemble classification models were evaluated.

1. Random Forest Classifier

Random Forest combines multiple decision trees and aggregates their predictions.

The model was first evaluated using the standard classification threshold of 0.50.

Feature importance was then analyzed to identify which variables contributed most strongly to the model's predictions.

The model identified historical driver performance and starting position as important predictors.


2. Gradient Boosting Classifier

Gradient Boosting was used as a second ensemble approach.

Unlike Random Forest, where trees are built independently, Gradient Boosting builds trees sequentially, with each new tree attempting to correct errors made by previous trees.

Feature importance was also analyzed to investigate which variables were most influential for the model.


Hyperparameter Optimization

Gradient Boosting was optimized using:

GridSearchCV with 5-fold cross-validation

The following hyperparameters were investigated:

  • n_estimators
  • learning_rate
  • max_depth
  • min_samples_split
  • min_samples_leaf

The optimization focused on the F1-score, because correctly identifying the minority class is more important than maximizing accuracy alone.

The best parameter combination found during the search was:

learning_rate = 0.1
max_depth = 2
min_samples_leaf = 2
min_samples_split = 5
n_estimators = 100

The best cross-validation F1-score was approximately:

0.1865

The tuned model was subsequently evaluated on the completely unseen 2020–2024 test period.


Classification Threshold Optimization

By default, binary classification uses a probability threshold of 0.50.

However, with a strongly imbalanced target, this threshold does not necessarily provide the best balance between precision and recall.

Therefore, multiple thresholds between 0.20 and 0.60 were evaluated.

For each threshold, the following metrics were calculated:

  • Precision
  • Recall
  • F1-score

The threshold producing the highest F1-score was selected.

For the Random Forest model, the best threshold in the final evaluation was:

0.37

At this threshold, the Random Forest achieved for the winning class:

  • Precision: 0.46
  • Recall: 0.54
  • F1-score: 0.50

Compared with the standard threshold of 0.50, lowering the threshold increased recall and improved the F1-score.


Model Comparison

The models were evaluated on the unseen 2020–2024 test period.

The standard Random Forest initially achieved stronger results than the standard Gradient Boosting model.

After threshold optimization, the Random Forest improved its ability to identify race winners.

The final comparison focuses primarily on the minority class because identifying actual race winners is the central prediction objective.

Model Threshold Precision Recall F1-score Accuracy
Random Forest 0.37 0.46 0.54 0.50 ~95%
Gradient Boosting To be determined To be determined To be determined To be determined To be determined

The final Gradient Boosting values will be added after the optimized model has been evaluated.


Key Machine Learning Findings

Several observations emerged from the machine learning analysis.

Historical Performance Matters

Features such as:

  • previous win rate
  • previous podium rate
  • previous average position
  • previous constructor performance

were consistently important for the models.

This suggests that historical performance provides useful information when predicting future race outcomes.

Starting Position Matters

grid and pole_position were particularly important features.

Gradient Boosting placed especially strong importance on pole_position, demonstrating the predictive value of qualifying performance.

Accuracy Can Be Misleading

Because only around 4.22% of observations represent race wins, a model can achieve high accuracy while performing poorly at identifying winners.

For this reason, precision, recall and F1-score were prioritized when evaluating the models.

Threshold Selection Matters

Changing the classification threshold significantly changed the balance between precision and recall.

Lower thresholds identified more potential winners but also produced more false positives.

The optimal threshold was therefore selected based on the F1-score rather than accuracy.


Conclusion

This project demonstrates a complete data science and machine learning workflow using historical Formula 1 data.

The project combines multiple datasets to investigate historical race patterns and create meaningful predictive features.

The workflow included:

  1. Importing multiple Formula 1 datasets
  2. Data inspection and preparation
  3. Exploratory Data Analysis
  4. Data merging
  5. Feature Engineering
  6. Leakage-aware historical feature creation
  7. Time-based train/test splitting
  8. Missing-value handling
  9. Baseline classification
  10. Random Forest classification
  11. Gradient Boosting classification
  12. Hyperparameter optimization with GridSearchCV
  13. Classification threshold optimization
  14. Model comparison

The analysis demonstrates that predicting Formula 1 race winners is challenging due to the strong class imbalance and the complexity of race outcomes.

Historical driver and constructor performance, together with starting position, provide useful predictive information. However, no single feature is sufficient to determine a race winner.

The project also demonstrates an important machine learning principle:

A high accuracy does not necessarily mean that a classification model performs well, especially when the target classes are highly imbalanced.

Therefore, model evaluation should consider the specific objective of the prediction task and use appropriate metrics such as precision, recall and F1-score.


Technologies Used

  • Python
  • Pandas
  • NumPy
  • Matplotlib
  • Seaborn
  • Scikit-learn
  • Jupyter Notebook
  • Kaggle

Project Structure

The project is divided into three main stages:

Formula 1 Data
      ↓
Exploratory Data Analysis
      ↓
Feature Engineering
      ↓
Machine Learning
      ↓
Model Evaluation
      ↓
Threshold Optimization
      ↓
Model Comparison

The feature-engineered dataset is exported as:

formula_dataset_ml.csv

This dataset serves as the input for the machine learning workflow.

About

Formula 1 data analysis and race win prediction using EDA, feature engineering, Random Forest, and Gradient Boosting.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages