This project analyzes historical Formula 1 race data from 1950 to 2024 and explores whether historical driver and constructor performance can be used to predict race wins.
The project follows a complete data science workflow:
- Import and inspect multiple CSV datasets
- Perform Exploratory Data Analysis (EDA)
- Merge and combine data from different sources
- Create historical performance features through Feature Engineering
- Prepare the data for machine learning
- Train classification models
- Optimize model hyperparameters using GridSearchCV
- Optimize classification thresholds
- Compare model performance using precision, recall, F1-score and accuracy
The main objective is not only to build a predictive model, but also to understand which factors are most strongly associated with winning a Formula 1 race.
The project uses the Formula 1 World Championship 1950–2020 dataset by Rohan Rao, extended with the available race data through 2024.
The dataset consists of multiple CSV files covering different aspects of Formula 1 history.
Main datasets used include:
- Races
- Drivers
- Constructors
- Results
- Qualifying
- Circuits
- Driver standings
- Constructor standings
- Lap times
- Pit stops
- Sprint results
- Status
The use of multiple related datasets makes it possible to combine race, driver, constructor and circuit information.
The exploratory analysis investigates historical patterns within Formula 1.
Key areas analyzed include:
- Number of races over time
- Number of participating drivers
- Race distribution by country
- Most frequently used circuits
- Grand Prix frequency
- Driver victories
- Constructor victories
- Nationality of race winners
- Starting grid positions
- Position changes during races
- Relationship between pole position and race wins
Starting from pole position provides a significant advantage.
However, pole position does not guarantee victory.
Approximately 42.34% of pole positions resulted in a race win in the analyzed dataset.
This demonstrates that qualifying performance is an important factor, but race results are influenced by additional factors.
Drivers frequently gained or lost positions between their starting grid position and their final race position.
This highlights that starting position and final race result are related, but not identical.
Historically successful constructors such as Ferrari, McLaren, Mercedes and Red Bull account for a large share of Formula 1 race victories and championship points.
A relatively small number of circuits and countries have hosted a large proportion of Formula 1 races.
Monza, Monaco and Silverstone are among the circuits with the highest number of hosted races.
The Feature Engineering stage transforms historical race information into variables that can be used by machine learning models.
A central principle of the feature engineering process is avoiding data leakage.
Historical performance features are calculated using information available before the current race.
This allows the machine learning model to simulate a realistic prediction scenario.
The following historical driver features were created:
previous_racesprevious_winsprevious_win_rateprevious_pointsprevious_avg_positionprevious_podiumsprevious_podium_rate
These features describe the driver's historical performance before the current race.
Historical constructor performance was also incorporated:
previous_constructor_winsprevious_constructor_pointsprevious_constructor_racesprevious_constructor_win_rate
Current race-related features include:
gridpole_positionyear
The target variable is:
win
where:
0= driver did not win1= driver won the race
The machine learning task is a binary classification problem.
The objective is:
Predict whether a Formula 1 driver will win a given race.
The target variable is highly imbalanced because only a small proportion of all race results represent victories.
Approximately:
- 95.78% → no win
- 4.22% → win
Because of this imbalance, accuracy alone is not sufficient for evaluating model performance.
Particular attention is therefore given to:
- Precision
- Recall
- F1-score
for the winning class.
Because Formula 1 is a chronological dataset, a random train/test split was avoided.
Instead, a time-based split was used:
| Dataset | Years |
|---|---|
| Training | 1950–2019 |
| Testing | 2020–2024 |
This approach simulates a realistic prediction scenario where historical data is used to predict future races.
The final datasets contain:
- 26,759 observations
- 14 machine learning features
The feature previous_avg_position contains missing values for drivers who had not participated in a previous race.
These missing values were handled using:
SimpleImputer with median strategy
The imputer was fitted only on the training data and then applied to the test data.
This prevents information from the test period from influencing the preprocessing process.
A DummyClassifier was used as a baseline model.
The baseline always predicts the most frequent class.
Although this produces a relatively high accuracy due to the class imbalance, it fails to identify race winners.
This demonstrates why accuracy alone is not an appropriate metric for this prediction task.
Two ensemble classification models were evaluated.
Random Forest combines multiple decision trees and aggregates their predictions.
The model was first evaluated using the standard classification threshold of 0.50.
Feature importance was then analyzed to identify which variables contributed most strongly to the model's predictions.
The model identified historical driver performance and starting position as important predictors.
Gradient Boosting was used as a second ensemble approach.
Unlike Random Forest, where trees are built independently, Gradient Boosting builds trees sequentially, with each new tree attempting to correct errors made by previous trees.
Feature importance was also analyzed to investigate which variables were most influential for the model.
Gradient Boosting was optimized using:
GridSearchCV with 5-fold cross-validation
The following hyperparameters were investigated:
n_estimatorslearning_ratemax_depthmin_samples_splitmin_samples_leaf
The optimization focused on the F1-score, because correctly identifying the minority class is more important than maximizing accuracy alone.
The best parameter combination found during the search was:
learning_rate = 0.1
max_depth = 2
min_samples_leaf = 2
min_samples_split = 5
n_estimators = 100
The best cross-validation F1-score was approximately:
0.1865
The tuned model was subsequently evaluated on the completely unseen 2020–2024 test period.
By default, binary classification uses a probability threshold of 0.50.
However, with a strongly imbalanced target, this threshold does not necessarily provide the best balance between precision and recall.
Therefore, multiple thresholds between 0.20 and 0.60 were evaluated.
For each threshold, the following metrics were calculated:
- Precision
- Recall
- F1-score
The threshold producing the highest F1-score was selected.
For the Random Forest model, the best threshold in the final evaluation was:
0.37
At this threshold, the Random Forest achieved for the winning class:
- Precision: 0.46
- Recall: 0.54
- F1-score: 0.50
Compared with the standard threshold of 0.50, lowering the threshold increased recall and improved the F1-score.
The models were evaluated on the unseen 2020–2024 test period.
The standard Random Forest initially achieved stronger results than the standard Gradient Boosting model.
After threshold optimization, the Random Forest improved its ability to identify race winners.
The final comparison focuses primarily on the minority class because identifying actual race winners is the central prediction objective.
| Model | Threshold | Precision | Recall | F1-score | Accuracy |
|---|---|---|---|---|---|
| Random Forest | 0.37 | 0.46 | 0.54 | 0.50 | ~95% |
| Gradient Boosting | To be determined | To be determined | To be determined | To be determined | To be determined |
The final Gradient Boosting values will be added after the optimized model has been evaluated.
Several observations emerged from the machine learning analysis.
Features such as:
- previous win rate
- previous podium rate
- previous average position
- previous constructor performance
were consistently important for the models.
This suggests that historical performance provides useful information when predicting future race outcomes.
grid and pole_position were particularly important features.
Gradient Boosting placed especially strong importance on pole_position, demonstrating the predictive value of qualifying performance.
Because only around 4.22% of observations represent race wins, a model can achieve high accuracy while performing poorly at identifying winners.
For this reason, precision, recall and F1-score were prioritized when evaluating the models.
Changing the classification threshold significantly changed the balance between precision and recall.
Lower thresholds identified more potential winners but also produced more false positives.
The optimal threshold was therefore selected based on the F1-score rather than accuracy.
This project demonstrates a complete data science and machine learning workflow using historical Formula 1 data.
The project combines multiple datasets to investigate historical race patterns and create meaningful predictive features.
The workflow included:
- Importing multiple Formula 1 datasets
- Data inspection and preparation
- Exploratory Data Analysis
- Data merging
- Feature Engineering
- Leakage-aware historical feature creation
- Time-based train/test splitting
- Missing-value handling
- Baseline classification
- Random Forest classification
- Gradient Boosting classification
- Hyperparameter optimization with GridSearchCV
- Classification threshold optimization
- Model comparison
The analysis demonstrates that predicting Formula 1 race winners is challenging due to the strong class imbalance and the complexity of race outcomes.
Historical driver and constructor performance, together with starting position, provide useful predictive information. However, no single feature is sufficient to determine a race winner.
The project also demonstrates an important machine learning principle:
A high accuracy does not necessarily mean that a classification model performs well, especially when the target classes are highly imbalanced.
Therefore, model evaluation should consider the specific objective of the prediction task and use appropriate metrics such as precision, recall and F1-score.
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- Scikit-learn
- Jupyter Notebook
- Kaggle
The project is divided into three main stages:
Formula 1 Data
↓
Exploratory Data Analysis
↓
Feature Engineering
↓
Machine Learning
↓
Model Evaluation
↓
Threshold Optimization
↓
Model Comparison
The feature-engineered dataset is exported as:
formula_dataset_ml.csv
This dataset serves as the input for the machine learning workflow.