Building the dataset¶
Iteration through all races of the selected seasons. We build two datasets:
laps_clean.parquet- one row = one clean lap (for lap time regression),driver_race.parquet- one row = one driver in one race (for podium classification).
import os, warnings
import fastf1
import numpy as np
import pandas as pd
warnings.simplefilter('ignore')
fastf1.set_log_level('ERROR')
os.makedirs('../cache', exist_ok=True)
os.makedirs('../data', exist_ok=True)
fastf1.Cache.enable_cache('../cache')
SEASONS = [2023, 2024]
SEASONS
[2023, 2024]
Step 1 - qualifying times¶
The first mistake was rooted in a misunderstanding of the dataset.
The starting position and race result are in the race session ('R'). When we also try to fetch the qualifying times,
we run into the following problem:
demo = fastf1.get_session(2024, 1, 'R')
demo.load(telemetry=False, weather=False, messages=False)
demo.results[['Abbreviation', 'Q1', 'Q2', 'Q3']].head()
| Abbreviation | Q1 | Q2 | Q3 | |
|---|---|---|---|---|
| 1 | VER | NaT | NaT | NaT |
| 11 | PER | NaT | NaT | NaT |
| 55 | SAI | NaT | NaT | NaT |
| 16 | LEC | NaT | NaT | NaT |
| 63 | RUS | NaT | NaT | NaT |
In the race session results, the qualifying times Q1/Q2/Q3 are empty. It turns out they live in a separate session 'Q':
demo = fastf1.get_session(2024, 1, 'Q')
demo.load(telemetry=False, weather=False, messages=False)
demo.results[['Abbreviation', 'Q1', 'Q2', 'Q3']].head()
| Abbreviation | Q1 | Q2 | Q3 | |
|---|---|---|---|---|
| 1 | VER | 0 days 00:01:30.031000 | 0 days 00:01:29.374000 | 0 days 00:01:29.179000 |
| 16 | LEC | 0 days 00:01:30.243000 | 0 days 00:01:29.165000 | 0 days 00:01:29.407000 |
| 63 | RUS | 0 days 00:01:30.350000 | 0 days 00:01:29.922000 | 0 days 00:01:29.485000 |
| 55 | SAI | 0 days 00:01:29.909000 | 0 days 00:01:29.573000 | 0 days 00:01:29.507000 |
| 11 | PER | 0 days 00:01:30.221000 | 0 days 00:01:29.932000 | 0 days 00:01:29.537000 |
If we were to build the quali_gap_s feature from the 'R' session, it would be completely empty!
It would then be ignored by the model - not contributing at all to model quality. This was an error we made in the first version of the project.
def quali_gap(season, rnd):
"""
O ile sekund najlepsze okrążenie kierowcy w kwalifikacjach
było wolniejsze od najszybszego okrążenia w całych kwalifikacjach?
"""
q = fastf1.get_session(season, rnd, 'Q')
q.load(telemetry=False, weather=False, messages=False)
r = q.results.copy()
for col in ['Q1', 'Q2', 'Q3']:
r[col + '_s'] = pd.to_timedelta(r[col]).dt.total_seconds()
r['quali_best_s'] = r[['Q1_s', 'Q2_s', 'Q3_s']].min(axis=1)
r['quali_gap_s'] = r['quali_best_s'] - r['quali_best_s'].min()
return r[['Abbreviation', 'quali_gap_s']]
quali_gap(2024, 1).head()
| Abbreviation | quali_gap_s | |
|---|---|---|
| 1 | VER | 0.014 |
| 16 | LEC | 0.000 |
| 63 | RUS | 0.320 |
| 55 | SAI | 0.342 |
| 11 | PER | 0.372 |
Step 2 - iterating through races¶
A single season consists of over 20 races held in different countries.
In this loop, all races (Grands Prix) from the selected seasons are collected, and for each one we gather three types of data into a single set:
- laps - each completed lap along with weather data; this is the basis for predicting lap time
- results - final position, starting position and team of each driver; the basis for podium classification
- qualifying time gap - the output of the method above, as a separate column; by how many seconds the driver was slower than the fastest qualifying lap.
Races are identified by a key composed of the season and the round (season_round).
The goal of this step is to merge the data into specific tables. We can save these tables to parquet files - which will be used by the next notebooks.
all_laps, all_dr = [], []
for season in SEASONS:
sched = fastf1.get_event_schedule(season, include_testing=False)
rounds = sched[sched['RoundNumber'] > 0][['RoundNumber', 'EventName']]
print(f'== sezon {season}: {len(rounds)} rund ==')
for _, row in rounds.iterrows():
rnd, name = int(row['RoundNumber']), row['EventName']
race_id = f'{season}_{rnd:02d}'
try:
s = fastf1.get_session(season, rnd, 'R')
s.load()
except Exception as e:
print(f' [pomijam] {race_id} {name}: {e}')
continue
laps = s.laps
if laps is None or len(laps) == 0:
print(f' [pomijam] {race_id} {name}: brak okrążeń'); continue
w = laps.get_weather_data().reset_index(drop=True)
laps = laps.reset_index(drop=True)
laps = pd.concat([laps, w.drop(columns=['Time'])], axis=1)
laps['Season'], laps['Round'] = season, rnd
laps['EventName'], laps['race_id'] = name, race_id
all_laps.append(laps)
res = s.results.copy()
res['Season'], res['Round'] = season, rnd
res['EventName'], res['race_id'] = name, race_id
try:
res = res.merge(quali_gap(season, rnd), on='Abbreviation', how='left')
except Exception as e:
print(f' [quali skip] {race_id}: {e}'); res['quali_gap_s'] = np.nan
all_dr.append(res)
print(f' [ok] {race_id} {name:<26} okr.={len(laps):>4} quali={res.quali_gap_s.notna().sum():>2}/{len(res)}')
laps_raw = pd.concat(all_laps, ignore_index=True)
res_raw = pd.concat(all_dr, ignore_index=True)
print('\nlaps_raw:', laps_raw.shape, '| results_raw:', res_raw.shape)
== sezon 2023: 22 rund ==
[ok] 2023_01 Bahrain Grand Prix okr.=1056 quali=20/20
[ok] 2023_02 Saudi Arabian Grand Prix okr.= 943 quali=20/20
[ok] 2023_03 Australian Grand Prix okr.=1003 quali=19/20
[ok] 2023_04 Azerbaijan Grand Prix okr.= 962 quali=20/20
[ok] 2023_05 Miami Grand Prix okr.=1138 quali=20/20
[ok] 2023_06 Monaco Grand Prix okr.=1515 quali=20/20
[ok] 2023_07 Spanish Grand Prix okr.=1312 quali=20/20
[ok] 2023_08 Canadian Grand Prix okr.=1317 quali=20/20
[ok] 2023_09 Austrian Grand Prix okr.=1354 quali=20/20
[ok] 2023_10 British Grand Prix okr.= 971 quali=20/20
[ok] 2023_11 Hungarian Grand Prix okr.=1252 quali=20/20
[ok] 2023_12 Belgian Grand Prix okr.= 816 quali=20/20
[ok] 2023_13 Dutch Grand Prix okr.=1343 quali=20/20
[ok] 2023_14 Italian Grand Prix okr.= 957 quali=20/20
[ok] 2023_15 Singapore Grand Prix okr.=1088 quali=20/20
[ok] 2023_16 Japanese Grand Prix okr.= 880 quali=19/20
[ok] 2023_17 Qatar Grand Prix okr.=1006 quali=20/20
[ok] 2023_18 United States Grand Prix okr.=1014 quali=20/20
[ok] 2023_19 Mexico City Grand Prix okr.=1282 quali=19/20
[ok] 2023_20 São Paulo Grand Prix okr.=1108 quali=20/20
[ok] 2023_21 Las Vegas Grand Prix okr.= 946 quali=20/20
[ok] 2023_22 Abu Dhabi Grand Prix okr.=1157 quali=19/20 == sezon 2024: 24 rund ==
[ok] 2024_01 Bahrain Grand Prix okr.=1129 quali=20/20
[ok] 2024_02 Saudi Arabian Grand Prix okr.= 901 quali=19/20
[ok] 2024_03 Australian Grand Prix okr.= 998 quali=19/19
[ok] 2024_04 Japanese Grand Prix okr.= 907 quali=20/20
[ok] 2024_05 Chinese Grand Prix okr.=1032 quali=20/20
[ok] 2024_06 Miami Grand Prix okr.=1111 quali=20/20
[ok] 2024_07 Emilia Romagna Grand Prix okr.=1238 quali=19/20
[ok] 2024_08 Monaco Grand Prix okr.=1237 quali=20/20
[ok] 2024_09 Canadian Grand Prix okr.=1272 quali=20/20
[ok] 2024_10 Spanish Grand Prix okr.=1310 quali=20/20
[ok] 2024_11 Austrian Grand Prix okr.=1405 quali=20/20
[ok] 2024_12 British Grand Prix okr.= 960 quali=20/20
[ok] 2024_13 Hungarian Grand Prix okr.=1355 quali=20/20
[ok] 2024_14 Belgian Grand Prix okr.= 841 quali=20/20
[ok] 2024_15 Dutch Grand Prix okr.=1426 quali=19/20
[ok] 2024_16 Italian Grand Prix okr.=1008 quali=20/20
[ok] 2024_17 Azerbaijan Grand Prix okr.= 973 quali=20/20
[ok] 2024_18 Singapore Grand Prix okr.=1177 quali=20/20
[ok] 2024_19 United States Grand Prix okr.=1059 quali=20/20
[ok] 2024_20 Mexico City Grand Prix okr.=1215 quali=20/20
[ok] 2024_21 São Paulo Grand Prix okr.=1134 quali=20/20
[ok] 2024_22 Las Vegas Grand Prix okr.= 938 quali=20/20
[ok] 2024_23 Qatar Grand Prix okr.= 943 quali=20/20
[ok] 2024_24 Abu Dhabi Grand Prix okr.=1035 quali=20/20 laps_raw: (51024, 42) | results_raw: (919, 27)
Step 3 - cleaning laps for time prediction¶
How long does a lap take depending on tires, fuel, weather, and the car?
The goal of the first dataset is to answer this question. To determine this, we must get rid of laps whose time depends on something other than pure pace.
If our dataset kept laps without a measured time, or laps marked as unreliable (IsAccurate=False),
there would be a risk that the model would learn noise instead of real correlations.
The following filters were applied:
- Laps without a measured time (
LapTime): the model has nothing to learn from. - No pit-in or pit-out laps: the driver slows down to enter the pit lane, or is still getting up to speed after leaving it. That time does not reflect on-track pace.
- Green flag only (
TrackStatus == '1'): we drop laps completed under conditions that artificially inflate the time - for example, a yellow flag indicates danger ahead. Drivers must then slow down and are not allowed to overtake. - Laps marked as accurate (
IsAccurate == True): this is the FastF1 library's flag for measurement quality. It restricts the dataset to reliable measurements. - The 107% rule: within each race we keep only laps no slower than 107% of the fastest lap. This quickly cuts off the remaining strong outliers - for example running in traffic, fuel/tire saving, or minor technical problems.
Additionally, the SectorTime and Speed columns are intentionally not included. Since a lap consists of 3 sectors (sections of the track), the lap time is their sum.
Feeding these times to the model as features means it would learn nothing about pace. It would simply add the numbers and reproduce the exact answer. The model would look flawless.
The above behavior is an example of data leakage. That is - information about the answer, which we cannot actually know before the event, leaks into the input features. After all, we only learn the times of individual sectors after completing them. When making a prediction, we cannot use them.
KEEP = ['Season', 'Round', 'EventName', 'race_id', 'Driver', 'Team',
'LapNumber', 'Stint', 'Compound', 'TyreLife', 'FreshTyre',
'AirTemp', 'TrackTemp', 'Humidity', 'Pressure', 'WindSpeed', 'Rainfall',
'LapTime_s']
cl = laps_raw.copy()
cl = cl[cl['LapTime'].notna()]
cl = cl[cl['PitInTime'].isna() & cl['PitOutTime'].isna()]
cl = cl[cl['TrackStatus'].astype(str) == '1']
cl = cl[cl['IsAccurate'] == True]
cl['LapTime_s'] = cl['LapTime'].dt.total_seconds()
fastest = cl.groupby('race_id')['LapTime_s'].transform('min')
cl = cl[cl['LapTime_s'] <= 1.07 * fastest]
cl['FreshTyre'] = cl['FreshTyre'].astype(bool)
laps_clean = cl[KEEP].copy()
print('laps_clean:', laps_clean.shape)
print('okrążeń na wyścig (mediana):', int(laps_clean.groupby('race_id').size().median()))
laps_clean.head()
laps_clean: (39647, 18) okrążeń na wyścig (mediana): 861
| Season | Round | EventName | race_id | Driver | Team | LapNumber | Stint | Compound | TyreLife | FreshTyre | AirTemp | TrackTemp | Humidity | Pressure | WindSpeed | Rainfall | LapTime_s | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 3.0 | 1.0 | SOFT | 6.0 | False | 27.3 | 31.2 | 22.0 | 1016.7 | 0.6 | False | 98.006 |
| 3 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 4.0 | 1.0 | SOFT | 7.0 | False | 27.3 | 31.2 | 22.0 | 1016.7 | 0.4 | False | 97.976 |
| 4 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 5.0 | 1.0 | SOFT | 8.0 | False | 27.2 | 31.0 | 22.0 | 1016.7 | 1.0 | False | 98.035 |
| 5 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 6.0 | 1.0 | SOFT | 9.0 | False | 27.1 | 31.0 | 22.0 | 1016.9 | 0.6 | False | 97.986 |
| 6 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 7.0 | 1.0 | SOFT | 10.0 | False | 27.1 | 30.9 | 22.0 | 1016.9 | 0.6 | False | 98.021 |
laps_clean.to_parquet('../data/laps_clean.parquet', index=False)
print('zapisano ../data/laps_clean.parquet')
zapisano ../data/laps_clean.parquet
Step 4 - preparing the driver dataset for podium classification¶
Will a given driver in a given Grand Prix finish on the podium?
To perform classification, we added a podium column. It takes the value 1 when the driver finishes in the top three, and 0 otherwise.
As in the previous case, we must make sure to include only features known before the race. Such features include:
GridPosition- starting position. It is determined mainly by qualifying, but also reflects any penalties, so it doesn't always match the qualifying result.Team- teamquali_gap_s- time gap to the fastest qualifying lap - it tells by how much the driver was slower than the front-runners.
It is worth noting that GridPosition = 0 is treated as the last position - a value of 0 means a pit-lane start.
dr = res_raw.rename(columns={'Abbreviation': 'Driver', 'TeamName': 'Team'})
dr['GridPosition'] = dr['GridPosition'].replace(0, np.nan)
dr['GridPosition'] = dr['GridPosition'].fillna(
dr.groupby('race_id')['GridPosition'].transform('max') + 1)
dr = dr[dr['Position'].notna()].copy()
dr['podium'] = (dr['Position'] <= 3).astype(int)
cols = ['Season', 'Round', 'EventName', 'race_id', 'Driver', 'Team',
'GridPosition', 'Position', 'Points', 'Status', 'quali_gap_s', 'podium']
dr = dr[cols]
print('driver_race:', dr.shape)
print('quali_gap_s wypełnione:', dr['quali_gap_s'].notna().sum(), '/', len(dr))
print('odsetek podiów:', round(dr['podium'].mean(), 3))
dr.head()
driver_race: (918, 12) quali_gap_s wypełnione: 911 / 918 odsetek podiów: 0.15
| Season | Round | EventName | race_id | Driver | Team | GridPosition | Position | Points | Status | quali_gap_s | podium | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | VER | Red Bull Racing | 1.0 | 1.0 | 25.0 | Finished | 0.000 | 1 |
| 1 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | PER | Red Bull Racing | 2.0 | 2.0 | 18.0 | Finished | 0.138 | 1 |
| 2 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | ALO | Aston Martin | 5.0 | 3.0 | 15.0 | Finished | 0.628 | 1 |
| 3 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | SAI | Ferrari | 4.0 | 4.0 | 12.0 | Finished | 0.446 | 0 |
| 4 | 2023 | 1 | Bahrain Grand Prix | 2023_01 | HAM | Mercedes | 7.0 | 5.0 | 10.0 | Finished | 0.676 | 0 |
dr.to_parquet('../data/driver_race.parquet', index=False)
print('zapisano ../data/driver_race.parquet')
zapisano ../data/driver_race.parquet