Exploratory data analysis (EDA)¶
After creating a dataset - we should first check whether it is useful at all and whether it is possible to use it to train a model.
import pandas as pd
import numpy as np
import plotly.express as px
laps = pd.read_parquet('../data/laps_clean.parquet')
dr = pd.read_parquet('../data/driver_race.parquet')
print('laps:', laps.shape, '| driver_race:', dr.shape)
print('sezony:', sorted(laps.Season.unique()), '| torów:', laps.EventName.nunique())
laps: (39647, 18) | driver_race: (918, 12) sezony: [np.int64(2023), np.int64(2024)] | torów: 24
Distribution of lap times¶
fig = px.histogram(laps, x='LapTime_s', nbins=80, color='Season',
title='Rozkład czasów okrążeń (po czyszczeniu)')
fig.update_layout(xaxis_title='czas okrążenia [s]', yaxis_title='liczba okrążeń')
fig.show()
In the chart above, you can see that the distribution ranges from approximately 67 to 116 seconds.
The range is therefore almost 50 seconds. The histogram also shows several peaks. They come from the lap times on individual tracks and depend primarily on the length and type of each track. Each cluster corresponds approximately to one of the tracks. For example, Austria, with the shortest track, forms a peak at ~71s.
This could be a challenge when training a model to predict lap time. Since the result depends largely on which track we're driving on, the model will learn mainly to recognize the track, not what we really want to study (the effect of tires, fuel, or weather).
What proportion of lap-time variability does the track explain?¶
We noted above that the track is the decisive factor for lap time. This still needs to be quantified, though:
var_total = laps['LapTime_s'].var()
resid = laps.groupby('race_id')['LapTime_s'].transform(lambda s: s - s.median())
udzial_toru = 1 - resid.var() / var_total
print(f'Odchylenie std czasu (globalnie): {laps.LapTime_s.std():.2f} s')
print(f'Odchylenie std PO odjęciu mediany toru: {resid.std():.2f} s')
print(f'=> sam tor tłumaczy {udzial_toru:.1%} zmienności czasu okrążenia')
Odchylenie std czasu (globalnie): 10.53 s Odchylenie std PO odjęciu mediany toru: 1.13 s => sam tor tłumaczy 98.8% zmienności czasu okrążenia
med = laps.groupby('EventName')['LapTime_s'].median().sort_values()
fig = px.bar(med, title='Mediana czasu okrążenia per tor - to jest „tło", które dominuje',
labels={'value': 'mediana czasu [s]', 'EventName': 'tor'})
fig.update_layout(showlegend=False, height=520)
fig.show()
The track accounts for 98.8% of the variability.
This means that a model trained on lap time in its current form will learn mainly how long it takes to complete a lap on a given track. But we already know this before the race!
We should therefore think about how to represent this time differently so that the measure is normalized across tracks.
Tire degradation¶
How does lap time change as the tire wears?
Tire wear (TyreLife - the number of laps completed on a given set) is one of the features we take into account. Intuition tells us that the older the tire, the less grip and therefore the slower the lap:
g = laps.groupby(['Compound', 'TyreLife'])['LapTime_s']
deg = g.agg(['median', 'count']).reset_index()
deg = deg[deg['count'] >= 20]
fig = px.line(deg, x='TyreLife', y='median', color='Compound', markers=True,
title='Mediana czasu vs wiek opony per mieszanka',
labels={'median': 'mediana czasu [s]', 'TyreLife': 'wiek opony [okr.]'})
fig.show()
The result runs counter to intuition: the median time decreases with tire age.
Before concluding that tires get faster with age, let's check what this chart ignores: the point in the race. An old set of tires usually means a late phase of the race, and a race evolves over time - the fuel load changes, the track rubbers in. So let's examine how race progress relates to tire age and to lap time:
korelacja = laps[['TyreLife', 'LapNumber']].corr().iloc[0, 1]
print(f'korelacja TyreLife vs LapNumber: {korelacja:.2f}')
nachylenia = []
for race_id, g in laps.groupby('race_id'):
if len(g) < 20:
continue
a, _ = np.polyfit(g['LapNumber'].astype(float), g['LapTime_s'], 1)
nachylenia.append(a)
efekt_paliwa = np.median(nachylenia)
print(f'wypadkowy trend tempa (paliwo + degradacja + ewolucja toru): {efekt_paliwa:.3f} s/okrążenie')
print(f'w typowym wyścigu (~57 okrążeń) auto przyspiesza łącznie o ~{abs(efekt_paliwa) * 57:.1f} s')
print('-> mimo że opony się zużywają (spowalniają), auto przyspiesza:')
print(' efekt paliwa jest silniejszy niż degradacja i przykrywa ją w surowych danych')
korelacja TyreLife vs LapNumber: 0.53 wypadkowy trend tempa (paliwo + degradacja + ewolucja toru): -0.045 s/okrążenie w typowym wyścigu (~57 okrążeń) auto przyspiesza łącznie o ~2.5 s -> mimo że opony się zużywają (spowalniają), auto przyspiesza: efekt paliwa jest silniejszy niż degradacja i przykrywa ją w surowych danych
Two numbers. First, the correlation of 0.53: tire age and lap number really do go hand in hand - the older the set, the further into the race. Second, lap time falls with every successive lap, by 0.045 s on average, i.e. by about 2.5 s over a typical race. The car gets faster during the race despite its tires wearing out.
The simplest explanation is the fuel burning off: a car that burns around 100 kg of fuel over a race is clearly faster towards the end. We have no fuel-load data - lap number serves here as a proxy for fuel load, assuming fuel is consumed roughly evenly. The measured trend in fact blends several factors: not only fuel, but also tire degradation and track evolution. We can't separate them here, but their net result is clear: whatever speeds the car up is stronger than what slows it down, and in the raw data it masks tire degradation.
For this reason the first chart was misleading. To see degradation alone, we must remove this dominant trend. So within each race we fit a linear trend of lap time against lap number and subtract it. What remains is pace_resid - the deviation with the fuel trend and the track's baseline time removed:
def korekta_paliwa(g):
x = g['LapNumber'].values.astype(float)
y = g['LapTime_s'].values
a, b = np.polyfit(x, y, 1)
g['pace_resid'] = y - (a * x + b)
return g
laps_fc = laps.groupby('race_id', group_keys=False).apply(korekta_paliwa)
# tylko opony na sucho (INTERMEDIATE/WET to wyścigi w deszczu - inna historia)
suche = laps_fc[laps_fc['Compound'].isin(['SOFT', 'MEDIUM', 'HARD'])]
deg = (suche.groupby(['Compound', 'TyreLife'])['pace_resid']
.agg(['median', 'count']).reset_index())
deg = deg[deg['count'] >= 20]
fig = px.line(deg, x='TyreLife', y='median', color='Compound', markers=True,
title='Degradacja opony po usunięciu efektu paliwa',
labels={'median': 'odchylenie tempa [s]', 'TyreLife': 'wiek opony [okr.]'})
fig.add_hline(y=0, line_dash='dash')
fig.show()
/tmp/ipykernel_1208528/646244157.py:8: FutureWarning: DataFrameGroupBy.apply operated on the grouping columns. This behavior is deprecated, and in a future version of pandas the grouping columns will be excluded from the operation. Either pass `include_groups=False` to exclude the groupings or explicitly select the grouping columns after groupby to silence this warning.
laps_fc = laps.groupby('race_id', group_keys=False).apply(korekta_paliwa)
After removing the fuel effect, the picture flips and agrees with the initial assumption. Laps become slower with tire age.
The relationship is nonlinear and different for each compound - the soft compound (SOFT) degrades faster than the hard one (HARD). At this stage of the analysis, it's already clear that a single straight line (linear regression) is not ideal for this problem.
Classification: starting position vs podium¶
rate = dr.groupby('GridPosition')['podium'].mean().reset_index()
fig = px.bar(rate, x='GridPosition', y='podium',
title='Szansa na podium wg pozycji startowej',
labels={'podium': 'P(podium)', 'GridPosition': 'pozycja startowa'})
fig.show()
The chart above shows the relationship quite clearly - the further forward the car starts, the higher its chance of a podium.
print(f'odsetek podiów: {dr.podium.mean():.1%} (klasy NIEzbalansowane!)')
fig = px.histogram(dr, x='podium', color='podium',
title=f'Balans klas: podium={dr.podium.mean():.1%} reszty')
fig.show()
odsetek podiów: 15.0% (klasy NIEzbalansowane!)
You can see that podiums are only ~15% of cases. If a model assumed that there is never a podium - it would achieve about 85% accuracy! For this reason, we will need to look at other metrics.