Regression: lap time¶
How long does a lap take depending on tires, fuel, and weather?
Below we attempt to create a model that will answer this question.
import time
import numpy as np
import pandas as pd
import plotly.express as px
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.dummy import DummyRegressor
from sklearn.linear_model import LinearRegression, Ridge
from sklearn.neighbors import KNeighborsRegressor
from sklearn.ensemble import RandomForestRegressor, HistGradientBoostingRegressor
from sklearn.model_selection import GroupShuffleSplit
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
from sklearn.inspection import permutation_importance
laps = pd.read_parquet('../data/laps_clean.parquet')
laps['FreshTyre'] = laps['FreshTyre'].astype(int)
laps['circuit'] = laps['EventName']
print('laps:', laps.shape, '| sezony:', sorted(laps.Season.unique()),
'| torów:', laps.EventName.nunique(), '| wyścigów (race_id):', laps.race_id.nunique())
laps: (39647, 19) | sezony: [np.int64(2023), np.int64(2024)] | torów: 24 | wyścigów (race_id): 46
cat_cols = ['Compound', 'Team']
num_cols = ['TyreLife', 'FreshTyre', 'Stint', 'LapNumber',
'AirTemp', 'TrackTemp', 'Humidity', 'Pressure', 'WindSpeed', 'Rainfall']
def report(name, y_true, y_pred):
mae = mean_absolute_error(y_true, y_pred)
rmse = mean_squared_error(y_true, y_pred) ** 0.5
r2 = r2_score(y_true, y_pred)
print(f'{name:<28} MAE={mae:6.3f}s RMSE={rmse:6.3f}s R2={r2:7.3f}')
return dict(model=name, MAE=mae, RMSE=rmse, R2=r2)
def split_per_race(df):
gss = GroupShuffleSplit(n_splits=1, test_size=0.25, random_state=42)
return next(gss.split(df, groups=df['race_id'].values))
Each model will be evaluated using three metrics - the report function calculates all of them at once. It's worth getting to know them right away, as they will appear in our first attempt:
- MAE (Mean Absolute Error) - by how many seconds the model is off on average. The simplest to interpret:
MAE = 0.5 smeans "an error of around half a second on average". - RMSE (root mean square error) - similar to MAE, but penalizes large errors more heavily (squares the errors). When RMSE is noticeably larger than MAE, it's a sign that there are occasional large mistakes.
- $R^2$ (coefficient of determination) - what fraction of the target's variability the model explains, on a scale from
0to1. The key here is interpreting the reference points: $R^2 = 1$ is a perfect model, $R^2 = 0$ means you're doing no better than guessing the mean (the model adds nothing), and negative $R^2$ means the model is worse than guessing the mean. We'll see this last part firsthand shortly.
First attempt - absolute time on a single season¶
We start as simply as possible. We consider only the 2024 season, predict absolute LapTime_s, and add the track (circuit) as another categorical feature.
We split the data by race (GroupShuffleSplit by race_id) - so that laps from a single Grand Prix don't end up in both the training and the test set. Otherwise, the model would pick up on conditions specific to that race, and the evaluation would be inflated.
one = laps[laps['Season'] == 2024].reset_index(drop=True)
tr, te = split_per_race(one)
feat = ['circuit'] + cat_cols + num_cols
Xtr, Xte = one[feat].iloc[tr], one[feat].iloc[te]
ytr, yte = one['LapTime_s'].values[tr], one['LapTime_s'].values[te]
pre = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), ['circuit'] + cat_cols),
('num', SimpleImputer(strategy='median'), num_cols)])
lin = Pipeline([('pre', pre), ('reg', LinearRegression())]).fit(Xtr, ytr)
report('Akt 1: Linear (1 sezon)', yte, lin.predict(Xte))
Xtr_h, Xte_h = Xtr.copy(), Xte.copy()
for col in ['circuit'] + cat_cols:
Xtr_h[col] = Xtr_h[col].astype('category')
Xte_h[col] = Xte_h[col].astype('category')
mask = [col in (['circuit'] + cat_cols) for col in feat]
hgb = HistGradientBoostingRegressor(categorical_features=mask, random_state=42).fit(Xtr_h, ytr)
report('Akt 1: HistGB (1 sezon)', yte, hgb.predict(Xte_h))
Akt 1: Linear (1 sezon) MAE=12.249s RMSE=13.776s R2= -0.762
Akt 1: HistGB (1 sezon) MAE=12.016s RMSE=13.861s R2= -0.784
{'model': 'Akt 1: HistGB (1 sezon)',
'MAE': 12.01614779429148,
'RMSE': 13.861168193665701,
'R2': -0.7841030157067987}
The result is disappointing - $R^2$ is negative. This means the model is worse than predicting a constant mean.
As it turns out - both the linear model and gradient boosting return identical results, despite a significant difference in model capacity. If model capacity changes nothing - the problem must lie in how the task is framed.
Why the model failed¶
From notebook 02, we know that the track alone accounts for ~99% of the variability in lap time. Additionally, in a single season each track appears only once!
Since we split the data per race, tracks in the test set don't appear in training at all!
One-hot encoding then converts an unknown track into all zeros. The model has no idea where it's going. We should verify this directly, though:
One-hot encoding. Models work with numbers, not text, and circuit or Team are categories - names.
One-hot encoding converts each category into a separate column taking the value 0 or 1: for a given row, 1 is assigned only to the column for its category, while all others are zero.
This way Monza and Monaco become two independent columns that the model can weight separately.
Encoding such features as 0, 1, 2, ... would impose artificial ordering and distances. The model might think, for example, that track 3 is larger than track 2, or closer to it than to track 1.
tory_tr = set(one['circuit'].iloc[tr].unique())
tory_te = set(one['circuit'].iloc[te].unique())
print('1 SEZON:')
print(' torów w treningu:', len(tory_tr), '| w teście:', len(tory_te))
print(' tory testowe obecne też w treningu:', len(tory_te & tory_tr), ' <-- ZERO')
print(' rozrzut median czasu między torami:',
round(one.groupby('circuit')['LapTime_s'].median().std(), 1), 's')
print('\nModel dostaje tor, którego nie widział, i strzela średnią z innych torów -')
print('dlatego myli się o kilkanaście sekund (tyle wynosi rozrzut między torami).')
1 SEZON: torów w treningu: 18 | w teście: 6 tory testowe obecne też w treningu: 0 <-- ZERO rozrzut median czasu między torami: 10.3 s Model dostaje tor, którego nie widział, i strzela średnią z innych torów - dlatego myli się o kilkanaście sekund (tyle wynosi rozrzut między torami).
First fix - more data¶
The cheapest fix doesn't touch the model. We add a second season (2023). Now each track typically appears twice (once per season).
We use exactly the same model and the same validation. Only the amount of data changes.
tr, te = split_per_race(laps)
feat = ['circuit'] + cat_cols + num_cols
Xtr, Xte = laps[feat].iloc[tr], laps[feat].iloc[te]
ytr, yte = laps['LapTime_s'].values[tr], laps['LapTime_s'].values[te]
tory_tr = set(laps['circuit'].iloc[tr].unique())
tory_te = set(laps['circuit'].iloc[te].unique())
print('2 SEZONY: tory testowe obecne też w treningu:',
len(tory_te & tory_tr), 'z', len(tory_te), ' <-- teraz niezerowe\n')
pre = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), ['circuit'] + cat_cols),
('num', SimpleImputer(strategy='median'), num_cols)])
lin = Pipeline([('pre', pre), ('reg', LinearRegression())]).fit(Xtr, ytr)
report('Akt 3: Linear (2 sezony)', yte, lin.predict(Xte))
2 SEZONY: tory testowe obecne też w treningu: 11 z 12 <-- teraz niezerowe Akt 3: Linear (2 sezony) MAE= 2.017s RMSE= 3.713s R2= 0.904
{'model': 'Akt 3: Linear (2 sezony)',
'MAE': 2.0169686276421195,
'RMSE': 3.7134827563886827,
'R2': 0.9039316983258943}
$R^2$ improved from -0.76 to ~0.90.
However, such a high $R^2$ is somewhat unsettling - given the impact the track alone has on the data.
Second fix - changing the target¶
We already know that the track accounts for ~99% of the variability. If that's the case, then a model achieving $R^2=0.90$ learned mainly which track we're on.
In that case, we should somehow limit the track's influence. What can transfer between tracks is the pace within a race - that is, by how much a given lap is faster or slower than that race's median, depending on tires, fuel, and weather.
New target: LapTime_rel = LapTime_s - median_race_time. The track as a feature disappears - it's already built into the subtracted baseline:
laps['base'] = laps.groupby('race_id')['LapTime_s'].transform('median')
laps['LapTime_rel'] = laps['LapTime_s'] - laps['base']
print('cel względny - odchylenie std:', round(laps['LapTime_rel'].std(), 2),
's (typowe wahanie tempa w obrębie wyścigu)')
tr, te = split_per_race(laps)
Xtr, Xte = laps[cat_cols + num_cols].iloc[tr], laps[cat_cols + num_cols].iloc[te]
ytr, yte = laps['LapTime_rel'].values[tr], laps['LapTime_rel'].values[te]
Xtr_h, Xte_h = Xtr.copy(), Xte.copy()
for col in cat_cols:
Xtr_h[col] = Xtr_h[col].astype('category')
Xte_h[col] = Xte_h[col].astype('category')
mask = [col in cat_cols for col in cat_cols + num_cols]
hgb = HistGradientBoostingRegressor(categorical_features=mask, learning_rate=0.08,
max_iter=400, l2_regularization=1.0,
random_state=42).fit(Xtr_h, ytr)
report('Akt 4: HistGB (tempo względne)', yte, hgb.predict(Xte_h))
cel względny - odchylenie std: 1.13 s (typowe wahanie tempa w obrębie wyścigu)
Akt 4: HistGB (tempo względne) MAE= 0.627s RMSE= 0.808s R2= 0.486
{'model': 'Akt 4: HistGB (tempo względne)',
'MAE': 0.6272653916602322,
'RMSE': 0.8081744244273222,
'R2': 0.48555492588058535}
$R^2$ dropped from ~0.90 to ~0.5. The task is now more honest, though - we've removed the free signal from the track.
We consider this target correct for now and will use it for our algorithm benchmark.
Benchmark - algorithm comparison¶
With a correctly framed task (relative pace), we compare several algorithms on the same data and the same split. We test six approaches:
DummyRegressor- learns nothing, always predicts the mean. This is our baseline: any sensible model must beat it.LinearRegressionandRidge- linear models (Ridge is a variant with regularization, limiting coefficient size). They receive the data one-hot encoded, scaled, and with missing values filled in.kNN(k-nearest neighbors) - predicts based on similar laps from the training set.RandomForestandHistGradientBoosting- tree-based models, able to capture nonlinearities and interactions.HistGradientBoostingaccepts categories natively,RandomForestafter one-hot encoding.
We evaluate them with the quality metrics we've already learned (MAE, RMSE, $R^2$), and additionally record training time and prediction time. The latter matters because "how well the model performs" is not just accuracy, but also computational cost.
pre_lin = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', make_pipeline(SimpleImputer(strategy='median'), StandardScaler()), num_cols)])
rf_pre = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', SimpleImputer(strategy='median'), num_cols)])
def timed(model, Xtr_, ytr_, Xte_, yte_, name):
t = time.perf_counter(); model.fit(Xtr_, ytr_); fit_t = time.perf_counter() - t
t = time.perf_counter(); pred = model.predict(Xte_); pred_t = time.perf_counter() - t
r = report(name, yte_, pred); r['fit_s'] = fit_t; r['pred_s'] = pred_t
return r
rows = []
rows.append(timed(Pipeline([('pre', pre_lin), ('reg', DummyRegressor())]), Xtr, ytr, Xte, yte, 'Dummy (średnia)'))
rows.append(timed(Pipeline([('pre', pre_lin), ('reg', LinearRegression())]), Xtr, ytr, Xte, yte, 'Linear Regression'))
rows.append(timed(Pipeline([('pre', pre_lin), ('reg', Ridge(alpha=1.0))]), Xtr, ytr, Xte, yte, 'Ridge'))
rows.append(timed(Pipeline([('pre', pre_lin), ('reg', KNeighborsRegressor(n_neighbors=15))]), Xtr, ytr, Xte, yte, 'kNN (k=15)'))
rows.append(timed(Pipeline([('pre', rf_pre), ('reg', RandomForestRegressor(n_estimators=200, n_jobs=-1, random_state=42))]), Xtr, ytr, Xte, yte, 'Random Forest'))
rows.append(timed(HistGradientBoostingRegressor(categorical_features=mask, learning_rate=0.08, max_iter=400, l2_regularization=1.0, random_state=42), Xtr_h, ytr, Xte_h, yte, 'HistGradientBoosting'))
bench = pd.DataFrame(rows).set_index('model').round(3)
bench
Dummy (średnia) MAE= 0.912s RMSE= 1.127s R2= -0.000 Linear Regression MAE= 0.585s RMSE= 0.758s R2= 0.548 Ridge MAE= 0.585s RMSE= 0.758s R2= 0.548
kNN (k=15) MAE= 0.736s RMSE= 0.950s R2= 0.289
Random Forest MAE= 0.622s RMSE= 0.799s R2= 0.497
HistGradientBoosting MAE= 0.627s RMSE= 0.808s R2= 0.486
| MAE | RMSE | R2 | fit_s | pred_s | |
|---|---|---|---|---|---|
| model | |||||
| Dummy (średnia) | 0.912 | 1.127 | -0.000 | 0.033 | 0.006 |
| Linear Regression | 0.585 | 0.758 | 0.548 | 0.038 | 0.008 |
| Ridge | 0.585 | 0.758 | 0.548 | 0.036 | 0.009 |
| kNN (k=15) | 0.736 | 0.950 | 0.289 | 0.037 | 0.174 |
| Random Forest | 0.622 | 0.799 | 0.497 | 1.842 | 0.039 |
| HistGradientBoosting | 0.627 | 0.808 | 0.486 | 0.564 | 0.022 |
The table compares all models. The first observation concerns the baseline: Dummy achieves MAE = 0.912 s and an $R^2$ of zero, so any model that falls below this error has actually learned something.
The surprise is that the linear models turn out to be the best - LinearRegression and Ridge give an identical result (MAE = 0.585 s, $R^2 = 0.548$), ahead of both tree-based models (RandomForest 0.497, HistGradientBoosting 0.486). The same task that previously required trees to expose tire degradation becomes, once reframed as relative pace, "linear" enough that plain regression suffices. It is dominated by the fuel effect (LapNumber), which is linear, while tire degradation - though nonlinear - is a much weaker effect. With the limited feature set, the trees have nothing to add here and overfit slightly.
Ridge and LinearRegression give the same result because with this many observations, regularization has nothing left to constrain - the coefficients are stable even without it.
b = bench.reset_index()
fig = px.scatter(b, x='fit_s', y='MAE', text='model', color='model',
title='Dokładność (MAE, niżej=lepiej) vs czas treningu',
labels={'fit_s': 'czas treningu [s]', 'MAE': 'MAE [s]'})
fig.update_traces(textposition='top center', marker_size=12)
fig.update_layout(showlegend=False)
fig.show()
fig = px.bar(b.sort_values('MAE'), x='MAE', y='model', orientation='h',
title='Ranking modeli wg MAE (tempo względne)',
labels={'MAE': 'średni błąd bezwzględny [s]', 'model': ''})
fig.update_layout(height=380)
fig.show()
Both charts show the same thing from two angles. The ranking chart orders models by accuracy (MAE), while the scatter plot compares accuracy with training time - showing how much a given level of quality costs.
A clear lesson emerges: here the simplest model is both the best and the cheapest. LinearRegression has the lowest error and trains in hundredths of a second. By contrast:
RandomForestachieves a worse result but takes roughly forty times longer to train (~1.8 sversus~0.04 s),kNNperforms worst among the models that actually learn ($R^2 = 0.289$) and is also the slowest at prediction, since each query requires searching the entire training set for neighbors.
So - a more complex model is not always better. When a task is properly framed, a simpler solution is often simultaneously more accurate, faster, and easier to understand. Complexity should be introduced only when it genuinely improves results.
It's worth explaining what $R^2 = 0.55$ for the best model actually means. This number says the model explains 55% of the variability in relative pace - the remaining 45% comes from factors absent from our features (traffic on track, engine mode, a driver's form on the day) plus plain noise.
Let's translate this into seconds via MAE: the model is off by ~0.585 s on average, while simply guessing the mean (Dummy) is off by ~0.91 s - a roughly one-third reduction in error. That suffices to understand what drives pace (fuel, car performance, tires), but 0.585 s is still a lot for F1, where positions are fought over hundredths of a second. So the model is an analytical tool, not an oracle for individual laps - it is not fit for precise strategy work.
Model diagnostics¶
Since the models are of similar quality, for diagnostics we take HistGradientBoosting as a representative of the nonlinear approach - we want to see where the model hits and where it misses, and which features matter most to it. We start by comparing predicted values to actual ones.
best_name = bench['MAE'].idxmin()
print('najlepszy wg MAE:', best_name)
p_best = hgb.predict(Xte_h) # HGB jako reprezentatywny model nieliniowy
res_df = pd.DataFrame({'rzeczywiste': yte, 'przewidziane': p_best})
fig = px.scatter(res_df, x='rzeczywiste', y='przewidziane', opacity=0.25,
title='HistGB: przewidziane vs rzeczywiste tempo względne [s]')
lo, hi = res_df.rzeczywiste.min(), res_df.rzeczywiste.max()
fig.add_shape(type='line', x0=lo, y0=lo, x1=hi, y1=hi, line=dict(dash='dash', color='red'))
fig.show()
najlepszy wg MAE: Linear Regression
The chart compares predicted pace (vertical axis) with actual pace (horizontal); each point is one lap from the test set. The red dashed line marks perfect prediction - if the model made no errors, all points would lie exactly on it.
The cloud lies along this line (correlation ~0.70), confirming that the model has captured the main signal. However, there's a clear, systematic pattern: points are flattened around zero relative to the diagonal. For really slow laps (actual pace above +2 s), the model predicts too low, and for very fast ones - too high. The predicted range (-3 to 2.6 s) is narrower than the actual one (-4.3 to 3.3 s).
This effect is regression to the mean: the model, not knowing the cause of an extreme result, cautiously predicts values closer to the average. Extreme laps in F1 usually come from factors we don't have in the data - momentary traffic, a driver error, the aftermath of a safety car period. The model can't predict these, so it plays it safe and stays in the middle. This is the same limit we discussed for $R^2$: part of pace variability is simply absent from our features.
perm = permutation_importance(hgb, Xte_h, yte, n_repeats=8, random_state=42)
imp = pd.DataFrame({'cecha': cat_cols + num_cols,
'ważność': perm.importances_mean}).sort_values('ważność')
fig = px.bar(imp, x='ważność', y='cecha', orientation='h',
title='Ważność cech (permutation importance) - HistGB, tempo względne')
fig.update_layout(height=480)
fig.show()
print('Uwaga: LapNumber (proxy paliwa) jest skorelowany z TyreLife - ważności czytaj łącznie.')
Uwaga: LapNumber (proxy paliwa) jest skorelowany z TyreLife - ważności czytaj łącznie.
The importance ranking tells us a lot about the problem's nature. The strongest feature is LapNumber, a proxy for fuel - confirming the EDA observation that the falling fuel load is the largest time effect in a race. What may be surprising, though, is Team in second place, ahead of TyreLife.
Team approximates how fast the car itself is, and the design differences between teams are enormous. In 2023-24, median relative pace ranged from -0.77 s (Red Bull) to +0.72 s (Kick Sauber) - a spread of 1.5 s per lap, larger than the total tire degradation effect. So just knowing "which car it is" carries more information than tire age.
TyreLife is weakened further still: as we showed in EDA, it correlates with LapNumber. The lap number takes over part of its signal, so its standalone importance drops. Team has no such substitute - its information is unique.
Striking, on the other hand, is how little the weather features matter. AirTemp, TrackTemp, Humidity and WindSpeed have zero or even negative importance (shuffling them does not hurt the predictions), Rainfall is barely above zero, and Pressure (~0.02) is more of a track identifier than a weather variable - pressure hardly changes within a race, but it distinguishes circuits at different altitudes. This doesn't mean weather has no effect on lap time - most likely it means weather mainly differentiates races from one another, and that is something we don't measure here: we removed that component of variability together with the track when we switched to relative pace. Within a single race, temperature, humidity and pressure change little, and what does change (a cooling track, for instance) goes hand in hand with the lap number - once the race median is subtracted, weather has nothing left to explain. Rain and wind vary more often, but wet laps make up under 2% of the cleaned set, and wind - as we can see - contributes nothing. On top of that, with a per-race split the model sees temperature combinations in the test set that it never saw in training - the weather features then act more like noise than signal.
Returning to Team: remember that its strength depends on data homogeneity. It's a good indicator only within a single technical era - "Mercedes 2019" and "Mercedes 2024" are entirely different in pace despite the same name. We'll return to this shortly.
Will more seasons improve the model?¶
Since it was data that saved us in the first fix (adding a second season), the question arises: will more seasons raise the current ~0.55? FastF1 provides data from 2018 onwards, so we can check. We load a wider set (laps_all - seasons 2018-2024) and compute $R^2$ for a growing number of seasons, added from the newest backwards.
We built laps_all.parquet with the same cleaning code as laps_clean in notebook 01 (identical filters), run for seasons 2018-2024 and saved under a different name. Downloading seven seasons takes hours, so the ready-made parquet lives in data/ and the notebook only reads it.
Let's compute it:
laps_all = pd.read_parquet('../data/laps_all.parquet')
laps_all['FreshTyre'] = laps_all['FreshTyre'].astype(int)
seasons_all = sorted(laps_all['Season'].unique())
print('dostępne sezony:', seasons_all, '| okrążeń:', len(laps_all))
def ocena_podzbioru(df):
df = df.copy()
df['base'] = df.groupby('race_id')['LapTime_s'].transform('median')
df['rel'] = df['LapTime_s'] - df['base']
g = df['race_id'].values
tr_, te_ = next(GroupShuffleSplit(1, test_size=0.25, random_state=42).split(df, groups=g))
pre = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', make_pipeline(SimpleImputer(strategy='median'), StandardScaler()), num_cols)])
m = Pipeline([('pre', pre), ('reg', LinearRegression())]).fit(
df[cat_cols + num_cols].iloc[tr_], df['rel'].values[tr_])
p = m.predict(df[cat_cols + num_cols].iloc[te_])
return r2_score(df['rel'].values[te_], p), mean_absolute_error(df['rel'].values[te_], p)
rows = []
for k in range(1, len(seasons_all) + 1):
chosen = seasons_all[-k:]
r2, mae = ocena_podzbioru(laps_all[laps_all['Season'].isin(chosen)])
rows.append({'liczba_sezonów': k, 'zakres': f'{chosen[0]}-{chosen[-1]}', 'R2': r2, 'MAE': mae})
wynik_sez = pd.DataFrame(rows)
wynik_sez.round(3)
dostępne sezony: [np.int64(2018), np.int64(2019), np.int64(2020), np.int64(2021), np.int64(2022), np.int64(2023), np.int64(2024)] | okrążeń: 113622
| liczba_sezonów | zakres | R2 | MAE | |
|---|---|---|---|---|
| 0 | 1 | 2024-2024 | 0.573 | 0.615 |
| 1 | 2 | 2023-2024 | 0.548 | 0.585 |
| 2 | 3 | 2022-2024 | 0.509 | 0.616 |
| 3 | 4 | 2021-2024 | 0.540 | 0.568 |
| 4 | 5 | 2020-2024 | 0.465 | 0.637 |
| 5 | 6 | 2019-2024 | 0.449 | 0.642 |
| 6 | 7 | 2018-2024 | 0.500 | 0.646 |
fig = px.line(wynik_sez, x='liczba_sezonów', y='R2', markers=True,
title='R² regresji a liczba sezonów treningowych',
labels={'liczba_sezonów': 'liczba sezonów', 'R2': 'R² (tempo względne)'})
fig.update_yaxes(range=[0, 0.7])
fig.show()
$R^2$ does not grow with more seasons - it fluctuates and even declines slightly. The best result comes from a single season, the newest one. Why doesn't adding data help this time, when it helped earlier?
The key seems to be data homogeneity. In 2022, F1 underwent a major regulation change (new aerodynamics, larger tires), so cars from 2018-19 and 2023-24 behave entirely differently. Let's check whether mixing these two eras hurts - by comparing a model trained on each era separately with one trained on everything together:
era_nowa = laps_all[laps_all['Season'] >= 2022]
era_stara = laps_all[laps_all['Season'] < 2022]
for nazwa, df in [('era >=2022 (osobno)', era_nowa),
('era <2022 (osobno)', era_stara),
('wszystko zmieszane', laps_all)]:
r2, mae = ocena_podzbioru(df)
print(f'{nazwa:<22} R2={r2:.3f} MAE={mae:.3f}s (okrążeń: {len(df)})')
era >=2022 (osobno) R2=0.509 MAE=0.616s (okrążeń: 56217) era <2022 (osobno) R2=0.547 MAE=0.637s (okrążeń: 57405)
wszystko zmieszane R2=0.500 MAE=0.646s (okrążeń: 113622)
This confirms the diagnosis: a model trained on each era separately performs better than one trained on everything together. Mixing the two eras gives a worse result than either alone - adding seasons from a different technical era adds only noise, not signal. The model receives contradictory patterns (different tire degradation, a different balance of power between teams) under the same features. Team suffers most: it's a good indicator of car strength only within one era, and when years are mixed, the same team name means an entirely different pace.
The conclusion is important and counterintuitive: the problem is not the amount of data, but the lack of features. The ~0.55 ceiling is a limit of the current feature set, not a consequence of too small a sample - this is clear from the fact that seven seasons perform worse than one or two. Lap pace depends on factors we don't have in the data - traffic (following a rival), engine mode, or a driver's form on the day. To raise this ceiling, we need new features, not more seasons of the same data.
Saving predictions and benchmark results¶
We save the predictions and the results table to parquet files - they will be used by notebook 05 (error analysis).
We save the predictions of HistGradientBoosting - the model we used for diagnostics - rather than those of the benchmark winner (LinearRegression). This costs little (MAE = 0.627 s vs 0.585 s), and in return the error analysis in notebook 05 concerns the same model whose feature importances and predicted-vs-actual plot we have already examined.
out = laps.iloc[te][['race_id', 'EventName', 'Compound', 'TyreLife', 'LapNumber']].copy()
out['actual_rel'] = yte
out['pred_rel'] = p_best
out['base_laptime'] = laps['base'].values[te]
out.to_parquet('../data/reg_predictions.parquet', index=False)
bench.to_parquet('../data/reg_benchmark.parquet')
print('zapisano reg_predictions.parquet', out.shape, 'oraz reg_benchmark.parquet')
zapisano reg_predictions.parquet (10477, 8) oraz reg_benchmark.parquet
Summary¶
Three observations:
- A naive model on one season achieved $R^2 = -0.76$ - worse than the mean. The fault lay not with the model, but with the task: the track (
~99%of the variability) was unique to each race, so with a per-race split the test tracks remained unknown. - Adding a second season raised $R^2$ to
~0.90. Sometimes the fix is data, not the model. - A high $R^2$ can be deceptive, though - the model mainly learned the track, which we know in advance anyway. After changing the target to pace relative to a race, $R^2$ dropped to
~0.5, but the model began learning what matters: the fuel effect, car performance (Team), and tire degradation. The weather features turned out to be nearly useless for a target framed this way - within a race the weather hardly changes, so it presumably differentiates races from one another rather than laps within a single race.
The benchmark additionally showed that with a properly framed task, a simple linear model can match tree-based models while running far faster.