Error analysis¶
Where and why do our models get it wrong?
The insights come from both the regression model's residuals and the classifier's mistakes.
import numpy as np
import pandas as pd
import plotly.express as px
reg = pd.read_parquet('../data/reg_predictions.parquet')
clf = pd.read_parquet('../data/clf_predictions.parquet')
reg['residual'] = reg['actual_rel'] - reg['pred_rel']
print('reg:', reg.shape, '| clf:', clf.shape)
reg: (10477, 9) | clf: (918, 14)
Regression - residual distribution¶
We analyze the predictions of HistGradientBoosting saved in notebook 03 - the model used there for diagnostics as the representative of the nonlinear approach (the benchmark winner, LinearRegression, had an MAE lower by ~0.04 s).
A residual is the difference between actual and predicted values (for relative pace). A good model produces a symmetric cloud around zero, without systematic bias.
mae = reg['residual'].abs().mean()
bias = reg['residual'].mean()
print(f'MAE = {mae:.3f} s | bias = {bias:+.3f} s')
fig = px.histogram(reg, x='residual', nbins=80,
title=f'Rozkład residuów (MAE={mae:.3f}s, bias={bias:+.3f}s)',
labels={'residual': 'residuum [s]'})
fig.add_vline(x=0, line_dash='dash')
fig.show()
MAE = 0.627 s | bias = +0.044 s
The histogram shows the residual distribution - the horizontal axis shows the residual (the difference between actual and predicted relative pace in seconds), and bar height represents the number of laps falling in each bin. The dashed vertical line at zero marks the point around which a good model should be symmetric.
We see two things. First, MAE = 0.627 s - the model errs by over six tenths of a second on average, consistent with the HistGradientBoosting result in the benchmark from notebook 03 (0.627 s; the winning linear model had 0.585 s).
Second, and diagnostically more important, bias = +0.044 s is close to zero. This is the mean of the residuals: if it were clearly positive, the model would systematically under-predict pace; if negative - over-predict it. Here it is close to zero, so the model has no systematic shift.
The distribution shape is unimodal and concentrated around zero, but with noticeable tails in both directions - these are the individual, severely mispredicted laps we'll examine below.
Is there a pattern? Residuals vs. tire age¶
If the error grew with tire age, it would mean the model didn't fully capture degradation. We want to see a cloud around zero (without a clear trend).
fig = px.scatter(reg, x='TyreLife', y='residual', color='Compound', opacity=0.3,
title='Residua vs wiek opony (chcemy chmurę wokół 0, bez trendu)',
labels={'TyreLife': 'wiek opony [okr.]', 'residual': 'residuum [s]'})
fig.add_hline(y=0, line_dash='dash')
fig.show()
corr = reg['TyreLife'].corr(reg['residual'])
print(f'korelacja TyreLife - residuum: {corr:+.3f}')
wiek = pd.cut(reg['TyreLife'], [0, 5, 10, 20, 30, 100])
reg.groupby(wiek, observed=True)['residual'].agg(mediana='median', srednia='mean', n='size').round(3)
korelacja TyreLife - residuum: -0.116
| mediana | srednia | n | |
|---|---|---|---|
| TyreLife | |||
| (0, 5] | 0.116 | 0.143 | 1659 |
| (5, 10] | 0.167 | 0.192 | 2253 |
| (10, 20] | -0.007 | 0.007 | 3645 |
| (20, 30] | -0.049 | -0.066 | 2040 |
| (30, 100] | -0.111 | -0.115 | 880 |
Each point is one test lap: on the horizontal axis is tire age (TyreLife, the number of laps driven on a given set), on the vertical axis is the residual, and color indicates the compound. The dashed line at zero represents perfect prediction.
Does the error grow with tire age? If the cloud tilted upward or downward with older tires, it would mean the model didn't fully capture degradation and left a wear-related pattern in the residuals.
We see, however, a cloud distributed essentially symmetrically around zero across the full width of the axis, without a clear tilt. The correlation between TyreLife and the residual, computed above, confirms this: -0.116, i.e. near zero - the trend is weak (it explains about 1% of the residual variance), though with ten thousand laps it is statistically detectable. This is exactly what we wanted: the model captured tire degradation well enough that only a faint trace of it remains in the residuals.
The table above shows this subtle residual pattern: laps on fresh tires are slightly underestimated (median residual +0.12 to +0.17 s for tires up to ten laps old - the model predicts slightly too fast a pace), and laps on very old tires slightly overestimated (median -0.11 s above thirty laps). But this is minor - at a correlation of about 0.12, it is practically insignificant. The symmetry observation thus holds: what remains in the residuals is mostly random scatter, not missed degradation.
by_comp = (reg.assign(abs_err=reg['residual'].abs())
.groupby('Compound')['abs_err'].agg(['mean', 'count']).reset_index())
fig = px.bar(by_comp, x='Compound', y='mean',
title='Średni błąd bezwzględny per mieszanka opon',
labels={'mean': 'MAE [s]', 'Compound': 'mieszanka'})
fig.show()
by_comp
| Compound | mean | count | |
|---|---|---|---|
| 0 | HARD | 0.615374 | 5728 |
| 1 | MEDIUM | 0.620592 | 3624 |
| 2 | SOFT | 0.709311 | 1125 |
The bars show the model's mean absolute error split by tire compound - bar height is MAE in seconds for laps driven on each compound. The table below adds the number of laps for each type.
We see a clear difference: on soft tires the model errs most (SOFT MAE = 0.709 s), distinctly more than on hard (HARD 0.615 s) and medium (MEDIUM 0.621 s). The explanation lies in the count column.
There are only 1125 laps on SOFT, while HARD has 5728 and MEDIUM 3624. Soft tires mean less data to learn from, plus they're used differently: in short, aggressive stints (a stint being a run on a single set of tires), where they degrade faster and pace changes more sharply from lap to lap.
Fewer examples and higher internal variability combine to make this a harder-to-predict and thus more error-prone data segment.
This is the first clue as to where to focus further model improvements.
Worst-predicted laps¶
Now we'll review the model's worst predictions.
worst = reg.reindex(reg['residual'].abs().sort_values(ascending=False).index).head(15)
worst[['EventName', 'Compound', 'TyreLife', 'LapNumber',
'actual_rel', 'pred_rel', 'residual']].round(2)
| EventName | Compound | TyreLife | LapNumber | actual_rel | pred_rel | residual | |
|---|---|---|---|---|---|---|---|
| 679 | Azerbaijan Grand Prix | HARD | 3.0 | 14.0 | 3.35 | -0.25 | 3.60 |
| 4766 | Australian Grand Prix | HARD | 2.0 | 43.0 | 1.81 | -1.58 | 3.39 |
| 6245 | Japanese Grand Prix | HARD | 7.0 | 39.0 | 2.89 | -0.47 | 3.37 |
| 2605 | Spanish Grand Prix | MEDIUM | 9.0 | 48.0 | 1.73 | -1.64 | 3.36 |
| 5513 | Australian Grand Prix | HARD | 6.0 | 42.0 | 2.57 | -0.78 | 3.35 |
| 8353 | Singapore Grand Prix | HARD | 31.0 | 59.0 | 2.15 | -1.16 | 3.31 |
| 8787 | Singapore Grand Prix | SOFT | 8.0 | 54.0 | 1.95 | -1.27 | 3.22 |
| 8737 | Singapore Grand Prix | SOFT | 19.0 | 56.0 | 1.89 | -1.29 | 3.18 |
| 10347 | Qatar Grand Prix | HARD | 9.0 | 43.0 | 1.71 | -1.37 | 3.08 |
| 6105 | Japanese Grand Prix | SOFT | 18.0 | 52.0 | 2.49 | -0.57 | 3.06 |
| 1801 | Miami Grand Prix | HARD | 35.0 | 40.0 | 2.14 | -0.90 | 3.04 |
| 719 | Azerbaijan Grand Prix | HARD | 3.0 | 9.0 | 2.60 | -0.43 | 3.03 |
| 5559 | Australian Grand Prix | HARD | 30.0 | 38.0 | 2.59 | -0.41 | 3.00 |
| 10280 | Qatar Grand Prix | HARD | 21.0 | 56.0 | -3.59 | -0.59 | -3.00 |
| 4528 | Dutch Grand Prix | SOFT | 2.0 | 11.0 | 2.57 | -0.39 | 2.96 |
The table lists the fifteen laps where the model erred the most (sorted by residual magnitude). For each, we see the track, compound, tire age, lap number, the actual and predicted pace, and the residual itself.
Nearly all these laps were actually slow (actual_rel on the order of +1.7 to +3.4 s above the race median)!
The model consistently predicted pace near zero or even negative (pred_rel from -0.25 to -1.64 s).
This is the behavior already noted in notebook 03 (the predicted/actual plot). Unable to find the cause of an extremely slow lap, the model guesses something close to the average pace. Similar behavior appears for the Qatar Grand Prix lap with a residual of -3.00, where the pace was exceptionally fast (-3.59 s) but the model expected an average one.
What these cases have in common is that they are extreme outliers driven by factors outside our features - traffic on track, transient race situations, the aftermath of safety car periods. There's no signal in the data to predict them - they unfortunately top the error list.
# wnioski regresji LICZONE z danych
worst_track = reg.assign(ae=reg.residual.abs()).groupby('EventName')['ae'].mean().idxmax()
worst_comp = by_comp.loc[by_comp['mean'].idxmax(), 'Compound']
tail = (reg['residual'].abs() > 3 * reg['residual'].abs().median()).mean()
print('Tor z największym średnim błędem:', worst_track)
print('Mieszanka z największym błędem: ', worst_comp)
print(f'Odsetek okrążeń z błędem >3x mediany (ogon): {tail:.1%}')
Tor z największym średnim błędem: Qatar Grand Prix Mieszanka z największym błędem: SOFT Odsetek okrążeń z błędem >3x mediany (ogon): 6.3%
The track with the largest mean error is Qatar Grand Prix, the most difficult compound to predict is SOFT, and the share of laps with an error exceeding three times the median is 6.3%.
This means roughly six laps in every hundred are predicted distinctly worse than a typical one.
The model thus handles most laps well, with its trouble concentrated in a small but recognizable group: a tough track (Qatar), a fragile compound (SOFT), and rare extremes whose causes aren't in the features.
The limit of the model's quality therefore lies not in the average lap, but in this narrow tail of atypical situations.
Classification - where the model errs¶
The main theme is errors near the boundary. That is, drivers in fourth or fifth place whom the model predicted to finish on the podium, and vice versa.
clf['err_type'] = np.select(
[(clf.podium == 1) & (clf.pred_logit == 1),
(clf.podium == 0) & (clf.pred_logit == 0),
(clf.podium == 0) & (clf.pred_logit == 1),
(clf.podium == 1) & (clf.pred_logit == 0)],
['TP (trafione podium)', 'TN (trafione nie-podium)',
'FP (typował podium, nie było)', 'FN (przegapił podium)'],
default='?')
clf['err_type'].value_counts()
err_type TN (trafione nie-podium) 624 FP (typował podium, nie było) 156 TP (trafione podium) 123 FN (przegapił podium) 15 Name: count, dtype: int64
Classification - anatomy of errors¶
Above, the confusion matrix from notebook 04 reappears, this time as a table.
TP(true positive) is a correctly predicted podium,TN(true negative) is a correctly predicted non-podium,FP(false positive) is when the model predicted a podium that didn't happen,FN(false negative) is a missed podium.
TN = 624, TP = 123, and among errors FP = 156 far exceeds FN = 15.
This means the model prefers to raise false alarms (predict a podium that doesn't happen) rather than miss true podiums.
As mentioned in notebook 04, this is a result of class_weight='balanced'.
With only ~15% podiums in the data, this option forces the model to take the rare class seriously, so it readily flags podium candidates.
This shows up in two metrics:
- Recall (sensitivity, what fraction of true podiums was caught) is
123 / (123 + 15) = 0.89- the model misses very few podiums. - Precision (what fraction of predicted podiums proved correct) is only
123 / (123 + 156) = 0.44- less than half of the predictions pan out.
miss = clf[clf.podium != clf.pred_logit].sort_values('proba_logit', ascending=False)
miss[['EventName', 'Driver', 'Team', 'GridPosition', 'Position',
'proba_logit', 'podium', 'err_type']].round(3).head(20)
| EventName | Driver | Team | GridPosition | Position | proba_logit | podium | err_type | |
|---|---|---|---|---|---|---|---|---|
| 642 | Austrian Grand Prix | VER | Red Bull Racing | 1.0 | 5.0 | 0.964 | 0 | FP (typował podium, nie było) |
| 704 | Belgian Grand Prix | PER | Red Bull Racing | 2.0 | 7.0 | 0.959 | 0 | FP (typował podium, nie było) |
| 497 | Australian Grand Prix | VER | Red Bull Racing | 1.0 | 19.0 | 0.954 | 0 | FP (typował podium, nie było) |
| 823 | Mexico City Grand Prix | VER | Red Bull Racing | 2.0 | 6.0 | 0.953 | 0 | FP (typował podium, nie było) |
| 682 | Hungarian Grand Prix | VER | Red Bull Racing | 3.0 | 5.0 | 0.939 | 0 | FP (typował podium, nie było) |
| 903 | Abu Dhabi Grand Prix | VER | Red Bull Racing | 4.0 | 6.0 | 0.919 | 0 | FP (typował podium, nie było) |
| 774 | Azerbaijan Grand Prix | PER | Red Bull Racing | 4.0 | 17.0 | 0.910 | 0 | FP (typował podium, nie było) |
| 723 | Dutch Grand Prix | PER | Red Bull Racing | 5.0 | 6.0 | 0.904 | 0 | FP (typował podium, nie było) |
| 862 | Las Vegas Grand Prix | VER | Red Bull Racing | 5.0 | 5.0 | 0.900 | 0 | FP (typował podium, nie było) |
| 317 | Japanese Grand Prix | PER | Red Bull Racing | 5.0 | 19.0 | 0.895 | 0 | FP (typował podium, nie było) |
| 541 | Miami Grand Prix | PER | Red Bull Racing | 4.0 | 4.0 | 0.888 | 0 | FP (typował podium, nie było) |
| 358 | United States Grand Prix | LEC | Ferrari | 1.0 | 20.0 | 0.888 | 0 | FP (typował podium, nie było) |
| 843 | São Paulo Grand Prix | NOR | McLaren | 1.0 | 6.0 | 0.886 | 0 | FP (typował podium, nie było) |
| 801 | United States Grand Prix | NOR | McLaren | 1.0 | 4.0 | 0.886 | 0 | FP (typował podium, nie było) |
| 378 | Mexico City Grand Prix | PER | Red Bull Racing | 5.0 | 20.0 | 0.878 | 0 | FP (typował podium, nie było) |
| 762 | Azerbaijan Grand Prix | VER | Red Bull Racing | 6.0 | 5.0 | 0.871 | 0 | FP (typował podium, nie było) |
| 124 | Spanish Grand Prix | SAI | Ferrari | 2.0 | 5.0 | 0.862 | 0 | FP (typował podium, nie było) |
| 721 | Dutch Grand Prix | PIA | McLaren | 3.0 | 4.0 | 0.858 | 0 | FP (typował podium, nie było) |
| 907 | Abu Dhabi Grand Prix | PIA | McLaren | 2.0 | 10.0 | 0.855 | 0 | FP (typował podium, nie było) |
| 424 | Abu Dhabi Grand Prix | PIA | McLaren | 3.0 | 6.0 | 0.855 | 0 | FP (typował podium, nie było) |
The table shows the model's errors sorted in descending order of assigned podium probability.
Each row is a case where the prediction missed reality - we see the track, driver, team, grid and finishing positions, probability, and error type.
The entire top of the list consists of false alarms (FP) with very high probability (proba_logit from 0.86 to 0.96).
They share one pattern: they're well-known drivers from leading teams (Red Bull Racing, Ferrari, McLaren) who started up front (GridPosition from 1 to 6).
Some finished just off the podium (Position 4-7).
Quite a few names, however, appear with finishing positions of 17, 19 or 20 - races in which something completely unpredictable happened (a failure, a collision, a penalty).
The common thread is that, based on pre-race features, they looked like sure podium contenders, and the model rightly gave them a high probability. Unfortunately, the race itself - from losing one or two places to a complete breakdown - overturned that forecast.
The model didn't pick random drivers here - it identified the right favorites, but race-day events not in the features decided the outcome.
Calibration - do the probabilities mean anything?¶
Good calibration means that among cases with a predicted probability of ~0.7, roughly 70% are actually podiums.
In the plot below, the points should cluster near the diagonal.
clf['bin'] = pd.cut(clf['proba_logit'], bins=np.linspace(0, 1, 11))
cal = (clf.groupby('bin', observed=True)
.agg(srednie_P=('proba_logit', 'mean'),
udzial_podiow=('podium', 'mean'),
n=('podium', 'size')).dropna().reset_index())
fig = px.scatter(cal, x='srednie_P', y='udzial_podiow', size='n',
title='Krzywa kalibracji (regresja logistyczna)',
labels={'srednie_P': 'średnie P(podium)', 'udzial_podiow': 'realny udział podiów'})
fig.add_shape(type='line', x0=0, y0=0, x1=1, y1=1, line=dict(dash='dash', color='red'))
fig.show()
ece = (cal['n'] / cal['n'].sum() * (cal['srednie_P'] - cal['udzial_podiow']).abs()).sum()
print(f'ECE = {ece:.3f} | średnie P(podium) = {clf.proba_logit.mean():.3f} | realny odsetek podiów = {clf.podium.mean():.3f}')
cal.round(3)
ECE = 0.153 | średnie P(podium) = 0.304 | realny odsetek podiów = 0.150
| bin | srednie_P | udzial_podiow | n | |
|---|---|---|---|---|
| 0 | (0.0, 0.1] | 0.027 | 0.007 | 449 |
| 1 | (0.1, 0.2] | 0.142 | 0.051 | 59 |
| 2 | (0.2, 0.3] | 0.255 | 0.041 | 49 |
| 3 | (0.3, 0.4] | 0.346 | 0.024 | 42 |
| 4 | (0.4, 0.5] | 0.450 | 0.150 | 40 |
| 5 | (0.5, 0.6] | 0.551 | 0.156 | 45 |
| 6 | (0.6, 0.7] | 0.648 | 0.238 | 42 |
| 7 | (0.7, 0.8] | 0.749 | 0.352 | 71 |
| 8 | (0.8, 0.9] | 0.854 | 0.568 | 74 |
| 9 | (0.9, 1.0] | 0.952 | 0.830 | 47 |
The plot examines calibration - whether the model's probabilities can be read literally.
We group the probabilities into ten bins - for each, the horizontal axis shows the mean P(podium) and the vertical axis shows the actual podium rate. Dot size is the bin count. The dashed diagonal is perfect calibration: among cases with a predicted probability of ~0.7, roughly 70% should end up on the podium.
Points lie clearly below the diagonal - the model systematically overestimates probabilities. When it says 0.55, a podium actually occurs in only ~16% of cases; at 0.75, in ~35%.
ECE (expected calibration error), the bin-count-weighted average calibration error, is 0.153 - the same value the experiment at the end of notebook 04 gave for the raw balanced model (printed there to four decimal places as 0.1535). The mean predicted P(podium) is 0.304 against an actual podium share of 0.150: the model promises a podium roughly twice as often as one actually occurs. class_weight='balanced' artificially inflates the probabilities of the minority class in order to catch it. A side effect is the model's overconfidence. Practical takeaway: these probabilities must not be read literally as "an X% chance."
But there's another side. The curve rises almost monotonically - a higher P(podium) corresponds to genuinely higher odds. The only exceptions are two small bins at 0.2-0.4 (~40-50 cases each), where the podium rate dips slightly; with counts that small this is more likely noise than a pattern. So the model ranks candidates well (a higher score means a better candidate), but misprices absolute probability.
We have good discrimination (ROC-AUC = 0.92) with poor calibration. The largest bin is 0-0.1 (449 cases, nearly half the data, actual podium rate ~0.01), where the assignment is accurate - the model confidently filters out hopeless cases.
# najwieksza niespodzianka: podium z dalekiej pozycji startowej
upset = clf[(clf.podium == 1)].sort_values('GridPosition', ascending=False)
print('Podia z najdalszych pozycji startowych (model słusznie je przegapia):')
print(upset[['EventName', 'Driver', 'GridPosition', 'Position', 'proba_logit']].head(5).to_string(index=False))
Podia z najdalszych pozycji startowych (model słusznie je przegapia):
EventName Driver GridPosition Position proba_logit
Abu Dhabi Grand Prix LEC 19.0 3.0 0.006586
São Paulo Grand Prix VER 17.0 1.0 0.107751
Saudi Arabian Grand Prix VER 15.0 2.0 0.175216
Austrian Grand Prix PER 15.0 3.0 0.266306
São Paulo Grand Prix GAS 13.0 3.0 0.084385
The final table picks out real surprises: podiums scored from the furthest-back grid positions. This is an extreme test for the model - can it admit that some outcomes can't be predicted from pre-race features alone?
LECfromP19toP3with a probability of just0.007,VERfromP17to victory withproba 0.108,VERfromP15toP2(0.175).
These are races where in-race factors decided everything. The model rightly assigned these drivers very low probabilities, since nothing at the start suggested a podium.
Final conclusions¶
Regression (lap pace)
- After switching the target to relative pace, the model works well. The largest errors remain on laps with atypical conditions that survived filtering - for example, traffic or the aftermath of a safety car.
- Error magnitude depends on the tire compound - hard and soft tires differ in the spread of their degradation.
Classification (podium)
- The model captures the rule "starting up front boosts podium odds" but errs where F1 can be unpredictable: rain, safety cars, mechanical issues, bold strategies.
- These errors are a natural limit of pre-race prediction - they can't be eliminated without in-race data, which would be data leakage.
- The
balancedmodel ranks candidates well (ROC-AUC = 0.92), but its probabilities are overestimated (ECE = 0.153,0.30on average against an actual0.15). As the experiment at the end of notebook 04 showed, class weights barely change the ordering of candidates (ROC-AUC0.920-0.922), only the operating point at the0.5threshold, and calibration (Platt or isotonic) removes the overestimation with practically no change in order (ROC-AUCunchanged,PR-AUCslightly lower). Ranking needs only discrimination; calibration becomes necessary only when the probabilities are to be read literally.
Methodological takeaway
- Two failures from earlier notebooks ($R^2 < 0$ and
quali_gap_sentirely empty) were instructive - they showed the outcome depends more on problem formulation and data quality than on algorithm choice. - The timing benchmark showed that with small datasets, simpler models can be both better and cheaper.