Classification: will the driver finish on the podium?¶
Can we predict the podium based on what we know before the race starts?
We use pre-race features: starting position, team, and time gap to the fastest qualifying lap.
import time
import numpy as np
import pandas as pd
import plotly.express as px
import plotly.graph_objects as go
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.dummy import DummyClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.ensemble import RandomForestClassifier, HistGradientBoostingClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from sklearn.metrics import (accuracy_score, confusion_matrix, classification_report,
roc_auc_score, average_precision_score,
roc_curve, precision_recall_curve)
dr = pd.read_parquet('../data/driver_race.parquet')
dr = dr.dropna(subset=['GridPosition']).reset_index(drop=True)
print('wierszy:', len(dr), '| podiów:', int(dr.podium.sum()), f'({dr.podium.mean():.1%})')
print('sezony:', sorted(dr.Season.unique()))
wierszy: 918 | podiów: 138 (15.0%) sezony: [np.int64(2023), np.int64(2024)]
Features and validation method¶
We use only pre-race features. The quali_gap_s feature is included only if it is meaningfully populated - in the original project version it was completely empty (we described this in notebook 01, Step 1).
Since there is little data and the classes are imbalanced, instead of a single split we use stratified cross-validation (StratifiedKFold) and evaluate the model on out-of-fold predictions (each row's prediction comes from a model that did not see it during training).
Let us explain stratified cross-validation, as it is key to fair evaluation on such a small dataset. Cross-validation splits the data into k = 5 equal parts. We train the model on four of them and test on the fifth - and repeat this five times, each time setting aside a different part as the test set. As a result, each row is in the test set exactly once, and its prediction always comes from a model that did not see it during training (hence out-of-fold). This is much fairer than a single split, because we evaluate the model on all data, not on one randomly chosen quarter.
Stratification also ensures that each of the five parts has the same percentage of podiums (~15%). Without it, with such a rare class, one part could end up without a single podium - and evaluation on it would be worthless. Stratification guarantees that the minority class is represented evenly in each fold.
cat_cols = ['Team']
num_cols = ['GridPosition']
if dr['quali_gap_s'].notna().sum() > 0.5 * len(dr):
num_cols.append('quali_gap_s')
dr['quali_gap_s'] = dr['quali_gap_s'].fillna(dr['quali_gap_s'].median())
print('quali_gap_s UŻYTE jako cecha (wypełnione:', dr.quali_gap_s.notna().sum(), ')')
else:
print('quali_gap_s pominięte - za dużo braków')
features = cat_cols + num_cols
X = dr[features].copy()
y = dr['podium'].values
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
print('cechy:', features)
quali_gap_s UŻYTE jako cecha (wypełnione: 918 ) cechy: ['Team', 'GridPosition', 'quali_gap_s']
Why accuracy alone is misleading in this scenario¶
Imagine the simplest possible "model": it always claims there will be no podium. Let's see what accuracy it achieves:
zawsze_nie = np.zeros_like(y)
print(f'„zawsze NIE podium" - accuracy = {accuracy_score(y, zawsze_nie):.1%}')
print('...model, który nie przewiduje NICZEGO, ma ~85% trafności.')
print('Dlatego accuracy jest tu bezużyteczne - liczą się precision/recall, ROC-AUC, PR-AUC.')
„zawsze NIE podium" - accuracy = 85.0% ...model, który nie przewiduje NICZEGO, ma ~85% trafności. Dlatego accuracy jest tu bezużyteczne - liczą się precision/recall, ROC-AUC, PR-AUC.
Data leakage¶
If we wanted to artificially boost the score, we could add a feature from the race itself - for example, the points scored:
# CELOWY BLAD: dorzucamy 'Points'
X_leak = dr[features + ['Points']].copy()
pre_leak = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', StandardScaler(), num_cols + ['Points'])])
leak_model = Pipeline([('pre', pre_leak),
('clf', LogisticRegression(class_weight='balanced', max_iter=1000))])
proba_leak = cross_val_predict(leak_model, X_leak, y, cv=cv, method='predict_proba')[:, 1]
print(f'Z wyciekiem (Points): ROC-AUC = {roc_auc_score(y, proba_leak):.3f} <- idealne')
Z wyciekiem (Points): ROC-AUC = 1.000 <- idealne
The result looks perfect - ROC-AUC close to 1.0.
Points (Points) are awarded for the race result - whoever scored many points by definition finished high up. The model does not predict the podium so much as read it off an almost ready-made answer. As in the previous notebook - we stick only to pre-race features.
Main model - logistic regression¶
We return to honest features. The class_weight='balanced' parameter compensates for the numerical advantage of the "non-podium" class, and we evaluate the model on out-of-fold predictions.
pre = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', StandardScaler(), num_cols)])
logit = Pipeline([('pre', pre),
('clf', LogisticRegression(class_weight='balanced', max_iter=1000))])
proba_logit = cross_val_predict(logit, X, y, cv=cv, method='predict_proba')[:, 1]
pred_logit = (proba_logit >= 0.5).astype(int)
print('Regresja logistyczna (cechy przedwyścigowe):')
print(f' ROC-AUC = {roc_auc_score(y, proba_logit):.3f} PR-AUC = {average_precision_score(y, proba_logit):.3f}')
print(classification_report(y, pred_logit, target_names=['nie podium', 'podium']))
Regresja logistyczna (cechy przedwyścigowe):
ROC-AUC = 0.921 PR-AUC = 0.703
precision recall f1-score support
nie podium 0.98 0.80 0.88 780
podium 0.44 0.89 0.59 138
accuracy 0.81 918
macro avg 0.71 0.85 0.73 918
weighted avg 0.90 0.81 0.84 918
ROC-AUC = 0.921 and PR-AUC = 0.703 - what does this mean?
recall (sensitivity) for the podium class is 0.89 - out of all actual podiums, the model catches nearly nine out of ten.
precision is only 0.44 - when the model declares "podium", it is right about half the time.
recall tells us what fraction of true podiums we managed to catch, while precision tells us what fraction of our podium predictions turned out correct. f1-score = 0.59 is the harmonic mean of both, one metric reconciling sensitivity and precision.
class_weight='balanced' makes the model eager to predict a podium so as not to miss a real one (high recall), and it pays for this with frequent false alarms (low precision).
For the non-podium class, the proportions are reversed and favorable - precision = 0.98, recall = 0.80. Note that overall accuracy = 0.81 is even lower than the ~85% of the "always no" model from the first pitfall.
Yet this model is incomparably more useful. The "always no" model has 85% accuracy, but its podium recall is 0 - it does not detect a single podium, so it is useless for anything. Our model catches 89% of podiums, so it actually answers the question "who has a chance at the podium", at the price of false alarms. accuracy is misleading precisely because it rewards passivity toward the numerous non-podium class: the idle "always no" looks better by this measure than a model that actually predicts something.
How should we understand ROC-AUC and PR-AUC?¶
ROC-AUC randomly picks one driver who was on the podium and one who was not.
ROC-AUC is the probability that the model assigned a higher chance to the correct one. Let's verify:
pod = proba_logit[y == 1] # prawdopodobieństwa dla prawdziwych podiów
nie = proba_logit[y == 0] # prawdopodobieństwa dla nie-podiów
roznice = pod[:, None] - nie[None, :] # każda para (podium, nie-podium)
auc_z_par = (roznice > 0).mean() + 0.5 * (roznice == 0).mean()
print(f'par (podium x nie-podium): {roznice.size}')
print(f'model dał wyższą szansę podium w: {auc_z_par:.1%} par')
print(f'ROC-AUC ze sklearn: {roc_auc_score(y, proba_logit):.3f} (to samo)')
par (podium x nie-podium): 107640 model dał wyższą szansę podium w: 92.1% par ROC-AUC ze sklearn: 0.921 (to samo)
So ROC-AUC = 0.5 would mean a coin flip (the model does not distinguish podiums from the rest), and 1.0 a perfect model. Our ~0.92 means that in more than nine out of ten cases the model correctly indicates which of two drivers had a better chance at the podium.
And PR-AUC? It is the area under the precision-recall curve. It measures how well the model detects a rare class - here the podium - at different cutoff thresholds: for each threshold we read a pair (precision, recall), and PR-AUC collects them into one number. The baseline for it is the percentage of podiums itself (~0.15) - that is what random guessing would achieve - so our ~0.70 is a result well above chance.
Unlike ROC-AUC, PR-AUC does not reward the model for easily getting the numerous non-podium class right: it focuses solely on how accurately and completely we catch the rare podiums. Therefore, with imbalanced classes, it is a harsher and more demanding measure than ROC-AUC.
Benchmark - classifier comparison¶
We compare several algorithms on the same data and the same validation. We record ROC-AUC, PR-AUC, and full cross-validation time. DummyClassifier, always predicting the majority class, is the baseline.
def make(est, scale=True):
steps = [('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols)]
steps.append(('num', StandardScaler() if scale else 'passthrough', num_cols))
return Pipeline([('pre', ColumnTransformer(steps)), ('clf', est)])
models = {
'Dummy (większość)': make(DummyClassifier(strategy='most_frequent')),
'Logistic Regression': make(LogisticRegression(class_weight='balanced', max_iter=1000)),
'kNN (k=15)': make(KNeighborsClassifier(n_neighbors=15)),
'Random Forest': make(RandomForestClassifier(n_estimators=300, class_weight='balanced',
n_jobs=-1, random_state=42), scale=False),
'HistGradientBoosting': make(HistGradientBoostingClassifier(
learning_rate=0.05, max_iter=200, l2_regularization=1.0,
class_weight='balanced', random_state=42), scale=False),
}
rows = []
for name, model in models.items():
t = time.perf_counter()
try:
proba = cross_val_predict(model, X, y, cv=cv, method='predict_proba')[:, 1]
auc = roc_auc_score(y, proba); ap = average_precision_score(y, proba)
except Exception:
pred = cross_val_predict(model, X, y, cv=cv)
auc = roc_auc_score(y, pred); ap = average_precision_score(y, pred)
dt = time.perf_counter() - t
rows.append(dict(model=name, ROC_AUC=auc, PR_AUC=ap, cv_time_s=dt))
print(f'{name:<22} ROC-AUC={auc:.3f} PR-AUC={ap:.3f} czas_CV={dt:.2f}s')
bench = pd.DataFrame(rows).set_index('model').round(3)
bench
Dummy (większość) ROC-AUC=0.500 PR-AUC=0.150 czas_CV=0.03s
Logistic Regression ROC-AUC=0.921 PR-AUC=0.703 czas_CV=0.05s kNN (k=15) ROC-AUC=0.918 PR-AUC=0.666 czas_CV=0.04s
Random Forest ROC-AUC=0.907 PR-AUC=0.607 czas_CV=1.72s
HistGradientBoosting ROC-AUC=0.912 PR-AUC=0.643 czas_CV=0.82s
| ROC_AUC | PR_AUC | cv_time_s | |
|---|---|---|---|
| model | |||
| Dummy (większość) | 0.500 | 0.150 | 0.032 |
| Logistic Regression | 0.921 | 0.703 | 0.048 |
| kNN (k=15) | 0.918 | 0.666 | 0.043 |
| Random Forest | 0.907 | 0.607 | 1.723 |
| HistGradientBoosting | 0.912 | 0.643 | 0.816 |
The table compares all classifiers on the same data and with the same cross-validation. Each row is one model, and the columns are ROC-AUC, PR-AUC, and cv_time_s (full 5-fold validation time in seconds).
The baseline is Dummy (majority). Its ROC-AUC = 0.500 is chance level - a coin flip. A model that doesn't distinguish podium from the rest at all. Meanwhile, PR-AUC = 0.150 is the percentage of podiums in the data.
A useful model should beat both of these baselines - and all the other models do, with ROC-AUC ranging from 0.907 to 0.921. A detailed interpretation of the differences between models follows below the chart.
b = bench.reset_index()
fig = px.scatter(b, x='cv_time_s', y='ROC_AUC', text='model', color='model',
title='ROC-AUC (wyżej=lepiej) vs czas walidacji',
labels={'cv_time_s': 'czas 5-fold CV [s]', 'ROC_AUC': 'ROC-AUC'})
fig.update_traces(textposition='top center', marker_size=12)
fig.update_layout(showlegend=False)
fig.show()
On a small dataset (two seasons, only 138 podiums) simple logistic regression performs as well as or better than tree models - and is faster and easier to interpret.
Main model diagnostics¶
We look at three things: the confusion matrix, the ROC and Precision-Recall curves, and finally the model coefficients - to see which features most strongly influence the chance of a podium.
cm = confusion_matrix(y, pred_logit)
fig = px.imshow(cm, text_auto=True, color_continuous_scale='Blues',
x=['pred: nie', 'pred: podium'],
y=['rzecz.: nie', 'rzecz.: podium'],
title='Macierz pomyłek - regresja logistyczna')
fig.update_layout(height=420, width=520)
fig.show()
The confusion matrix breaks down all 918 predictions into four cells. Rows are the actual state, columns are the model's prediction. With podium as the "positive" class, the cells have standard names:
- TN (true negative, correct non-podium):
624- the model said "no" and was right, - FP (false positive, false alarm):
156- the model predicted a podium that did not happen, - FN (false negative, missed):
15- a true podium that the model missed, - TP (true positive, correct podium):
123- the model predicted a podium and there was one.
From the chart we can recompute the numbers in the report above. Podium recall is TP / (TP + FN) = 123 / 138 = 0.89 - out of 138 actual podiums only 15 escaped. Podium precision is TP / (TP + FP) = 123 / 279 = 0.44 - the 123 hits come with as many as 156 false alarms. The model is vigilant, but overzealous: it almost never misses a podium, yet along the way it warns of many that won't happen.
fpr, tpr, _ = roc_curve(y, proba_logit)
prec, rec, _ = precision_recall_curve(y, proba_logit)
fig = go.Figure()
fig.add_trace(go.Scatter(x=fpr, y=tpr, name='ROC (logit)', mode='lines'))
fig.add_trace(go.Scatter(x=[0, 1], y=[0, 1], line=dict(dash='dash'), name='losowy'))
fig.update_layout(title='Krzywa ROC', xaxis_title='FPR (1 - swoistość)', yaxis_title='TPR (czułość)')
fig.show()
fig2 = px.area(x=rec, y=prec, title='Krzywa Precision-Recall',
labels={'x': 'Recall (czułość)', 'y': 'Precision'})
fig2.add_hline(y=y.mean(), line_dash='dash', annotation_text='baza (odsetek podiów)')
fig2.show()
The above charts show how the model behaves at every possible cutoff threshold, not just the default 0.5.
The ROC curve plots sensitivity (TPR, vertical axis - what fraction of true podiums are caught) against the rate of false alarms (FPR, horizontal axis - what fraction of non-podiums are incorrectly labeled as podium). The dashed diagonal is the random model: on it every gain in sensitivity is paid for by an equal increase in false alarms. The more the curve bulges toward the upper left corner, the better - we catch many true podiums at little cost. The area under this curve is ROC-AUC = 0.921!
The Precision-Recall curve shows how precision (how often podium predictions are correct) changes with recall (sensitivity) as the threshold shifts. Here the reference line is a horizontal dashed line at the percentage of podiums (~0.15) - that is what random guessing would achieve. Our curve lies well above that line, and the area under the curve is PR-AUC = 0.703.
logit.fit(X, y)
ohe = logit.named_steps['pre'].named_transformers_['cat']
names = list(ohe.get_feature_names_out(cat_cols)) + num_cols
coefs = logit.named_steps['clf'].coef_[0]
imp = pd.DataFrame({'cecha': names, 'współczynnik': coefs}).sort_values('współczynnik')
fig = px.bar(imp, x='współczynnik', y='cecha', orientation='h',
title='Współczynniki regresji logistycznej (log-szanse)')
fig.update_layout(height=520)
fig.show()
print('Dodatni współczynnik = zwiększa szansę na podium. Ujemny GridPosition = im dalej z tyłu, tym gorzej.')
Dodatni współczynnik = zwiększa szansę na podium. Ujemny GridPosition = im dalej z tyłu, tym gorzej.
This is the most important chart for understanding what drives the model. Each bar is one logistic regression coefficient, expressed as log-odds of a podium. The model has 14 features (twelve one-hot-encoded teams plus GridPosition and quali_gap_s) and we can now see them all on the chart.
Interpretation: a positive coefficient raises the chance of a podium, a negative one lowers it, and the longer the bar, the stronger the feature's effect.
The most strongly negative single numerical feature is GridPosition (-1.88).
As intuition suggests - the further back a driver starts, the smaller the chance of a podium. For teams, the picture is equally clear and they provide the strongest swings.
Red Bull Racing dominates with a coefficient of +2.42 - the strongest positive signal in the entire model, reflecting this car's advantage in the 2023-24 seasons. Also positive are the other top teams: McLaren (+1.25), Ferrari (+1.13), and Mercedes (+0.71). On the opposite side are the tail-end teams: Haas (-1.43), Williams (-1.41), and RB (-1.24), whose bars clearly go negative. The entire gradient of team coefficients reflects the real balance of power on the grid.
quali_gap_s has a small positive coefficient (+0.37), though intuitively a larger gap to the fastest lap should lower the chance.
This is due to strong correlation with GridPosition - qualifying directly sets the grid, so much of this signal is already contained in the starting position, and quali_gap_s adds only a weak, unstable increment.
Chances of a podium are primarily determined by starting position, reinforced by the strength of the team.
Correlation between quali_gap_s and GridPosition¶
An interesting detail on the chart is quali_gap_s, with a fairly small positive coefficient (+0.37). It would seem that a gap to the fastest lap should lower the chance.
The reason may be correlation with GridPosition. After all, qualifying sets the grid - we should check whether the signal is perhaps already contained in the starting position.
from scipy.stats import pearsonr, spearmanr
r_pear = pearsonr(dr['GridPosition'], dr['quali_gap_s'])[0]
r_spear = spearmanr(dr['GridPosition'], dr['quali_gap_s'])[0]
print(f'korelacja Pearsona (liniowa): {r_pear:.3f}')
print(f'korelacja Spearmana (monotoniczna): {r_spear:.3f}')
fig = px.scatter(dr, x='GridPosition', y='quali_gap_s', opacity=0.35,
title='Strata w kwalifikacjach a pozycja startowa',
labels={'GridPosition': 'pozycja startowa',
'quali_gap_s': 'strata do najszybszego okrążenia [s]'})
med = dr.groupby('GridPosition')['quali_gap_s'].median().reset_index()
fig.add_scatter(x=med['GridPosition'], y=med['quali_gap_s'], mode='lines+markers',
name='mediana', line=dict(color='red'))
# przycinamy oś Y - ~4% skrajnych (kierowcy bez czystego czasu w kwalifikacjach) psuje czytelność
fig.update_yaxes(range=[-0.5, 6])
fig.show()
korelacja Pearsona (liniowa): 0.440 korelacja Spearmana (monotoniczna): 0.853
The Spearman correlation is 0.85 - a strong relationship: the further back the starting position, the larger the gap in qualifying. This matches intuition.
The Pearson correlation, however, is much lower (0.44), because the relationship is not linear. Additionally, extreme values distort it (a few percent of drivers who did not set a clean time in qualifying and have a gap of more than ten seconds - we trimmed the vertical axis for readability). The red median line gets steeper and steeper: drivers in the front rows are separated by only a few hundredths (median gap for P1-P3 is ~0.15 s), but in the midfield and at the back the differences grow sharply (P4-P10 is ~0.68 s, and P11-P20 is ~1.63 s).
For this reason, it is hard for the model to assign a stable sign to quali_gap_s. Almost all information contained in this feature is already available through GridPosition.
Will more seasons help?¶
We asked the same question for regression - and it is worth repeating here, because the situation is different. Classification has far less data at its disposal: not tens of thousands of laps, but a few hundred races, of which podiums make up only ~15%. Here more seasons means more podiums - and that is exactly what the model lacks. Let's see if it helps.
We use a broader set driver_race_all (seasons 2018-2024) and calculate quality for an increasing number of seasons, added from the newest backwards. The set was built with the same logic and the same filters as driver_race in notebook 01, but covering seven seasons - just like laps_all in notebook 03:
dr_all = pd.read_parquet('../data/driver_race_all.parquet').dropna(subset=['GridPosition'])
seasons_all = sorted(dr_all['Season'].unique())
print('dostępne sezony:', seasons_all, '| wierszy:', len(dr_all),
'| podiów:', int(dr_all.podium.sum()))
def ocena_clf(df, use_team=True):
df = df.copy()
cat = ['Team'] if use_team else []
num = ['GridPosition']
if df['quali_gap_s'].notna().sum() > 0.5 * len(df):
num.append('quali_gap_s')
df['quali_gap_s'] = df['quali_gap_s'].fillna(df['quali_gap_s'].median())
Xx, yy = df[cat + num], df['podium'].values
steps = ([('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat)]
if cat else [])
steps.append(('num', StandardScaler(), num))
m = Pipeline([('pre', ColumnTransformer(steps)),
('clf', LogisticRegression(class_weight='balanced', max_iter=1000))])
cv5 = StratifiedKFold(5, shuffle=True, random_state=42)
pr = cross_val_predict(m, Xx, yy, cv=cv5, method='predict_proba')[:, 1]
return roc_auc_score(yy, pr), average_precision_score(yy, pr)
rows = []
for k in range(1, len(seasons_all) + 1):
chosen = seasons_all[-k:]
sub = dr_all[dr_all['Season'].isin(chosen)]
auc, ap = ocena_clf(sub)
rows.append({'liczba_sezonów': k, 'zakres': f'{chosen[0]}-{chosen[-1]}',
'podia': int(sub.podium.sum()), 'ROC_AUC': auc, 'PR_AUC': ap})
wynik_clf = pd.DataFrame(rows)
wynik_clf.round(3)
dostępne sezony: [np.int64(2018), np.int64(2019), np.int64(2020), np.int64(2021), np.int64(2022), np.int64(2023), np.int64(2024)] | wierszy: 2956 | podiów: 444
| liczba_sezonów | zakres | podia | ROC_AUC | PR_AUC | |
|---|---|---|---|---|---|
| 0 | 1 | 2024-2024 | 72 | 0.929 | 0.722 |
| 1 | 2 | 2023-2024 | 138 | 0.921 | 0.703 |
| 2 | 3 | 2022-2024 | 204 | 0.927 | 0.696 |
| 3 | 4 | 2021-2024 | 270 | 0.923 | 0.679 |
| 4 | 5 | 2020-2024 | 321 | 0.921 | 0.679 |
| 5 | 6 | 2019-2024 | 384 | 0.923 | 0.687 |
| 6 | 7 | 2018-2024 | 444 | 0.924 | 0.681 |
fig = px.line(wynik_clf, x='liczba_sezonów', y=['ROC_AUC', 'PR_AUC'], markers=True,
title='Jakość klasyfikacji a liczba sezonów',
labels={'liczba_sezonów': 'liczba sezonów', 'value': 'wartość metryki',
'variable': 'metryka'})
fig.update_yaxes(range=[0, 1])
fig.show()
Although the number of podiums grows with each added season, quality practically stands still - ROC-AUC hovers around 0.92, and PR-AUC ~0.70. Interestingly, as in regression, adding data does not improve the result. The reason here is different, though, which is best seen when we compare the technical eras (before and after the 2022 rule change) and check the role of the team:
print('era vs era vs zmieszane:')
for nazwa, df in [('era >=2022', dr_all[dr_all.Season >= 2022]),
('era <2022 ', dr_all[dr_all.Season < 2022]),
('zmieszane ', dr_all)]:
auc, ap = ocena_clf(df)
print(f' {nazwa} podiów={int(df.podium.sum()):>3} ROC-AUC={auc:.3f} PR-AUC={ap:.3f}')
print('\nrola cechy Team:')
for nazwa, ut in [('z Team ', True), ('bez Team', False)]:
auc, ap = ocena_clf(dr_all, use_team=ut)
print(f' {nazwa} ROC-AUC={auc:.3f} PR-AUC={ap:.3f}')
era vs era vs zmieszane: era >=2022 podiów=204 ROC-AUC=0.927 PR-AUC=0.696 era <2022 podiów=240 ROC-AUC=0.921 PR-AUC=0.699 zmieszane podiów=444 ROC-AUC=0.924 PR-AUC=0.681 rola cechy Team:
z Team ROC-AUC=0.924 PR-AUC=0.681 bez Team ROC-AUC=0.916 PR-AUC=0.661
In regression, mixing eras clearly hurt, because Team was the second most important feature and inconsistent across eras ("Mercedes 2019" has a different pace from "Mercedes 2024"). In classification it is different:
- mixing eras doesn't hurt - the result on all seasons is the same as on each era separately,
- removing
Teamchanges almost nothing - the model relies mainly onGridPositionandquali_gap_s.
And these two features are timeless: the rule "whoever starts up front has a chance at the podium" works the same in 2018 and 2024, regardless of regulations. That's why classification is robust to era changes, while regression suffered from them.
In regression the ceiling came from a lack of features describing pace. Here ROC-AUC ~0.92 is already a very high level, and the limit is set by the unpredictability of the race itself and its circumstances.
Saving the predictions¶
We save the main model's (balanced) predictions and the benchmark table to parquet files - notebook 05 (error analysis) will use them. The next two sections are an appendix: experiments with class weights and calibration on the same model.
out = dr.copy()
out['proba_logit'] = proba_logit
out['pred_logit'] = pred_logit
out.to_parquet('../data/clf_predictions.parquet', index=False)
bench.to_parquet('../data/clf_benchmark.parquet')
print('zapisano clf_predictions.parquet', out.shape, 'oraz clf_benchmark.parquet')
zapisano clf_predictions.parquet (918, 14) oraz clf_benchmark.parquet
Class weight selection and probability calibration¶
Throughout this notebook we repeatedly reached for class_weight='balanced', treating this setting as obvious - never checking whether alternatives might behave better. In this section we will try to investigate this.
Our main model - logistic regression on the features Team, GridPosition, and quali_gap_s, evaluated out-of-fold on StratifiedKFold - achieves
ROC-AUC = 0.921 and PR-AUC = 0.703 and this is the baseline against which we will measure each change.
We will explore two questions:
- How do different class weight settings affect the model? We'll compare no weights,
class_weight='balanced', and manually chosen weights. - Can the model's probabilities be read literally? Class weights inflate the minority-class probabilities, so there is a suspicion that the model with
class_weight='balanced'overstates the chance of a podium. We'll check how much - and whether calibration can fix it.
# Porównujemy różne ustawienia `class_weight` na tej samej walidacji out-of-fold.
# `None` - brak wag (klasy traktowane równo), `'balanced'` - wagi odwrotne do liczebności,
# oraz ręczne wagi {0:1, 1:k}, które k-krotnie mocniej karzą przeoczenie podium.
from sklearn.metrics import (roc_auc_score, average_precision_score, confusion_matrix,
precision_score, recall_score, f1_score)
def pre_cw():
# świeży preprocesor dla każdego wariantu (one-hot dla zespołu, skalowanie cech liczbowych)
return ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', StandardScaler(), num_cols)])
# warianty wag klasy: brak, automatyczne 'balanced' oraz trzy ręczne nasilenia
warianty = {
'None (brak wag)': None,
"'balanced'": 'balanced',
'{0:1, 1:3}': {0: 1, 1: 3},
'{0:1, 1:5}': {0: 1, 1: 5},
'{0:1, 1:10}': {0: 1, 1: 10},
}
wiersze = []
for nazwa, cw in warianty.items():
model = Pipeline([('pre', pre_cw()),
('clf', LogisticRegression(class_weight=cw, max_iter=1000))])
# predykcje out-of-fold - każdy wiersz oceniany przez model, który go nie widział
proba = cross_val_predict(model, X, y, cv=cv, method='predict_proba')[:, 1]
pred = (proba >= 0.5).astype(int) # decyzja przy domyślnym progu 0.5
tn, fp, fn, tp = confusion_matrix(y, pred).ravel()
wiersze.append({
'class_weight': nazwa,
'ROC_AUC': roc_auc_score(y, proba),
'PR_AUC': average_precision_score(y, proba),
'precision': precision_score(y, pred, zero_division=0),
'recall': recall_score(y, pred, zero_division=0),
'F1': f1_score(y, pred, zero_division=0),
'TP': tp, 'FP': fp, 'FN': fn, 'TN': tn,
})
tabela_cw = pd.DataFrame(wiersze).set_index('class_weight').round(3)
print('Porównanie opcji class_weight (próg 0.5, predykcje out-of-fold):')
tabela_cw
Porównanie opcji class_weight (próg 0.5, predykcje out-of-fold):
| ROC_AUC | PR_AUC | precision | recall | F1 | TP | FP | FN | TN | |
|---|---|---|---|---|---|---|---|---|---|
| class_weight | |||||||||
| None (brak wag) | 0.920 | 0.713 | 0.706 | 0.522 | 0.600 | 72 | 30 | 66 | 750 |
| 'balanced' | 0.921 | 0.703 | 0.441 | 0.891 | 0.590 | 123 | 156 | 15 | 624 |
| {0:1, 1:3} | 0.922 | 0.708 | 0.534 | 0.804 | 0.642 | 111 | 97 | 27 | 683 |
| {0:1, 1:5} | 0.922 | 0.703 | 0.454 | 0.891 | 0.601 | 123 | 148 | 15 | 632 |
| {0:1, 1:10} | 0.921 | 0.698 | 0.389 | 0.942 | 0.551 | 130 | 204 | 8 | 576 |
Regardless of the weights, ROC-AUC stays in the range 0.920-0.922, and PR-AUC 0.698-0.713.
ROC-AUC and PR-AUC measure ranking - how well the model orders drivers from smallest to largest chance of a podium.
They are computed from the entire range of probabilities, so they don't depend on the 0.5 threshold, and they depend on class weights only marginally: the weights change the fitted coefficients, but - as the table shows - the ordering of drivers barely moves.
class_weight shifts all probabilities up or down, but the order stays practically the same.
What changes? The precision-recall trade-off at the 0.5 threshold:
- no weights (
None):precision = 0.71, butrecall = 0.52. The model predicts podiums conservatively and confidently - out of102predictions it gets72hits (TP) and only30false alarms (FP) - but misses66of138true podiums (FN). 'balanced':precision = 0.44,recall = 0.89. The model catches123podiums and misses only15, but raises156false alarms.- stronger weight
{0:1, 1:10}:recall = 0.94(only8podiums missed), withprecision = 0.39and as many as204false alarms.
The highest F1 = 0.642 comes from the moderate variant {0:1, 1:3} - a compromise between the two extremes.
The more we weight the podium class, the more willing the model is to predict it: recall rises, precision falls, and TP and FP grow together.
Since ROC-AUC and PR-AUC are nearly constant, the model's ability to distinguish podiums stays the same throughout. The weights only express what we want.
Probability calibration¶
A high ROC-AUC says the model orders drivers well - but it says nothing about whether its probabilities are trustworthy. The suspicion is specific: class weighting, which rescues recall, also inflates all predictions, so the model with class_weight='balanced' should systematically overstate the chance of a podium.
The measure of this overestimation is ECE (Expected Calibration Error): we split the predictions into probability bins and average - weighting each bin by its size - the difference between the stated probability and the actual share of podiums.
Below we compute ECE directly, on out-of-fold predictions, for four variants: the raw balanced model, a plain unweighted logit, and two calibration techniques - Platt/sigmoid and isotonic regression (any non-decreasing transformation). We'll also see at what cost: whether improving calibration breaks the ranking measured by ROC-AUC and PR-AUC.
# Eksperyment: kalibracja prawdopodobieństw
# Wysoki ROC-AUC stwierdza, czy model dobrze *porządkuje* kierowców, ale NIE czy jego
# prawdopodobieństwa są wiarygodne.
# Mierzymy ECE: dzielimy prognozy na 10 koszyków i sprawdzamy, jak bardzo średnia deklarowana szansa
# odbiega od rzeczywistego odsetka podiów w koszyku. Im niżej, tym lepiej.
from sklearn.calibration import CalibratedClassifierCV
def ece(p, y, bins=10):
krawedzie = np.linspace(0, 1, bins + 1)
e = 0.0
for i in range(bins):
maska = (p >= krawedzie[i]) & (p < krawedzie[i + 1])
if maska.sum():
e += maska.sum() * abs(y[maska].mean() - p[maska].mean())
return e / len(y)
def baza(balanced):
pre_k = ColumnTransformer([
('cat', OneHotEncoder(handle_unknown='ignore', sparse_output=False), cat_cols),
('num', StandardScaler(), num_cols)])
return Pipeline([('pre', pre_k),
('clf', LogisticRegression(max_iter=1000,
class_weight='balanced' if balanced else None))])
warianty = {
'(a) balanced (baseline)': baza(True),
'(b) zwykły logit': baza(False),
'(c) balanced + Platt': CalibratedClassifierCV(baza(True), method='sigmoid', cv=5),
'(d) balanced + isotonic': CalibratedClassifierCV(baza(True), method='isotonic', cv=5),
}
wiersze = []
for nazwa, est in warianty.items():
p = cross_val_predict(est, X, y, cv=cv, method='predict_proba')[:, 1]
wiersze.append(dict(wariant=nazwa,
ROC_AUC=roc_auc_score(y, p),
PR_AUC=average_precision_score(y, p),
ECE=ece(p, y)))
kalib = pd.DataFrame(wiersze).set_index('wariant').round(4)
print(kalib)
kalib
ROC_AUC PR_AUC ECE wariant (a) balanced (baseline) 0.9213 0.7034 0.1535 (b) zwykły logit 0.9196 0.7128 0.0168 (c) balanced + Platt 0.9216 0.6955 0.0229 (d) balanced + isotonic 0.9198 0.6742 0.0201
| ROC_AUC | PR_AUC | ECE | |
|---|---|---|---|
| wariant | |||
| (a) balanced (baseline) | 0.9213 | 0.7034 | 0.1535 |
| (b) zwykły logit | 0.9196 | 0.7128 | 0.0168 |
| (c) balanced + Platt | 0.9216 | 0.6955 | 0.0229 |
| (d) balanced + isotonic | 0.9198 | 0.6742 | 0.0201 |
Is calibration needed?¶
| variant | ECE |
ROC-AUC |
PR-AUC |
|---|---|---|---|
balanced (raw) |
0.1535 |
0.921 |
0.703 |
plain logit (no balanced) |
0.0168 |
0.920 |
0.713 |
balanced + Platt (sigmoid) |
0.0229 |
0.922 |
0.696 |
balanced + isotonic |
0.0201 |
0.920 |
0.674 |
The suspicion is confirmed, and calibration works. The raw model with class_weight='balanced' has ECE = 0.1535, so it heavily overestimates podium chances: at the 0.5 threshold it predicts 279 podiums against 138 real ones (class-weight table above: TP + FP = 123 + 156), and its declared probabilities sit clearly above the actual share of podiums (~15% of the data). We'll examine this model's calibration curve bin by bin in notebook 05. Both Platt and isotonic regression pull ECE down to around 0.02.
Calibration doesn't move the ordering measured by ROC-AUC, but it isn't free. ROC-AUC stays at ~0.92 in all four variants, while PR-AUC drops slightly: from 0.703 to 0.696 (Platt) and 0.674 (isotonic). The Platt drop is within noise - the sigmoid is strictly monotone, so by itself it cannot change the order; the difference comes from averaging five models inside CalibratedClassifierCV. The real cost is paid by isotonic regression: it fits a step function, and each calibration set here has barely ~20 podiums, so the steps are wide - neighboring probabilities get merged and some ranking resolution within the rare class is lost.
Plain logistic regression (no balanced) is self-calibrated (ECE=0.017). This is a known property of a logit with an intercept: the mean fitted probability equals the positive-class share of the training data. If reliable percentages are key, in this case it seems better to drop balanced.
Verdict for our case¶
The goal is to rank podium candidates - who has the best chance - and to compare drivers with each other. Only order matters for this, and it is measured by ROC-AUC, which calibration leaves unchanged.
Therefore for our application calibration is optional - we don't read the output literally as this driver has a 44% chance, but rather as this one ranks higher than that one.
If we wanted to give a fan the true probability of a podium or plug those numbers into further calculations (e.g., the expected value of a bet) - then calibration would make sense and be necessary, because the raw balanced model would clearly overstate the chances (as we'll show in notebook 05 - roughly twofold on average).
Summary¶
- Accuracy alone is misleading with imbalanced classes - the "always no" model achieves
~85%without predicting anything. - Data leakage (the
Pointsfeature) gives a seemingly perfect result that is useless in practice - that's why we stick to features known before the race. - With this amount of data logistic regression beats tree models on quality, speed, and interpretability at once.
- The strongest signal turns out to be the starting position (
GridPosition) - which matches a Formula 1 fan's intuition;quali_gap_scarries almost the same information as the starting position, andTeamadds little (without this feature the model loses barely~0.01ROC-AUC). - More seasons don't help:
ROC-AUCsits at~0.92from one to seven seasons, and mixing technical eras does no harm, because the rule "whoever starts up front has a chance at the podium" is timeless. - Class weights don't change the model's ability to order drivers (
ROC-AUC0.920-0.922across all variants), only the trade-off betweenprecisionandrecallat the0.5threshold; the bestF1 = 0.642came from the moderate variant{0:1, 1:3}. - The
balancedmodel overestimates podium chances (ECE = 0.1535); Platt or isotonic calibration pullsECEdown to~0.02withROC-AUCunchanged, at the cost of a slight drop inPR-AUC(isotonic in particular). For our purpose - ranking podium candidates - calibration is optional.