Skip to main content

Resampling

Examples

Every panel on this page uses one data set of 500 points, drawn once from \(y = f(x) + \varepsilon\) with \(f(x) = 0.3 + 2x - 0.8x^2 - 0.4x^3\), \(x \sim \mathcal N(0, 1)\) and \(\varepsilon \sim \mathcal N(0, 1)\). The shuffle seed slider only changes how those 500 points are dealt out into the training and test sets. At seed 0 they are not shuffled at all.

The validation set approach

Shuffle the points, put half in a training set and half in a test set, fit on the first half and measure on the second.

A degree 8 polynomial fitted to whichever 250 points the shuffle put in the training set, with the other 250 as the test set.

The fitted curve and the test error depend strongly on which points landed in the training set. Over the 21 seeds the test RMSE runs from 1.06 to 2.38 — the worst split reports a model more than twice as bad as the best one, from the same 500 points and the same degree. The training error barely moves, 0.95 to 1.08, and sits on the irreducible error of 1 throughout.

Cross-validation

\(K\)-fold splits the data into \(K\) parts, fits on \(K-1\) of them and scores the one left out, \(K\) times, so every point is scored exactly once by a model that did not see it. Taking \(K = n\) leaves one point out at a time, which is leave-one-out cross-validation.

The same 500 points and the same degree 8 fit, scored by cross-validation. Each point is one of reshufflings 1 to 10, jittered sideways so they do not overlap; LOOCV has nothing to reshuffle and gives one number.

2-fold is two validation-set splits averaged, and it inherits most of their spread: 1.12 to 1.84 over the ten reshufflings. 4-fold is hardly better, with two reshufflings above 1.5. 10-fold stays between 1.10 and 1.21, close to the LOOCV value of 1.10.

Running a resampling strategy

The validation set approach is one call to split and one to score.

import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.metrics import root_mean_squared_error
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import PolynomialFeatures

def f(x):
    return 0.3 + 2 * x - 0.8 * x ** 2 - 0.4 * x ** 3

rng = np.random.default_rng(1)
x = rng.standard_normal(500)
y = f(x) + rng.standard_normal(500)
X = x.reshape(-1, 1)

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.5,
                                                    shuffle=True, random_state=0)

model = make_pipeline(PolynomialFeatures(degree=8), LinearRegression())
model.fit(X_train, y_train)
print("test RMSE:", round(root_mean_squared_error(y_test, model.predict(X_test)), 3))
1
The same 500 points as in the panels above. train_test_split shuffles them its own way, so this split is none of the panel’s seeds.
2
sklearn wants a two dimensional X, one row per point.
test RMSE: 1.054

Cross-validation is the same model handed to cross_val_score instead, with the data whole.

from sklearn.model_selection import cross_val_score

scores = -cross_val_score(model, X, y, cv=10,
                          scoring="neg_root_mean_squared_error")
print("per fold:", scores.round(2))
print(f"mean {scores.mean():.3f}   sd {scores.std():.3f}")
1
sklearn maximizes a score, so its error measures are negated. Flipping the sign back is what makes these numbers comparable to the ones above.
per fold: [1.1  1.1  1.13 1.11 1.13 1.27 0.94 1.13 0.96 0.97]
mean 1.085   sd 0.096

The spread across the ten folds is not an error bar on the mean — the folds share training data, so their scores are correlated.