Every panel on this page uses one data set of 500 points, drawn once from \(y = f(x) + \varepsilon\) with \(f(x) = 0.3 + 2x - 0.8x^2 - 0.4x^3\), \(x \sim \mathcal N(0, 1)\) and \(\varepsilon \sim \mathcal N(0, 1)\). The shuffle seed slider only changes how those 500 points are dealt out into the training and test sets. At seed 0 they are not shuffled at all.
The validation set approach
Shuffle the points, put half in a training set and half in a test set, fit on the first half and measure on the second.
A degree 8 polynomial fitted to whichever 250 points the shuffle put in the training set, with the other 250 as the test set.
The fitted curve and the test error depend strongly on which points landed in the training set. Over the 21 seeds the test RMSE runs from 1.06 to 2.38 — the worst split reports a model more than twice as bad as the best one, from the same 500 points and the same degree. The training error barely moves, 0.95 to 1.08, and sits on the irreducible error of 1 throughout.
Cross-validation
\(K\)-fold splits the data into \(K\) parts, fits on \(K-1\) of them and scores the one left out, \(K\) times, so every point is scored exactly once by a model that did not see it. Taking \(K = n\) leaves one point out at a time, which is leave-one-out cross-validation.
The same 500 points and the same degree 8 fit, scored by cross-validation. Each point is one of reshufflings 1 to 10, jittered sideways so they do not overlap; LOOCV has nothing to reshuffle and gives one number.
2-fold is two validation-set splits averaged, and it inherits most of their spread: 1.12 to 1.84 over the ten reshufflings. 4-fold is hardly better, with two reshufflings above 1.5. 10-fold stays between 1.10 and 1.21, close to the LOOCV value of 1.10.
Running a resampling strategy
The validation set approach is one call to split and one to score.