With \(C\) classes the model needs \(C\) probabilities that are positive and sum to one. We take one linear function per class, \(f_c(x) = \theta_{c0} + \theta_{c1}x_1 + \cdots + \theta_{cp}x_p\), and pass the vector of all \(C\) of them through the softmax, \[
P(Y = c\,|\,x) = s(f(x))_c = \frac{e^{f_c(x)}}{\sum_{k=1}^{C} e^{f_k(x)}} .
\] The parameters form a \(C \times (p + 1)\) matrix \(\theta\). This model is called multinomial logistic regression, or linear classification. Its loss is again the negative log-likelihood, \(-\frac1n\sum_i \log P(Y = y_i\,|\,x_i)\), the cross-entropy.
Four classes and two inputs, so three parameters per class. Every point of the plane is coloured by the class with the largest probability, \(\arg\max_c s(f(x))_c\). The arrows are the weights \((\theta_{c1}, \theta_{c2})\) of each class, drawn twice as long. The sliders \(a_0\), \(a_1\), \(a_2\) add the same number \(a_0\) to every intercept and the same vector \((a_1, a_2)\) to every arrow. The black point is \(x = (1, 0)\).
Every boundary between two regions is a straight line, because every \(f_c\) is linear. The boundary between classes \(c\) and \(c'\) is where \(f_c = f_{c'}\), a line perpendicular to the difference of their arrows. Far from the origin the intercepts no longer matter, and in every direction the class whose arrow reaches furthest in that direction wins. Moving \(a_0\), \(a_1\) or \(a_2\) changes the arrows but neither the regions nor the probabilities at the black point: only the differences between the classes matter.
Fitting it by gradient descent
torch has cross_entropy, which takes the values \(f_c(x_i)\) and the integer labels, applies the softmax and returns the mean negative log-likelihood.
Code
import numpy as npimport torchimport matplotlib.pyplot as pltrng = np.random.default_rng(5)C, p, n =3, 2, 600centers = np.array([[-2.0, 0.0], [1.5, 1.5], [1.0, -2.0]])y = rng.integers(0, C, n)X = centers[y] + rng.normal(0, 1.1, (n, p)) # three clouds of pointsXt = torch.tensor(X, dtype=torch.float64)yt = torch.tensor(y, dtype=torch.long)
theta = torch.zeros((C, p +1), dtype=torch.float64, requires_grad=True)def f(theta, X):return theta[:, 0] + X @ theta[:, 1:].Topt = torch.optim.Adam([theta], lr=0.05)for step inrange(400): opt.zero_grad() loss = torch.nn.functional.cross_entropy(f(theta, Xt), yt) loss.backward() opt.step()print("cross-entropy after 400 steps:", round(float(loss), 4))
1
The \(C\) linear functions for every point at once, an \(n \times C\) matrix.
Figure 29.1: The data and the decision regions of the fitted model.
Linear classification of MNIST
The MNIST images have \(28 \times 28 = 784\) pixels and ten classes, so a linear classifier has ten linear functions of 784 inputs, \(10 \times 785 = 7850\) parameters. We fit it on the 60000 training images with Adam on batches of 128, and evaluate it on the 10000 test images.