weight was read as text, because one entry has a unit in it. -999 is not a weight but a code for “not measured”. sex has four spellings of two categories. And the row with id 3 is there twice.
errors="coerce" turns anything that is still not a number into a missing value instead of stopping with an error.
3
The code for “not measured” becomes a real missing value, which the next section deals with.
id
weight
sex
0
1
71.0
male
1
2
68.5
female
2
3
NaN
female
4
4
80.0
male
5
5
75.0
male
A fifth kind of problem is one that no code finds: a predictor that is only known after the response. A diagnosis written at discharge predicts perfectly whether a patient was admitted, and it is useless for deciding who to admit. Check for each column when it becomes available.
Missing data
Two columns, one number and one category, each with a hole in it.
The simplest option is to drop the rows. This is fine if only a few rows are affected. It is not fine if the values are missing for a reason, because then the remaining rows are no longer a random sample.
datam.dropna()
age
gender
1
41.0
female
3
33.0
male
4
27.0
male
5
50.0
female
The other option is to fill the missing values in. Note that the imputer learns the value it fills in, so it belongs inside the fold.
from sklearn.impute import SimpleImputerimp = SimpleImputer(strategy="most_frequent")pd.DataFrame(imp.fit_transform(datam), columns=["age", "gender"])
age
gender
0
12.0
male
1
41.0
female
2
12.0
male
3
33.0
male
4
27.0
male
5
50.0
female
Whether a value is missing can itself carry information: a test that was not ordered, a question that was skipped. add_indicator=True keeps it as an extra column of zeros and ones.
Columns a and e are constant and are gone. Column d is 2 * c: two predictors with correlation 1 carry the same information twice, which makes the least squares solution non-unique. Without a penalty, drop one of them. With a penalty, as in ridge regression, the solution is unique again and both can stay.
Standardization
Standardization shifts the data so that the mean is 0 and scales it so that the standard deviation is 1. Some methods need it, for example regularization and \(k\) nearest neighbours. Others do not care.
The scaler learns the mean and the standard deviation from the data. If we fit it on all the data before splitting, the validation fold has already influenced the training set. Use a Pipeline.