Course
machine-learning-zoomcamp
Question
Why should I calculate the mean for filling missing values using only the training dataset instead of the full dataset?
Answer
The mean should be calculated only from the training data because the validation and test datasets should not be used when preparing the training data.
If we calculate the mean using the full dataset, information from the validation or test data can leak into the training process. This is called data leakage.
For example:
mean = X_train['horsepower'].mean()
X_train = X_train.fillna(mean)
X_val = X_val.fillna(mean)
X_test = X_test.fillna(mean)
We calculate the mean from X_train once, and then use the same value to fill missing values in the training, validation, and test datasets.
Checklist
Course
machine-learning-zoomcamp
Question
Why should I calculate the mean for filling missing values using only the training dataset instead of the full dataset?
Answer
The mean should be calculated only from the training data because the validation and test datasets should not be used when preparing the training data.
If we calculate the mean using the full dataset, information from the validation or test data can leak into the training process. This is called data leakage.
For example:
mean = X_train['horsepower'].mean()
X_train = X_train.fillna(mean)
X_val = X_val.fillna(mean)
X_test = X_test.fillna(mean)
We calculate the mean from X_train once, and then use the same value to fill missing values in the training, validation, and test datasets.
Checklist