Skip to content

[FAQ] Why should missing-value mean be calculated only from the training data? #435

Description

@Nilo0far

Course

machine-learning-zoomcamp

Question

Why should I calculate the mean for filling missing values using only the training dataset instead of the full dataset?

Answer

The mean should be calculated only from the training data because the validation and test datasets should not be used when preparing the training data.

If we calculate the mean using the full dataset, information from the validation or test data can leak into the training process. This is called data leakage.

For example:

mean = X_train['horsepower'].mean()

X_train = X_train.fillna(mean)
X_val = X_val.fillna(mean)
X_test = X_test.fillna(mean)

We calculate the mean from X_train once, and then use the same value to fill missing values in the training, validation, and test datasets.

Checklist

  • I have searched existing FAQs and this question is not already answered
  • The answer provides accurate, helpful information
  • I have included any relevant code examples or links
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions