Missing values are common in real-world datasets. Before deciding how to handle them, investigate where they occur, why they are missing and how each choice might affect your analysis.
Inspect missing values
In pandas, isna() identifies missing values. Summing the resulting Boolean values gives the number of missing entries in each column.
import pandas as pd
df = pd.DataFrame({
"score": [82, None, 91, 76],
"hours_studied": [3, 5, None, 2],
})
print(df.isna().sum())Remove rows only when appropriate
dropna() can remove rows or columns with missing values. This may be reasonable for a small number of unusable records, but it can also discard useful information or introduce bias.
# Create a separate cleaned version
complete_rows = df.dropna()Consider imputation
fillna() can replace missing values with a chosen value. The mean or median may be reasonable in some numerical datasets, but the method should match the variable and the analysis.
median_score = df["score"].median()
df["score_filled"] = df["score"].fillna(median_score)Avoid data leakage
For predictive modelling, fit imputation rules using the training data only, then apply those fitted rules to validation and test data. Computing statistics across the full dataset can leak information into model evaluation.
Document your choice
Record how much data was missing, which strategy you used and why. Compare alternatives when the choice could materially affect your findings.
Understanding the concept is the first step.
Explore related resources or learn about subject-focused tutoring designed to help you understand the material.
