DATA SCIENCE

How to Handle Missing Values in pandas

Understand missing data, inspect null values, and compare common pandas strategies with practical Python examples.

Student learning resource7 min read

Missing values are common in real-world datasets. Before deciding how to handle them, investigate where they occur, why they are missing and how each choice might affect your analysis.

Inspect missing values

In pandas, isna() identifies missing values. Summing the resulting Boolean values gives the number of missing entries in each column.

import pandas as pd

df = pd.DataFrame({
    "score": [82, None, 91, 76],
    "hours_studied": [3, 5, None, 2],
})

print(df.isna().sum())

Remove rows only when appropriate

dropna() can remove rows or columns with missing values. This may be reasonable for a small number of unusable records, but it can also discard useful information or introduce bias.

# Create a separate cleaned version
complete_rows = df.dropna()

Consider imputation

fillna() can replace missing values with a chosen value. The mean or median may be reasonable in some numerical datasets, but the method should match the variable and the analysis.

median_score = df["score"].median()
df["score_filled"] = df["score"].fillna(median_score)

Avoid data leakage

For predictive modelling, fit imputation rules using the training data only, then apply those fitted rules to validation and test data. Computing statistics across the full dataset can leak information into model evaluation.

Document your choice

Record how much data was missing, which strategy you used and why. Compare alternatives when the choice could materially affect your findings.

KEEP LEARNING

Understanding the concept is the first step.

Explore related resources or learn about subject-focused tutoring designed to help you understand the material.