Skip to main content

Command Palette

Search for a command to run...

Why Do Both Data Preprocessing and Feature Engineering Handle Missing Values?

Updated
•4 min read•View as Markdown
Why Do Both Data Preprocessing and Feature Engineering Handle Missing Values?
N
I am an IT student passionate about software development, Artificial Intelligence, Machine Learning, and creating meaningful technology solutions. I enjoy exploring new technical concepts, transforming complex ideas into simple explanations, and documenting my learning journey through technical writing and projects. With experience in web development, UI/UX design, and programming, I am continuously improving my skills while building projects and exploring innovative ideas like educational technology platforms. I believe curiosity drives growth, and through my blogs, I aim to make technology easier to understand and inspire others to learn.

While studying the Machine Learning Pipeline, I came across two stages that confused me:

  • Data Preprocessing

  • Feature Engineering

At first, everything seemed clear.

Data Preprocessing is responsible for cleaning the data by handling missing values, removing duplicates, fixing formatting issues, and dealing with outliers.

But then, when I started reading about Feature Engineering, I noticed something unexpected.

It also mentioned cleaning the data and handling missing values.

That made me stop and think:

If Data Preprocessing already cleans the data, why does Feature Engineering clean it again? Aren't they doing the same thing?

If you've had the same question, you're not alone. I had exactly the same confusion.

After researching, I finally understood the difference.


What Data Preprocessing Really Does

Data Preprocessing focuses on making the dataset clean, consistent, and usable.

Its goal is not to improve the Machine Learning model. Instead, it prepares the raw data so that it can be processed correctly.

Typical tasks include:

  • Filling missing values using simple methods (mean, median, or mode)

  • Removing duplicate records

  • Fixing inconsistent formats

  • Handling incorrect or impossible values

  • Detecting and treating outliers

Think of Data Preprocessing as preparing ingredients before cooking.

Before you cook, you wash the vegetables, remove spoiled parts, and organize everything.

The ingredients are now clean and ready.


Then Why Does Feature Engineering Also Handle Missing Values?

This is where my confusion came from.

The key thing I learned is that Feature Engineering is not trying to clean messy data.

Instead, it modifies data to help the Machine Learning model learn better.

Sometimes that includes handling missing values again—but for a different reason.


The Difference Is the Purpose

Imagine we have a dataset with a missing Age value.

During Data Preprocessing

We simply replace the missing age with the average age of everyone.

Example:

Average age = 25

Missing age → 25

The goal is simply to remove the missing value so the dataset is complete.


During Feature Engineering

Instead of using the overall average, we use more meaningful information.

For example:

  • If the person's job is Doctor, estimate the age as 35.

  • If the person's job is Student, estimate the age as 20.

Now the missing value is filled using domain knowledge, which may help the model make better predictions.

Notice that both stages filled the missing value.

But they did it for completely different reasons.


Another Example

Suppose a dataset contains:

  • Height

  • Weight

Data Preprocessing

It only checks whether these values are valid and complete.

Feature Engineering

It creates a new feature:

BMI = Weight / Height²

The dataset is no longer just clean—it has become more informative for the model.


The Simple Difference

Data Preprocessing Feature Engineering
Makes data clean and usable Makes data more useful for the model
Focuses on data quality Focuses on model performance
Fixes technical problems Creates or improves features
Uses general cleaning techniques Uses domain knowledge and feature design

The Sentence That Solved My Confusion

When I finally understood this, everything became much clearer.

Data Preprocessing makes the data error-free. Feature Engineering makes the data more useful for the Machine Learning model.

Both stages may handle missing values, but the purpose is different.

  • Data Preprocessing removes technical problems so the dataset can be used.

  • Feature Engineering improves the dataset so the model can learn more effectively.

That small difference was the answer I was looking for.


Final Thoughts

This question came from my own AI and Machine Learning lecture while studying the Machine Learning Pipeline.

I realized that sometimes two stages seem to perform the same task, but understanding why they perform that task makes all the difference.

This is exactly why I started this blog—to document the questions that make me stop, think, and learn something new. If this explanation helped clear up your confusion too, then this learning journey is already worth sharing.