One Hot Encoding: Avoiding dummy variable trap and process unseen data with scikit learn

Viewed 602

I'm building a model, pretty much similiar to the well known House Price Prediction. I got to the point that I need to encode my nominal categorical variables by using scikit-learns OneHotEncoder. The so called "Dummy Variable Trap" is clear to me so I need to drop one of my OneHot encoded columns to avoid multicollinearity.

What's bothering me, is the way to handle unseen categories. In my understanding the unseen categories will be treated the same way as the "base category" (the category I dropped).

To make it clear have a look at this example:

This is my training data i use to fit my OneHotEncoder.

X_train:

index city
0 Munich
1 Berlin
2 Hamburg
3 Berlin

OneHotEncoding:

oh = OneHotEncoder(handle_unknown = 'ignore', drop = 'first')

oh.fit_transform(X_train)

Because of drop = 'first' the first column ('city_Munich') will be dropped.

index city_Berlin city_Hamburg
0 0 0
1 1 0
2 0 1
3 1 0

Now it's about to encode unseen data:

X_test:

index city
10 Munich
11 Berlin
12 Hamburg
13 Cologne

oh.transform(X_test)

index city_Berlin city_Hamburg
10 0 0
11 1 0
12 0 1
13 0 0

I guess you may see my problem. Row10 (Munich) is treated the same way as row13 (Cologne).

Either I run into the "dummy variable trap" when not dropping one column or I gonna treat unseen data as the "base category" which is in fact wrong.

Whats a proper way to deal with that? Is there any option in the OneHotEncoder class to add a new column for unseen categories like "city_unseen"?

0 Answers
Related