I'm building a model, pretty much similiar to the well known House Price Prediction. I got to the point that I need to encode my nominal categorical variables by using scikit-learns OneHotEncoder. The so called "Dummy Variable Trap" is clear to me so I need to drop one of my OneHot encoded columns to avoid multicollinearity.
What's bothering me, is the way to handle unseen categories. In my understanding the unseen categories will be treated the same way as the "base category" (the category I dropped).
To make it clear have a look at this example:
This is my training data i use to fit my OneHotEncoder.
X_train:
| index | city |
|---|---|
| 0 | Munich |
| 1 | Berlin |
| 2 | Hamburg |
| 3 | Berlin |
OneHotEncoding:
oh = OneHotEncoder(handle_unknown = 'ignore', drop = 'first')
oh.fit_transform(X_train)
Because of drop = 'first' the first column ('city_Munich') will be dropped.
| index | city_Berlin | city_Hamburg |
|---|---|---|
| 0 | 0 | 0 |
| 1 | 1 | 0 |
| 2 | 0 | 1 |
| 3 | 1 | 0 |
Now it's about to encode unseen data:
X_test:
| index | city |
|---|---|
| 10 | Munich |
| 11 | Berlin |
| 12 | Hamburg |
| 13 | Cologne |
oh.transform(X_test)
| index | city_Berlin | city_Hamburg |
|---|---|---|
| 10 | 0 | 0 |
| 11 | 1 | 0 |
| 12 | 0 | 1 |
| 13 | 0 | 0 |
I guess you may see my problem. Row10 (Munich) is treated the same way as row13 (Cologne).
Either I run into the "dummy variable trap" when not dropping one column or I gonna treat unseen data as the "base category" which is in fact wrong.
Whats a proper way to deal with that? Is there any option in the OneHotEncoder class to add a new column for unseen categories like "city_unseen"?