Removing columns with sklearn's OneHotEncoder

Viewed 6891
from sklearn.preprocessing import LabelEncoder as LE, OneHotEncoder as OHE
import numpy as np

a = np.array([[0,1,100],[1,2,200],[2,3,400]])


oh = OHE(categorical_features=[0,1])
a = oh.fit_transform(a).toarray()

Let's assume first and second column are categorical data. This code does one hot encoding, but for the regression problem, I would like to remove first column after encoding categorical data. In this example, there are two and I could do it manually. But what if you have many categorical features, how would you solve this problem?

4 Answers

For that I use a Wrapper like that which is also usable in Pipelines:

class DummyEncoder(BaseEstimator, TransformerMixin):

    def __init__(self, n_values='auto'):
        self.n_values = n_values

    def transform(self, X):
        ohe = OneHotEncoder(sparse=False, n_values=self.n_values)
        return ohe.fit_transform(X)[:,:-1]

    def fit(self, X, y=None, **fit_params):
        return self

This is one of the limitation of the One-Hot encoder in Sklearn when dealing with building models. The best way to do that if you have multiple categorical variables is first to use LabelEncoder to identify the unique labels for each categorical variables and then utilize them to generate the indexes to delete. As an example, if you have the data in a numpy array as X, with categorical variables in FIRST_IDX, SECOND_IDX, THIRD_IDX columns, first encode them using LabelEncoder.

labelencoder_X_1 = LabelEncoder()
X[:, FIRST_IDX] = labelencoder_X_1.fit_transform(X[:, FIRST_IDX])

labelencoder_X_2 = LabelEncoder()
X[:, SECOND_IDX] = labelencoder_X_2.fit_transform(X[:, SECOND_IDX])

labelencoder_X_3 = LabelEncoder()
X[:, THIRD_IDX] = labelencoder_X_3.fit_transform(X[:, THIRD_IDX])

Then apply One-Hot Encoder, which will create the representation at the beginning of the array for all categorical variables, one after the other.

onehotencoder = OneHotEncoder(categorical_features=[FIRST_IDX, SECOND_IDX, THIRD_IDX])

X = onehotencoder.fit_transform(X).toarray()

Finally, eliminate the first entry for each categorical variable by leveraging the size of unique values for each categorical variables and using the cumulative sum (here the cumulative sum gives you the first entry index of the categorical variables) in numpy.

index_to_delete = np.cumsum([0,
               len(labelencoder_X_1.classes_),
               len(labelencoder_X_2.classes_),
               len(labelencoder_X_3.classes_)
               ])
index_to_keep = [i for i in range(X.shape[1]) if i not in index_to_delete]

X = X[:, index_to_keep]

Now X contains the data ready to be used in any modeling task.

Related