How to handle categorical variables in MissForest() during missing value imputation?

Viewed 17

I am working on a regression problem using Bengaluru house price prediction dataset. I was trying to impute missing values in bath and balcony using MissForest(). Since documentation says that MissForest() can handle categorical variables using 'cat_vars' parameter, I tried to use 'area_type' and 'locality' features in the imputer fit_transform method by passing their index, as shown below:

df_temp.info()

        Data columns (total 13 columns):
     #   Column        Non-Null Count  Dtype  
    ---  ------        --------------  -----  
     0   area_type     10296 non-null  object 
     1   location      10296 non-null  object 
     2   bath          10245 non-null  float64
     3   balcony       9810 non-null   float64
     4   rooms         10296 non-null  int64  
     5   tot_sqft_1    10296 non-null  float64


    imputer = MissForest()
    imputer.fit_transform(df_temp, cat_vars=[0,1])

But I am getting the below error: 'Cannot convert str to float: 'Super built up Area''

Could you please let me know why this could be? Do we need to encode the categorical variables using one hot encoding?

0 Answers
Related