Using miceforest for imputation and getting an error on RandomState.choice()

Viewed 398

I'm working with a Kaggle dataset playing with different imputation techniques, and I'm getting an error from the miceforest package that I don't understand.

Full Colab here, but the gist is that the data I'm using are the numeric features from my X_train dataset.

X = df_agg.drop(['TARGET','SK_ID_CURR'], axis = 1)
y = df_agg.TARGET

X_train_raw, X_test_raw, y_train, y_test = train_test_split(
  X, y, test_size=0.10, random_state=42, stratify=y)


X_train_raw, X_dev_raw, y_train, y_dev = train_test_split(
  X_train_raw, y_train,
  test_size=1/9.,
  random_state=42,
  stratify=y_train
)
num_features = X_train_raw.select_dtypes(include=['int64', 'float64']).columns 



kernel = mf.MultipleImputedKernel(
  data=X_train_raw[num_features],
  save_all_iterations=True,
  random_state=1991
)

And here's the error:

ValueError                                Traceback (most recent call last)
<ipython-input-20-b6bd23e78d87> in <module>
      2   data=X_train_raw[num_features],
      3   save_all_iterations=True,
----> 4   random_state=1991
      5 )

~/anaconda3/envs/tensorflow2_latest_p37/lib/python3.7/site-packages/miceforest/MultipleImputedKernel.py in __init__(self, data, datasets, variable_schema, mean_match_candidates, save_all_iterations, save_models, random_state)
     56                 save_all_iterations=save_all_iterations,
     57                 save_models=save_models,
---> 58                 random_state=random_state,
     59             )
     60         )

~/anaconda3/envs/tensorflow2_latest_p37/lib/python3.7/site-packages/miceforest/KernelDataSet.py in __init__(self, data, variable_schema, mean_match_candidates, save_all_iterations, save_models, random_state)
     88             mean_match_candidates=mean_match_candidates,
     89             save_all_iterations=save_all_iterations,
---> 90             random_state=random_state,
     91         )
     92 

~/anaconda3/envs/tensorflow2_latest_p37/lib/python3.7/site-packages/miceforest/ImputedDataSet.py in __init__(self, data, variable_schema, mean_match_candidates, save_all_iterations, random_state)
     87             self.imputation_values[var] = {
     88                 0: self._random_state.choice(
---> 89                     data[var].dropna(), size=self.na_counts[var]
     90                 )
     91             }

mtrand.pyx in numpy.random.mtrand.RandomState.choice()

ValueError: 'a' cannot be empty unless no samples are taken
1 Answers

I'm the maintainer of this package - can you link me the kaggle data you are running this on? As a quick check - can you make sure that there isn't a column in data that has 100% missing values when you pass it to this function? This error occurs when you pass an empty object to numpy.random.choice(), so data[var].dropna() must be empty.

Related