ValueError: feature_names mismatch: when using XGB within RandomizedSearchCV but not in XGB alone

Viewed 145

I am using RandomizedSearchCV on an xgboost with early stopping within an imblearn pipeline. If I run the code using RandomizedSearchCV, I receive the following error:

File "/anaconda3/lib/python3.6/site-packages/xgboost/core.py", line 2131, in _validate_features
    data.feature_names))
ValueError: feature_names mismatch: ['f0', 'f1', 'f2'...

Here is the code

classifier = XGBClassifier(
                    n_estimators=200,
                    bootstrap=True,
                    objective = 'binary:logistic',
                    random_state=0,
                    verbosity=1
                    )

pipeline = Pipeline([
        
            ('union', FeatureUnion(  #Feature union vertically merges text, numeric, and categorical data for model ingestion
                transformer_list=[
                    ('categorical',Pipeline([ 
                        ('selector', ItemSelector(key=['region'])),
                        ('onehotencoder',encoder)
                    ])),
                    ('bow', Pipeline([ 
                        ('selector', ItemSelector(key='message')),
                        ('clean',CleanText()),
                        ('tfidf', tfidf_vectorizer)
                    ])),
                    ('text_stats', Pipeline([ 
                        ('selector', ItemSelector(key='message')),
                        ('stats', TextStats()),  
                        ('vect', DictVectorizer()) 
                    ]))
                ]
            )),
        
            ('model', classifier) # Modeling step
        ],verbose=2)
        
        pipeline_temp = Pipeline(pipeline.steps[:-1])
        pipeline_temp.fit(X_train,y_train)
        eval_set = [(pipeline_temp.transform(X_test),y_test)]
        
        param_dist = {  
            "model__max_depth": st.randint(32, 96),
            "model__learning_rate": st.uniform(0.05,0.4),
            "model__subsample": st.beta(10,1),
            "model__colsample_bytree": st.uniform(0.4,0.8),
            "model__gamma": st.uniform(0,5),
            "model__reg_alpha": st.expon(0, 50),
            "model__min_child_weight": [1,3,5,9]
        }
        searcher = RandomizedSearchCV(estimator=pipeline,
                                        param_distributions = param_dist,
                                        n_iter=10,
                                        cv=2,
                                        refit=False,
                                        verbose=3,
                                        n_jobs=-1,
                                        error_score='raise')
        
        searcher.fit(X_train,y_train,
                     model__early_stopping_rounds=20,
                     model__eval_metric="logloss",
                     model__eval_set=eval_set,
                     model__verbose=1)
        }

If I alter the code to use pipeline.fit without the RandomizedSearchCV, the code executes without error. All github issues and Stackoverflow posts seem to point to an xgboost issue with DMatrix, but since it runs outside of RandomizedSearchCV, this seems odd. The reason I chose to use the parameter error_score='raise' is because I noticed the cv_results_ showed NaNs for all test_score values.

Versions:

I tested this out in two dev environments.

Python 3.6.10 (same issue in 3.7.7)

sklearn 0.23.1

xgboost 1.1.1 (same issue in 0.90)

imblearn 0.6.2 (same issue in 0.7.0)

0 Answers
Related