Selecting best k features using Chi-Square test

Viewed 94

I have been trying to implement Chi-Square feature selection, wherein I select the best k features or the features that are highly dependent to the Label.

So far I am doing this:

from scipy.stats import chi2_contingency

for col in all_cols:
    contingency_table = pd.crosstab(data[col] , y)
    stat, _, _ , _ = chi2_contingency(contingency_table.values)

Then I am selecting the top features as the ones having higher stat values. Since sklearn already provides this feature using SelectKBest(chi2,...). So, is my implementation correct or in sync with the pre-built approach?

0 Answers
Related