I have been trying to implement Chi-Square feature selection, wherein I select the best k features or the features that are highly dependent to the Label.
So far I am doing this:
from scipy.stats import chi2_contingency
for col in all_cols:
contingency_table = pd.crosstab(data[col] , y)
stat, _, _ , _ = chi2_contingency(contingency_table.values)
Then I am selecting the top features as the ones having higher stat values.
Since sklearn already provides this feature using SelectKBest(chi2,...).
So, is my implementation correct or in sync with the pre-built approach?