I have a dataframe and I did some features selection (following a guide) to drop some columns:
What I did:
X = df.drop('goal', axis=1).select_dtypes(exclude=['object'])
y = df['goal']
Then I selected columns using mutual_info_gain:
from sklearn.feature_selection import mutual_info_regression, mutual_info_classif
info_gain = mutual_info_classif(X, y)
And finally:
columns_to_keep = []
for score, f_name in sorted(zip(info_gain, X.columns), reverse=True)[:50]:
print(f_name, score)
columns_to_keep.append(f_name)
df_info_gain = X[columns_to_keep]
So now df_info_gain has all 50 features I selected but not the goal column y
What I need:
What's the correct way to bring back the goal column to this new df_info_gain dataframe?