Starting with an arbitrary dataframe, I would like to return a dataframe with only those columns which have more than one distinct value.
I have:
X = df.nunique()
like:
Id 5
MSSubClass 3
MSZoning 1
LotFrontage 5
LotArea 5
Street 1
Alley 0
LotShape 2
Then I converted this from a series to a dataframe:
X = X.to_frame(name = 'dcount')
Then I used a where clause to only return values > 1:
X.where(X[['dcount']]>1)
which looks like:
dcount
Id 5.0
MSSubClass 3.0
MSZoning NaN
LotFrontage 5.0
LotArea 5.0
Street NaN
Alley NaN
LotShape 2.0
...
But I now want only those column_names (in the index of X) which don't have dcount = 'NaN', so that I can ultimately go back to my original dataframe df and define it as:
df=df[[list_of_columns]]
How should this be done? I've tried a dozen ways and it's a PitA. I suspect there's a way to do it in 1 or 2 lines of code.