When I run the following code:
import numpy as np
import pandas as pd
df = pd.DataFrame({"S": ["xx", np.nan, np.nan], "V": [1.1,2.1,3.1]})
a = df.groupby(["S"], as_index=False, dropna=False).sum()
print(a.dtypes)
I get the following result, as expected:
S object
V float64
However, if I select a subset of rows that happen only to contain nan values:
df2 = df.iloc[1:,:]
a2 = df2.groupby(["S"], as_index=False, dropna=False).sum()
print(a2.dtypes)
the S column in the resulting DataFrame is converted to a float which breaks subsequent code.
S float64
V float64
dtype: object
Is there any way of preventing this inferring of types and always keep the original type of S?