Convert pandas column (containing floats and NaN values) from float64 to nullable int8

Viewed 514

I have a large dataframe that looks somewhat like this:

    a   b   c
0   2.2 6.0 0.0
1   3.3 7.0 NaN
2   4.4 NaN 3.0
3   5.5 9.0 NaN

Columns b and c contain float values that are either postive, natural numbers or NaN. However, they are stored as float64, which is a problem, since (without going into further detail) this dataframe is the input of a pipeline that requires these to be integers, so and I want to store them as such. The output should look like this:

    a   b   c
0   2.2 6   0
1   3.3 7   NaN
2   4.4 NaN 3
3   5.5 9   NaN

I read in the pandas documentation that nullable integers are only supported in the pandas datatype "Int8" (note: this is different from np.int8), so naturally, I attempted this:

df = df.astype({'b':pd.Int8Dtype(), 'c':pd.Int8Dtype()})

This works when I run it in my Jupyter notebook, but when I integrate it within a larger function, I get this error:

TypeError: cannot safely cast non-equivalent float64 to int8

I understand why I get the error, since x == int(x), will be False for NaN values, so the program thinks this conversion is unsafe, even though all values are either NaN or natural number. So next, I tried:

'df = df.astype({'b':pd.Int8Dtype(), 'c':pd.Int8Dtype()}, errors='ignore')

I figured that this would get rid of the 'unsafe conversion' problem, since I am 100% sure all float64 values are natural numbers. However, when I use this line, all of my numbers are still stored as floats! Infuriating!

Does anyone have a workaround for this?

1 Answers

I ran into exactly the same issue which led me to this page. I do not have a genuinely good solution for this issue and am seeking for one myself... but I did find a workaround. Before going into that I would like to answer to the comment posted on the original question that: allowing to have NA or even None values assigned to series of such 'simple' types as int8 is the whole point of trying to make these dtype conversions. It is possible to perform the typical operations such as isna() (and so on) on series of these dtypes (see pd.IntXDtype() where 'X' stands for the number of bits). The advantage I explore by using these dtypes is on memory footprint, eg:

In[56]: test_df = pd.Series(np.zeros(1_000_000), dtype=np.float64)

In[57]: test_df.memory_usage()
Out[57]: 8000128

In[58]: test_df = pd.Series(np.zeros(1_000_000), dtype=pd.Int8Dtype())

In[59]: test_df.memory_usage()
Out[59]: 2000128

In[60]: test_df.iloc[:500_000] = None

In[61]: test_df.memory_usage()
Out[61]: 2000128

In[62]: test_df.isna().sum()
Out[62]: 500000

So you get the best of both worlds.

Now the workarround:

In[33]: my_df
Out[33]: 
     a    s      d
0    0 -500 -1.000
1    1 -499 -0.998
2    2 -498 -0.996
3    3 -497 -0.994
4    4 -496 -0.992

In[34]: my_df.dtypes
Out[34]: 
a      int64
s      int64
d    float64
dtype: object

In[35]: df_converted_to_int_first = my_df.astype(
   ...:     dtype={
   ...:         'a': np.int8,
   ...:         's': np.int16,
   ...:         'd': np.float16,
   ...:     },
   ...: )

In[36]: df_converted_to_int_first
Out[36]: 
     a    s         d
0    0 -500 -1.000000
1    1 -499 -0.998047
2    2 -498 -0.996094
3    3 -497 -0.994141
4    4 -496 -0.992188

In[37]: df_converted_to_int_first.dtypes
Out[37]: 
a       int8
s      int16
d    float16
dtype: object

In[38]: df_converted_to_special_int_after = df_converted_to_int_first.astype(
   ...:     dtype={
   ...:         'a': pd.Int8Dtype(),
   ...:         's': pd.Int16Dtype(),
   ...:     }
   ...: )

In[39]: df_converted_to_special_int_after.dtypes
Out[39]: 
a       Int8
s      Int16
d    float16
dtype: object

In[40]: df_converted_to_special_int_after.a.iloc[3] = None

In[41]: df_converted_to_special_int_after
Out[41]: 
       a     s         d
0      0  -500 -1.000000
1      1  -499 -0.998047
2      2  -498 -0.996094
3   <NA>  -497 -0.994141
4      4  -496 -0.992188

This is still not an acceptable solution in my opinion... but as mentioned above ir constitutes a workaround which is asked in the original question.

EDIT Some test that was missing, from np.float64 to pd.Int8Dtype():

In[67]: my_df.astype(
   ...:     dtype={
   ...:         'a': np.int8,
   ...:         's': np.int16,
   ...:         'd': np.int16,
   ...:     },
   ...: ).astype(    
   ...:     dtype={
   ...:         'a': np.int8,
   ...:         's': np.int16,
   ...:         'd': pd.Int8Dtype(),
   ...:     },
   ...: ).dtypes

Out[67]: 
a     int8
s    int16
d     Int8
dtype: object
Related