Why can't a pandas dataframe value of NaN be used as a dictionary key?

Viewed 717

I'm trying to use the elements of the values column in the below data frame as keys in a dictionary.

In [1]: import numpy as np
   ...: import pandas as pd
   ...: rng = pd.date_range('2021-06-01', periods=4)
   ...: values = [1, -1, 0, np.nan]
   ...: df = pd.DataFrame(values, index=rng, columns=['values'])

In [2]: df
Out[2]:
            values
2021-06-01     1.0
2021-06-02    -1.0
2021-06-03     0.0
2021-06-04     NaN

The goal is to map the elements of the values column to a set of new values in a separate column to produce the below data frame:

            values new_values
2021-06-01     1.0    A
2021-06-02    -1.0    B
2021-06-03     0.0    C
2021-06-04     NaN    D 

So I created a dictionary with the keys as the elements in the values column.

In [3]: repl = {1: 'A', 0: 'B', -1: 'C',np.nan: 'D'}
In [4]: df['rule'] = df['Val'].apply(lambda x: repl[x])

'NaN', however, is creating a key error (despite it being hashable).

KeyError                                  Traceback (most recent call last)
<ipython-input-143-2e9d3caa7f9c> in <module>
----> 1 df['rule'] = df['Val'].apply(lambda x: repl[x])

~/opt/miniconda3/envs/PyAlgo/lib/python3.7/site-packages/pandas/core/series.py in apply(self, func, convert_dtype, args, **kwds)
   4136             else:
   4137                 values = self.astype(object)._values
-> 4138                 mapped = lib.map_infer(values, f, convert=convert_dtype)
   4139
   4140         if len(mapped) and isinstance(mapped[0], Series):

pandas/_libs/lib.pyx in pandas._libs.lib.map_infer()

<ipython-input-143-2e9d3caa7f9c> in <lambda>(x)
----> 1 df['rule'] = df['Val'].apply(lambda x: repl[x])

KeyError: nan

Obviously, I can create the column manually for this simple example. However, this is just a minimally re-produceable example. In reality, I have a much larger data frame with many more potential keys.

Two questions:

  1. why does 'NaN' generate a key error despite it being hashable?
  2. what is the best way to solve this? One possibility is to set 'NaN' values to another value like -999 in the original data frame?
2 Answers

You can use df["column"].map(dict)

>>> df["new_values"] = df["values"].map(repl)
>>> df
            values new_values
2021-06-01     1.0          A
2021-06-02    -1.0          C
2021-06-03     0.0          B
2021-06-04     NaN          D

I think the explanation has to do with the fact that the way that python determines if a key is in a dictionary is 1. hashing the key and checking it against dictionary and then 2. checking to make sure that the key it is looking for is the key that it found in the dictionary.

The problem is that although np.nan is np.nan returns True, np.float64(np.nan) is np.float(np.nan) returns False. Similarly, np.float64(np.nan) is np.nan returns False.

My guess is that the reason your apply function does not work is that the lambda function you create is trying to find np.float64(np.nan) (or something similar from your DataFrame) in the dictionary repl and not finding it. Even if your original data just contained np.nan, it seems like pandas converts it to numpy.float64 type.

E.g.

a = pd.DataFrame([[np.nan, 0], [1,1]])
a[0][0], type(a[0][0]), type(np.nan)
>> nan, numpy.float64, float

map on the other hand takes a dictionary as an argument and is specifically designed to handle the case where some values are missing or equal to np.nan ( see: https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.Series.map.html)

For more on using nan as keys in dictionaries see this question: NaNs as key in dictionaries

Related