Filling in the missing value does not work properly in Python

Viewed 56

I have a dataset that has a number of numeric variables and a number of ordinal nominal variables. To fill in the missing value I want to use the mode for nominal variables. The following code does not fill some of the nominal value. Please advise why the code is wrong.

df = pd.read_csv(sample.csv')
nominal_data = df.select_dtypes(include=[np.object])
nominalColumns= list(set(nominal_data.columns))
df[nominalColumns]=df[nominalColumns].fillna(df[nominalColumns].mode())


age | class
------------
 1 |  no
 2 |  yes
 3 |  NAN
 4 |  yes
 5 |  no
 6 |  NAN
 7 |  no
 8 |  yes
 9 |  no
10 |  NAN
2 Answers

Since your column can only store two values, why don't use extract the value to use from the mode ? Providing a Series object (which is the return of .mode()) may not work as (from my experience), Series objects are expected to provide the replacement value for each index (which is not the case with .mode() function).

Also, since you want to replace values over a column, use axis=1.

Maybe you could try :

df[nominalColumns]=df[nominalColumns].fillna(df[nominalColumns].mode().iloc[0], axis=1)

DataFrame.mode returns a DataFrame (see the docs) because each column may have several modes. When you call df.fillna(x) where df and x are both DataFrames with length n and m, only the first m rows of df will be filled with the corresponding values in x. Thus you need to call .iloc[0] after .mode() to get the first row as a Series. See example below.

import pandas as pd

df = pd.DataFrame({
    "a": ["foo", "bar", "foo", None, "foobar"],
    "b": ["foo", "bar", "bar", "foobar", None],
    "c": [None, None, None, "foo", "bar"],
    "d": [1, 2, 3, 4, 5]
})

nominal_data = df.select_dtypes(include=[object])
nominal_columns= list(set(nominal_data.columns))

print(df[nominal_columns].mode())

     b    c    a
0  bar  bar  foo
1  NaN  foo  NaN

Column c has two modes (foo and bar), hence the two rows. Columns a and b only have one mode, hence the NaN on the second row.

Using only the first row of modes fills NAs as you expect:

print(df[nominal_columns].fillna(df[nominal_columns].mode().iloc[0]))

        b    c       a
0     foo  bar     foo
1     bar  bar     bar
2     bar  bar     foo
3  foobar  foo     foo
4     bar  bar  foobar
Related