Boolean masking a list of different lengths within each individual row

Viewed 81

I have the following dataframe:

df = pd.DataFrame({
    'tags': [
        [{'id': 1401}, {'id': 1801}],
        [{'id': 502}, {'id': 703}, {'id': 1801}],
        [{'id': 1801}]
    ]
})

I am only interested in the 'id': 1801 value in the 'tags' column and would like to create a new column which contains True if 'id': 1801 is present or False if it is not.

Any help would be greatly appreciated

2 Answers

We can explode the tags column then using str accesor get the value of id and compare it with 1801 to create a boolean mask followed by any on level=0 to reduce:

df['flag'] = df['tags'].explode().str['id'].eq(1801).any(level=0)

If dataframe is big and performance needs to be considred then we can use list comprehension which will outperform all of the available pandas based solution

df['flags'] = [any(d['id'] == 1801 for d in l) for l in df['tags']]

                                       tags  flag
0              [{'id': 1401}, {'id': 1801}]  True
1  [{'id': 502}, {'id': 703}, {'id': 1801}]  True
2                            [{'id': 1801}]  True

Use lambda function with any if performance is important with get for test also if id missing:

df = pd.DataFrame({
    'tags': [
        [{'id': 1401}, {'id': 1801}],
        [{'id': 502}, {'id': 703}, {'id': 1801}],
        [{'id': 1}]
    ]
})

df['new'] = df['tags'].apply(lambda x: any(y.get('id') == 1801 for y in x))
print (df)
                                       tags    new
0              [{'id': 1401}, {'id': 1801}]   True
1  [{'id': 502}, {'id': 703}, {'id': 1801}]   True
2                               [{'id': 1}]  False

df = pd.concat([df] * 1000, ignore_index=True)

In [275]: %timeit df['tags'].explode().str['id'].eq(1801).any(level=0)
8.09 ms ± 265 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)

In [276]: %timeit df['tags'].apply(lambda x: any(y.get('id') == 1801 for y in x))
2.64 ms ± 6.3 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)

If always id exist per tags:

In [283]: %timeit [any(d['id'] == 1801 for d in l) for l in df['tags']]
2.44 ms ± 215 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
Related