Filter dataframe that has list as column values with external list and delete outsiders

Viewed 384

I've got a dataframe like this

Sample_ID   Main_Sample_ID
1ABC        [2052, 2402]   
2CBA        [228]  

and an external list with allowed values:

allowed = [2402]

What I'm trying to do is filtering those rows which have allowed values and deleting those which don't, deleting either the internal list values that are not allowed too.

At the end, I'd like to get the result:

Sample_ID   Main_Sample_ID
1ABC        [2402]   

I tried it with:

sample_type_ids_list = self._full_structure['Main_Sample_ID'].tolist()
for sample_type_ids in sample_type_ids_list:
    for sample_type_id in sample_type_ids:
        info_by_type_df['flag'] = info_by_type_df.apply(lambda x: int(sample_type_id in allowed), axis=1)

I also tried with .loc and .isin() but without success...

Could you help me? Thanks in advance!

3 Answers

You can keep the items in the allowed list as follows, and then drop empty lists.

# change list in every row to empty if id are not present in `allowed`
# if in allowed list, then keep it
df = df.apply(lambda row: [id for id in row['Main_Sample_ID'] if id in allowed], axis=1)

# drop rows with empty lists
df = df[df.apply(len) > 0]

You can assign a list comprehension. This is only superficially a Pandas question because your current data structure only permits Python-level loops:

df = pd.DataFrame({'Sample_ID': ['1ABC', '2CBA'],
                   'Main_Sample_ID': [[20152, 2402], [228]]})

df['Main_Sample_ID'] = [[i for i in lst if i == 2402] for lst in \
                        df['Main_Sample_ID'].values.tolist()]

df = df[df['Main_Sample_ID'].str.len() > 0]

print(df)

  Main_Sample_ID Sample_ID
0         [2402]      1ABC

Using custom function with numpy arrays:

def func(values):
    l = np.array(values)[np.isin(values,allowed)]
    if l.size>0:
        return l
        #if list require return l.tolist()
    else:
        return np.nan

df.Main_Sample_ID = df.Main_Sample_ID.apply(func)
df = df.dropna()

print(df)
  Sample_ID Main_Sample_ID
0      1ABC         [2402]
Related