I must check to a boolean value in rows, get the name of the columns to build a list and create a new column with this list of strings.
I have this code, it works perfectly but it's very slow (I used 125 000 rows for 20 columns). Do you have an idea to optimize this code?
import pandas as pd
def exlusions(x: pd.Series) -> pd.Series:
reasons = [name for name in x.index if x[name] == True]
return ",".join(reasons)
dct_data = {
'A' : [False, False, True, False, True, False],
'B' : [False, False, False,False, False, False],
'C' : [False, True, False,False, False, False],
'D' : [False, False, True, False, False, False],
'Client' : ['Paul', 'Nick', 'Josh', 'Flo', 'Julia', 'Lucia']
}
df = pd.DataFrame(dct_data)
df = df[list(df.select_dtypes(include='bool').columns) + ['Client']]
df = df[df[list(df.select_dtypes(include='bool').columns)].any(1)]
df['Exclusions'] = df.apply(lambda x: exlusions(x), axis=1)
df