Let's say we have this dataframe in Pandas. In my case, I got it as the result of a pivot_table() with the aggfunc as lambda x: x, because list or set didn't do the right thing with this kind of data.
import pandas as pd
df = pd.DataFrame(
data=[
[None, "1,2,3", None],
["3,4,5", None, "1,4,5"],
[None, "1,3,6", None],
],
index=["YYZ", "YEG", "BRU"],
columns=["ANA", "JAL", "KLM"],
)
df
I'd like to parse it to change the strings with commas in between to sets. I managed to do it naïvely like this:
for column in df.columns:
nulls = df[column].isnull()
for idx in df.loc[nulls, column].index:
df.at[idx, column] = set()
for idx in df.loc[~nulls, column].index:
df.at[idx, column] = set(df.at[idx, column].split(","))
df
which gives this:
ANA JAL KLM
YYZ {} {3, 2, 1} {}
YEG {5, 4, 3} {} {5, 4, 1}
BRU {} {6, 3, 1} {}
What's the right way of doing it in Pandas?