I am trying to apply for loops inside a Pandas dataframe to access two columns at a time. My piece of code works perfectly for a single column. But when applying to multiple columns, it is throwing : "ValueError : too many values to unpack (expected 2)"
My code snippet is as follows -
for col1, col2 in df.columns:
if col1.startswith('ColumnName1') and col2.startswith('ColumnName2') and df[col2].notnull()*1:
new_df = df.groupby([col1, col2]).agg({'ColumnName3': 'unique'}).reset_index()
elif col1.startswith('ColumnName1') and col2.startswith('ColumnName2') and not df[col2].notnull()*1:
new_df = df.groupby(col1).agg({'ColumnName3': 'unique'}).reset_index()
The small problem is the column names are too large and not under control, because this dataframe has multiheader columns, so after merging they are creating some random filling names. Hence the ".startswith". The column names are much larger.
I am trying to perform a groupby of column 3 based on columns 1 and 2, if column 2 is not null, else a groupby using column1 when column 2 is null.
Can anyone tell me where am I wrong here, or what am I missing here?