import pandas as pd
df = pd.read_csv('chim_work.csv')
df_col = df[['ID #','Init Acct Type','Subs Acct Type','Max Days Diff']]
df_drop_null = df_col.dropna()
df_group = df_drop_null.groupby('ID #')
for i, d in df_group:
dfn = d.drop(columns=['ID #'])
print(i)
print(dfn)
This code gives me my ID#s attached to 3 column DataFrames.
I want to figure out which DataFrames have duplicates, the id# of the duplicates and the count. Then create new labels for them.
So the output would be:
A
5 Duplicates
Id: 101, 102, 105, 107, 120