My dataframe look like the below:
# initialize list of lists
data = [['tom', 10], ['nick', 15], ['juli', 14],['tom', 10], ['juli', 15] ]
# Create the pandas DataFrame
df = pd.DataFrame(data, columns = ['Name', 'Age'])
Name Age
0 tom 10
1 nick 15
2 juli 14
3 tom 10
4 juli 15
I want to group by the 'Name', count the 'Age' and unique count of 'Age'.
Using pandas I got the result:
Age
count nunique
Name
juli 2 2
nick 1 1
tom 2 1
Pandas code :
types = ['count', 'nunique']
df.groupby('Name').agg({'Age': types})
How can i achieve this in Dask?
In dask, I can do either count or nunique...
ddf = daskdf.from_pandas(df, npartitions=4)
ddf.groupby('Name').Age.count().to_frame().compute()
Age
Name
nick 1
tom 2
juli 2