Pandas in Python and Dplyr in R are both flexible data wrangling tools. For example, in R, with dplyr one can do the following;
custom_func <- function(col1, col2) length(col1) + length(col2)
ChickWeight %>%
group_by(Diet) %>%
summarise(m_weight = mean(weight),
var_time = var(Time),
covar = cov(weight, Time),
odd_stat = custom_func(weight, Time))
Notice how in one statement;
- I can aggregate over multiple columns in one line.
- I can apply different functions over these multiple columns in one line.
- I can use functions that take into account two columns.
- I can throw in custom functions for any of these.
- I can declare new column names for these aggregations.
Is such a pattern also possible in pandas? Note that I am interested in doing this in a short statement (so not creating three different dataframes and then joining them).