I have data with the following structure, where each responder is assigned a task that could either have a status being TRUE or FALSE.
month Responder Status Department
2020-02-01 A TRUE 1
2020-02-01 B FALSE 1
2020-02-01 B TRUE 1
2020-02-01 C TRUE 1
2020-02-01 C TRUE 1
2020-03-01 D FALSE 2
2020-03-01 E FALSE 1
2020-03-01 B FALSE 1
2020-03-01 F FALSE 2
2020-03-01 F TRUE 2
2020-03-01 F TRUE 2
I want to output a data frame so that each responder is given a probability of having Status = FALSE. I would like to group these results by month and department as follows:
month Responder Prob_False N n
2020-02-01 A 0 1 0
2020-02-01 B 0.5 2 1
2020-02-01 C 0 2 0
2020-03-01 B 1 1 1
2020-03-01 D 1 1 1
2020-03-01 E 1 1 1
2020-03-01 F 0.333 3 1
Where N is the total number of tasks assigned to the responder for that month and n is the number of tasks that had a FALSE status, grouped by the month and the responder.
I am trying to use the group_by and summarize functions in dplyr, but I guess I am not grasping the correct application for this particular problem.