I've got a dataset that looks like this:
id date outcome
ny21 2021-03-01 0
ny21 2021-02-01 1
ny21 2021-01-01 0
ch67 2021-08-09 0
I can calculate the time difference between rows grouped by id in the following way:
unlist(tapply(df$date, INDEX = df$id, FUN = function(x) c(0, `units<-`(diff(x), "days"))))
However, I also want to calculate the time difference between rows grouped by id, where outcome == 1. If the current row has outcome == 1, and no such previous row exists, I want to put 0. So, something like this:
id date outcome days_since_outcome_1
ny21 2021-03-01 0 28
ny21 2021-02-01 1 0
ny21 2021-01-01 0 0
ch67 2021-08-09 0 0
How can I do this?
EDIT:
In some situations, there are more than one outcome == 1 per id. In this case, I would like the code to calculate the difference since the latest match (not the first match). Something like this:
id date outcome days_since_outcome_1
ny21 2021-05-01 1 30 # this is days since 2021-04-01
ny21 2021-04-01 1 59 # this is days since 2021-02-01
ny21 2021-03-01 0 28 # this is days since 2021-02-01 (first match)
ny21 2021-02-01 1 0
ny21 2021-01-01 0 0
ch67 2021-08-09 0 0