I have a tibble with a lot of groups, and I want to do group-wise operations on it (highly simplified mutate below).
z <- tibble(k1 = rep(seq(1, 600000, 1), 5),
category = sample.int(2, 3000000, replace = TRUE)) %>%
arrange(k1, category)
t1 <- z %>%
group_by(k1) %>%
mutate(x = if_else(category == 1 & lead(category) == 2, "pie", "monkey")) %>%
ungroup()
This operation is very slow, but if I instead do grouping "manually", the process is hard to read, more annoying to write, but much (20x) faster.
z %>%
mutate(x = if_else(category == 1 & lead(category) == 2 & k1 == lead(k1), "pie", "monkey"),
x = if_else(category == 1 & k1 != lead(k1), NA_character_, x))
So clearly there is some way with keys to speed up the process. Is there a better way to do this? I tried with data.table, but it was still much slower than the manual technique.
zDT <- z %>% data.table::as.data.table()
zDT[, x := if_else(category == 1 & lead(category) == 2, "pie", "monkey"), by = "k1"]
Any advice for a natural, fast way to do this operation?