My task is to remove outliers from the grouped data frame of approx. 3.500.000 rows and to perform a linear fit on each group after that. I have already found a way to speed up my linear fit by using the fastLmPure function from the RcppEigen package instead of lm from stats, which made my calculations like >100x faster.
On the other hand, the cooks.distance function I'm using only works together with lm or glm, which means that I can't use either fastLmPure or fastLm outputs to calculate outliers (at least directly).
This is the formula I'm using at the moment to calculate Cook's distance:
df %>%
mutate(cooks_dist = cooks.distance(lm(y ~ x + 1)))
The code above needs ~45 minutes to process the said data frame, while fastLmPure's fit is done within 15 seconds.