I'm trying to sample from a data frame but with the condition, that the sample represents the distribution in terms of a certain criterion (in my case. The data frame is structured like this:
df <- data.frame(Locaton = c(A, B, B, B, C, C, ...),
Veg_Species = c(X, Y, Z, Z, Z, Z...),
Date_Diff = c(2, 5, 2, 0, 4, 4...))
It is important to know, that the number of a Veg_Species differs. That means X has 25 occurrences, Y 45 and Z 78 for example. And now I want to sample from the different Veg_Species based on the distribution in terms of Date_Diff of the smallest sample. In that case that would mean sampling from every species in terms of Date_diff distribution from X.
I thought that I can do that with dplyr:
sample.species <- df %>% filter(Veg_Species == 'Z') %>% sample_n(25, replace = TRUE)
But that obviously only samples randomly from all Veg_Species with the name Z.
How can I take the distribution into account too?
For a more detailed example, click here.