I have a df such as
mydata <- data.frame(variable = runif(142),
block = sample(x = c(1,2,3,4,5), 142, replace = TRUE))
I'm trying to sample 80 and 20% of each block value, without repeating, and adding each fraction to new dfs called train (80%) and test (20%). Important: Sometimes my blocks will not have exactly 80-20, but I'm trying to get as close as possible to this value.
How to proceed?
I was using sample_frac but wasn't able to avoid repeating and joining the data after.