You are explicitly calling for 1 random number each time: runif(1, ...). Instead, use runif(n(), ...). Realize that it isn't called once for each row, it is run once for all rows that meet that condition. In my example below, there are three rows in May, but runif is called as runif(1,..) and that single number is applied to all three rows.
Sample data:
set.seed(42)
df <- data.frame(day = as.Date("2022-01-01") + sample(364, size=10)) %>%
arrange(day) %>%
mutate(month = as.POSIXlt(day)$mon + 1L)
df
# day month
# 1 2022-02-19 2
# 2 2022-03-16 3
# 3 2022-05-03 5
# 4 2022-05-09 5
# 5 2022-05-27 5
# 6 2022-06-03 6
# 7 2022-08-17 8
# 8 2022-10-31 10
# 9 2022-11-18 11
# 10 2022-12-31 12
Broken:
library(dplyr)
set.seed(42)
df %>%
mutate(
new_day = case_when(
month == 2 ~ floor(runif(1, 1, 28)),
month %in% c(9, 4, 6, 11) ~ floor(runif(1, 1, 30)),
TRUE ~ floor(runif(1, 1, 31))
)
)
# day month new_day
# 1 2022-02-19 2 25
# 2 2022-03-16 3 9
# 3 2022-05-03 5 9
# 4 2022-05-09 5 9
# 5 2022-05-27 5 9
# 6 2022-06-03 6 28
# 7 2022-08-17 8 9
# 8 2022-10-31 10 9
# 9 2022-11-18 11 28
# 10 2022-12-31 12 9
To demonstrate that runif is being called once for all rows that meet each criterion, I'll add message to each. If we could rely on runif(1,..), then we should see "30d" printed to the console 7 times and "31d" twice, but we don't.
set.seed(42)
df %>%
mutate(
new_day = case_when(
month == 2 ~ { message("Feb: ", length(month)); floor(runif(1, 1, 28)); },
month %in% c(9, 4, 6, 11) ~ { message("30d: ", length(month)); floor(runif(1, 1, 30)); },
TRUE ~ { message("31d: ", length(month)); floor(runif(1, 1, 31)); }
)
)
# Feb: 10
# 30d: 10
# 31d: 10
# day month new_day
# 1 2022-02-19 2 25
# 2 2022-03-16 3 9
# 3 2022-05-03 5 9
# 4 2022-05-09 5 9
# 5 2022-05-27 5 9
# 6 2022-06-03 6 28
# 7 2022-08-17 8 9
# 8 2022-10-31 10 9
# 9 2022-11-18 11 28
# 10 2022-12-31 12 9
This demonstrates that when we're 'inside' the RHS of one of the conditions, it is a call for all rows of the frame. Notice that each time we call runif, it sees all values of month (we have 10 rows in df).
Instead, use n() (number of rows in each call):
set.seed(42)
df %>%
mutate(
new_day = case_when(
month == 2 ~ floor(runif(n(), 1, 28)),
month %in% c(9, 4, 6, 11) ~ floor(runif(n(), 1, 30)),
TRUE ~ floor(runif(n(), 1, 31))
)
)
# day month new_day
# 1 2022-02-19 2 25
# 2 2022-03-16 3 5
# 3 2022-05-03 5 30
# 4 2022-05-09 5 29
# 5 2022-05-27 5 3
# 6 2022-06-03 6 28
# 7 2022-08-17 8 12
# 8 2022-10-31 10 28
# 9 2022-11-18 11 14
# 10 2022-12-31 12 26
This means that we pull 30 random numbers in this case_when, 10 for each condition. While this is not a concern here (large draws on entropy can be slow), you can mitigate by pre-pulling random data and then scaling accordingly.
set.seed(42)
df %>%
mutate(
rand = runif(n(), 0, 1),
new_day = case_when(
month == 2 ~ ceiling(rand*28),
month %in% c(9, 4, 6, 11) ~ ceiling(rand*30),
TRUE ~ ceiling(rand*31)
)
)
# day month rand new_day
# 1 2022-02-19 2 0.9148060 26
# 2 2022-03-16 3 0.9370754 30
# 3 2022-05-03 5 0.2861395 9
# 4 2022-05-09 5 0.8304476 26
# 5 2022-05-27 5 0.6417455 20
# 6 2022-06-03 6 0.5190959 16
# 7 2022-08-17 8 0.7365883 23
# 8 2022-10-31 10 0.1346666 5
# 9 2022-11-18 11 0.6569923 20
# 10 2022-12-31 12 0.7050648 22
(noting the shift from floor to ceiling). There are further ways the code can be refactored, but I think this is generally good-enough.