Let's say I have a dataset from a regular school in which students from different living areas are tested in math, English, and science. You need to do a retest if your score is 1SD below the mean and you'll fail if your score is 2SD below the mean.
I can easily compute the means, standard deviation, and these cutoffs. I'm using the nest from the tidyverse package. However, I would like to discover how many students were 1SD below and 2SD below the mean.
However, I don't know how to do these count calculations to these results in an easy way.
Please check the dataset and the code I'm using to achieve the descriptive results:
library(tidyverse)
set.seed(123)
ds <- data.frame(quest = c(2,4,6),
living_area = c("rural","urban","mixed"),
math_sum = rnorm(120, 10,1),
english_sum = rnorm(120, 10,1),
science_sum = rnorm(120, 10,1)
)
ds %>%
select(quest, ends_with("sum")) %>% #get variable names
pivot_longer(-quest) %>% #tranform into long format
nest_by(quest, name) %>% #nest
mutate(
n = map_dbl(data, ~nrow(data.frame(.))), #compute sample size
mean = map_dbl(data, ~mean(.)), #get the means
sd = map_dbl(data, ~sd(.)), #get sd
below = mean-sd, #1 below
failed = mean-2*sd)
ds %>%
filter(quest == 2 & english_sum <= 9.19) %>% nrow()
ds %>%
filter(quest == 2 & english_sum <= 9.39) %>% nrow()
ds %>%
filter(quest == 2 & english_sum <= 8.73) %>% nrow()
