If you’re ever curious about what’s going on with a package, I would encourage you to have a quick look through the github. It may be a little overwhelming at first, but you can learn quite a bit by reviewing what others have done.
Looking at the yarrr github, we can see in line 750 the package calculates confidence intervals using a t-interval. You are calculating the 95% confidence interval using a z-interval. The methods are different, so we’ll get different results. For smaller sample sizes (less than 30 as a general rule), t-intervals are more appropriate.
The z-interval assumes we know the population standard error. For smaller sample sizes, you can imagine our estimates are going to be pretty varied and the standard deviation of the sample may not be representative of the population. The confidence intervals that result don’t account for this extra uncertainty and so they don’t have enough coverage to truly be a 95% confidence interval as we're violating the assumption that the population standard error is known.
Let’s do the calculation with your example data just to see the difference and make sure we are matching the package.
library(yarrr)
library(tidyverse)
df <- data.frame(x = c(1,2,3,4,5))
## t-interval uses n - 1 degrees of freedom
dof <- nrow(df) - 1
t <- qt(0.975, dof)
## z-interval. ~1.96 as you've used
z <- qnorm(0.975)
longdata <- df %>%
pivot_longer(cols = everything(), names_to = 'x', values_to = 'y')
datasum <- longdata %>%
summarize(avg = mean(y),
standard_dev = sd(y),
n = n(),
sem = standard_dev/sqrt(n),
## z-interval
ci_ll = avg - z*sem,
ci_ul = avg + z*sem,
## t-interval
lower = avg - t*sem,
upper = avg + t*sem)
datasum
#> # A tibble: 1 x 8
#> avg standard_dev n sem ci_ll ci_ul lower upper
#> <dbl> <dbl> <int> <dbl> <dbl> <dbl> <dbl> <dbl>
#> 1 3 1.58 5 0.707 1.61 4.39 1.04 4.96
So we can see the difference between the Z- and t-interval confidence intervals. The t-interval (appropriately) has more coverage. Does our calculation match the t-test?
t.test(longdata$y, conf.level = 0.95)$conf.int
#> [1] 1.036757 4.963243
#> attr(,"conf.level")
#> [1] 0.95
It does indeed. If we take a look at pirateplot, we’ll see our t-interval calculations match the plot.
yarrr::pirateplot(data=longdata, y ~ x,
inf.method ='ci',
bean.b.col = "white",
bean.f.col = "white")

Created on 2022-02-22 by the reprex package (v2.0.1)