I want to find the mean of one numeric variable for each percentile of another numeric variable. To essentially replicate this graph (Marian et al(2012) but for my own data:
I have tried the following:
tapply(quantile(CLEARPOND$word_frequency, probs = c(.05, .10, .15, .20, .25,.30,.35,.40,.45,.50,.55,.60,.65,.70,.75,.80,.85,.90,.95)), CLEARPOND$Colthearts_N, mean)
which returns the following error:
Error in tapply(quantile(CLEARPOND$word_frequency, probs = c(0.5, 0.1, : arguments must have same length
Is there anyway to fix this/ do this in a more logical way?
I basically want to divide the variable word_frequency into bins of 5% increments. And then find the mean of Colthearts_N for each of those bins. I would also ideally like to plot this on a scatter plot.
My percentiles for word_frequency are as follows:
5% 10% 15% 20% 25% 30% 35% 40% 45% 50% 55% 60% 65% 70% 75% 80% 85% 90% 95% 1 1 1 1 1 1 1 1 1 1 1 1 2 2 2 3 4 6 11
Any help would be appreciated