Background
I've got this R dataframe, d:
d <- data.frame(ID = c("a","a","a","a","a","a","b","b"),
event = c("G12","G12","G12","B4","B4","A24","L5","L5"),
stringsAsFactors=FALSE)
It looks like this:
As you can see, it's got 2 distinct ID's in it, with each one having events, some of which repeat / are duplicated any number of times.
The Problem
I'd like to figure out what the average number of repeated event is per ID in this dataframe.
At a glance, you see that id= a has 2 events that repeat -- G12, which repeats twice (with 3 entries total) and B4, which repeats once (with 2 total entries). id= b has 1 event that repeats: L5. Note that how many times each repeat/duplicate occurs is irrelevant to me here; it just matters that there's at least one duplicate event per ID.
So the result I want is a simple tabulation of that mean:
(2 events that repeat + 1 event that repeats) / 2 people = 1.5
What I've tried
I've gotten kinda close thanks to posts like this, but I'm not quite there:
d %>% summarise(mean = mean(duplicated(event)))
This runs, but it doesn't factor in the fact that the duplication is occurring within ID (or at least, that's the way I see it).
