Dummy data:
set.seed(4)
name <- sample(LETTERS[1:8], 500, replace = T)
id <- round(runif(500, min=1, max=200))
df <- data.frame(name, id)
I want to check the % of unique id of B which are there for other remaining name
The expected output will be something like this:
name count pct_common
<chr> <int> <dbl>
1 A 17 29.3
2 C 18 31.0
3 D 16 27.6
4 E 22 37.9
5 F 14 24.1
6 G 16 27.6
7 H 20 34.5
My approach so far:
the_name <- 'B'
#Selecting the unique name, id combination for 'B'
df %>%
filter(name %in% the_name) %>%
distinct(name, id)-> list_id
#Checking which of these ids are already there for other names and then count them.
df %>%
filter( id %in% list_id$id) %>%
filter(!name %in% the_name) %>%
group_by(name) %>%
summarise(count=n()) %>%
mutate(pct_common= count/nrow(list_id)*100)
It is getting the job done but creating a separate data frame like this doesn't seem very elegant. Also, it is taking more time for a comparatively large data frame (Millions of observations).
Is there a better way to approach this problem?