I have a function which finds a sort of intersection of two strings of prose:
# Function to get intersection of words
str_intersect_by_word_list <- function(string1, string2){
map2_chr(str_split(string1, '\\s'), str_split(string2, '\\s'),
~str_c(intersect(.x, .y), collapse = " "))
}
and a table with strings to match on:
# Sample data
my_df <- tibble(
grp = rep(LETTERS[1:3], each = 3),
strng = c(
"Hi I'm Abe",
"Hi I'm Beau",
"Hi I'm Cat",
"Hi there I'm Doug",
"Hi there I'm Emily",
"Hi there I'm Finn",
"Hi it's nice to be here",
"Hi it's nice to meet you",
"Hi it's nice outside"
)
)
If I want to create a column with the common string, I can do so like this:
# This works as expected
my_df %>%
mutate(
common_string = my_df %>%
pull(strng) %>%
reduce(str_intersect_by_word_list)
)
which gives
# A tibble: 9 x 3
grp strng common_string
<chr> <chr> <chr>
1 A Hi I'm Abe Hi
2 A Hi I'm Beau Hi
3 A Hi I'm Cat Hi
4 B Hi there I'm Doug Hi
5 B Hi there I'm Emily Hi
6 B Hi there I'm Finn Hi
7 C Hi it's nice to be here Hi
8 C Hi it's nice to meet you Hi
9 C Hi it's nice outside Hi
I would like to create a column with the string that is common to each group. However, inside the grouping, I can only access either the entire columns worth of strng, which gives the same output as above, or the current value of strng, which causes an error since my function str_intersect_by_word_list expects two inputs.
I tried referencing cur_data_all, but I don't think this is how that function was intended to be used, and it gives me the same results as above anyway.
# This fails.
my_df %>%
group_by(grp) %>%
mutate(
grp_string = cur_data_all() %>%
pull(strng) %>%
reduce(str_intersect_by_word_list)
)
The expected output would be
# A tibble: 9 x 3
grp strng grp_string
<chr> <chr> <chr>
1 A Hi I'm Abe Hi I'm
2 A Hi I'm Beau Hi I'm
3 A Hi I'm Cat Hi I'm
4 B Hi there I'm Doug Hi there I'm
5 B Hi there I'm Emily Hi there I'm
6 B Hi there I'm Finn Hi there I'm
7 C Hi it's nice to be here Hi it's nice
8 C Hi it's nice to meet you Hi it's nice
9 C Hi it's nice outside Hi it's nice
How can I get the common words from each group?