Problem
I want to selectively split some strings in a column. My limited ability with regular expressions is a problem, but your thoughts would be welcome.
Reprex
My starting dataset:
media <- c("TV, Radio and Games", "Theatre, Festival", "Festival and Arts", "TV", "Theatre", "Radio")
type <- c("Indoors", "Outdoors", "Outdoors", "Indoors", "Outdoors", "Indoors")
df <- as.data.frame(cbind(media, type))
df
# media type
# 1 TV, Radio and Games Indoors
# 2 Theatre, Festival Outdoors
# 3 Festival and Arts Outdoors
# 4 TV Indoors
# 5 Theatre Outdoors
# 6 Radio Indoors
Output Wanted
media2 <- c("TV", "Radio", "Games", "Theatre", "Festival", "Festival and Arts", "TV", "Theatre", "Radio")
type2 <- c("Indoors", "Indoors", "Indoors", "Outdoors", "Outdoors", "Outdoors", "Indoors", "Outdoors", "Indoors")
df2 <- cbind(media2, type2)
df2
# media2 type2
# [1,] "TV" "Indoors"
# [2,] "Radio" "Indoors"
# [3,] "Games" "Indoors"
# [4,] "Theatre" "Outdoors"
# [5,] "Festival" "Outdoors"
# [6,] "Festival and Arts" "Outdoors"
# [7,] "TV" "Indoors"
# [8,] "Theatre" "Outdoors"
# [9,] "Radio" "Indoors"
Attempted Solution
df3 <- df %>%
tidyr::separate(media, c("media1", "media2"), ", ") %>%
tidyr::pivot_longer(1:2, names_to = "media_type", values_to = "media") %>%
select(media, type) %>%
filter(!is.na(media))
# media type
# 1 TV Indoors
# 2 Radio and Games Indoors
# 3 Theatre Outdoors
# 4 Festival Outdoors
# 5 Festival and Arts Outdoors
# 6 TV Indoors
# 7 Theatre Outdoors
# 8 Radio Indoors
This gets me some of the way there, but I also want "Radio and Games" split up, but not "Festival and Arts". I cannot split it using " and ", as that would split them both.